Method and system for generating three-dimensional moving object based on artificial intelligence
By using artificial intelligence technology to generate three-dimensional motion images from two-dimensional images, the problems of object motion conversion and occlusion recovery are solved, and accurate motion similarity calculation and text-driven motion image generation are achieved.
Patent Information
- Application Number
- CN202380094637.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-09-14
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies lack the ability to convert the movements of objects in two-dimensional images into three-dimensional motion images. In particular, it is difficult to restore the complete appearance when the object is occluded or cut, and the movement similarity calculation is not accurate enough. There is a lack of text-to-motion image generation methods.
Through an AI-based method, convolutional neural networks and recurrent neural networks are used to extract the posture and shape data of objects from two-dimensional images to generate three-dimensional motion images. The occluded or cut parts are restored through a three-dimensional mesh model, and metric learning is combined to calculate action similarity and analyze text to generate motion images.
It realizes the conversion from two-dimensional images to three-dimensional motion images, can automatically restore occluded or cut body parts, accurately calculate action similarity, and generate corresponding motion images based on text.
Smart Images

Figure CN120752677A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and device for generating a three-dimensional moving object based on artificial intelligence, and more specifically, to a method and device for generating a three-dimensional moving image using a two-dimensional image based on artificial intelligence.
[0002] Furthermore, the present invention relates to a method and apparatus for generating an image of an object. More specifically, it relates to a method and apparatus for generating an image of an object that can completely restore the overall appearance of the object by inferring occluded or cropped portions of the object based on artificial intelligence.
[0003] Furthermore, the present invention relates to a method and apparatus for calculating motion similarity of objects based on artificial intelligence, and more specifically, to a method and apparatus for calculating motion similarity between objects based on three-dimensional grid objects.
[0004] Furthermore, the present invention relates to a method and apparatus for generating a moving image from a sentence, and more particularly, to a method and apparatus for analyzing a sentence in a text format and automatically generating a moving image corresponding to the meaning of the sentence. Background Art
[0005] As the virtual reality market grows, demand for 3D (three-dimensional) image generation technology, a core element of the virtual reality market, is rapidly increasing. Companies both domestically and internationally view virtual reality, which recreates the real world within a virtual world, as a next-generation growth engine. Interest in 3D imaging technology, a core technology for virtual reality—converting the real world into 3D and presenting it within the virtual world—is also growing.
[0006] In the past, 3D imaging technology was mainly used in the gaming field, but in recent years, its application trend has been to expand to all industrial fields, and its application scope has expanded to various content industries such as AR / VR, movies, animation, and broadcasting.
[0007] However, despite the high level of interest in 3D imaging technology, its technical sophistication is still somewhat lacking. For example, companies building AI humans are focusing solely on developing technologies for generating virtual human faces using CG or deep fake generation methods, but lack the motion reproduction technology to recognize and reproduce natural movements.
[0008] In addition, with the widespread application of 3D imaging technology in all industrial fields, the market for 3D creation tools is also growing continuously. However, current 3D imaging creation tools only focus on creating static objects (buildings, interior decoration, facilities, etc.), and the technology research and development for converting dynamic objects into 3D and expressing them in three-dimensional space is still imperfect. Summary of the Invention
[0009] Technical issues
[0010] The technical problem to be solved by the embodiments of the present invention is to provide a method and device for generating three-dimensional motion images based on artificial intelligence, which can extract the movement of an object from a two-dimensional image based on artificial intelligence and realize it as a three-dimensional motion image.
[0011] Another technical problem to be solved by the embodiments of the present invention is to provide a method and device for generating three-dimensional motion images based on artificial intelligence, which can generate three-dimensional mesh data of an object using a two-dimensional image of the object based on artificial intelligence.
[0012] The technical problem to be solved by the embodiments of the present invention is to provide a method and device for generating an image of an object, which can automatically restore the occluded or cut body part by reasoning about the whole body of the object based on artificial intelligence, even if the body part of the object in the image is occluded or cut.
[0013] Another technical problem to be solved by an embodiment of the present invention is to provide a method and device for generating an image of an object, which utilizes two-dimensional data containing three-dimensional information, namely, a UV map, when restoring the occluded or cut body parts of the object. In this way, while utilizing the existing convolutional neural network-based model technology, the body parts can be restored more carefully and realistically.
[0014] The technical problem to be solved by the embodiments of the present invention is to provide the following artificial intelligence-based object motion similarity calculation method and device, which compares the posture and motion of the object by using a three-dimensional grid model, and can also take into account the physical characteristics of the object, thereby more accurately calculating the motion similarity.
[0015] Another technical problem to be solved by the embodiments of the present invention is to provide the following artificial intelligence-based object action similarity calculation method and device, which uses an artificial intelligence module for machine learning based on metric learning and can also perform similarity analysis on unlearned postures and actions.
[0016] The technical problem to be solved by the embodiments of the present invention is to provide a method and apparatus for automatically generating a motion image corresponding to the meaning of a sentence by analyzing a sentence provided in a text format.
[0017] Another technical problem to be solved by an embodiment of the present invention is to provide a method and device for generating a motion image from a sentence, which generates an overall motion image by combining partial images corresponding to the meanings of words in the sentence, and can insert appropriate intermediate images to make the connection between the partial images natural.
[0018] Workaround
[0019] In order to solve the technical problem, a method for generating three-dimensional motion images based on artificial intelligence is performed by a computing device and may include: acquiring a two-dimensional image; using an artificial intelligence module that has been machine-learned to identify an object from the two-dimensional image, and analyzing the movement of the identified object to generate three-dimensional modeling data of the object; and generating a three-dimensional motion image using the three-dimensional modeling data.
[0020] As an embodiment, the step of generating the three-dimensional modeling data may include: extracting posture data of the object from the two-dimensional image using a first artificial intelligence module, and generating three-dimensional motion data representing the motion of the object based on the posture data.
[0021] As one embodiment, the two-dimensional image may include multiple frames representing the movement of the object in a time series, and the first artificial intelligence module may include: a first submodule based on a convolutional neural network, which extracts posture data of the object from a first frame of the multiple frames; and a second submodule based on a recurrent neural network, which provides information about a second frame of the multiple frames to the first submodule, wherein the second frame may be a frame that is located before the first frame in the time series.
[0022] As an embodiment, the step of generating the three-dimensional motion image may include: generating a motion image representing the motion of the three-dimensional character object by merging the three-dimensional motion data into a pre-stored three-dimensional character object.
[0023] As an embodiment, the three-dimensional motion data may include: key point data corresponding to the joints of the object.
[0024] As an embodiment, the step of generating a motion image representing the motion of the three-dimensional character object may include a retargeting step of matching joint positions of the three-dimensional motion data with joint positions of the three-dimensional character object based on the key point data.
[0025] As an embodiment, the step of generating the three-dimensional modeling data may further include: extracting shape information of the object from the two-dimensional image using a second artificial intelligence module, and generating a three-dimensional mesh object representing the volume of the object based on the shape information.
[0026] As an embodiment, the step of generating the three-dimensional motion image may include: merging the three-dimensional mesh object into the three-dimensional motion data to generate a motion image representing the motion of the three-dimensional mesh object.
[0027] As one embodiment, the second artificial intelligence module may include: a three-dimensional reconstruction module that generates a three-dimensional point cloud representing the object from the two-dimensional image; and a mesh modeling module that performs three-dimensional mesh modeling and texture mapping based on the three-dimensional point cloud to generate the three-dimensional mesh object.
[0028] In order to solve the above technical problems, an artificial intelligence-based three-dimensional motion image generation device may include a processor, a memory for loading a computer program executed by the processor, and a storage device for storing the computer program. The computer program may include instructions for performing the following operations: acquiring a two-dimensional image; extracting pose data of an object from the two-dimensional image using an artificial intelligence module that has been machine-learned, and generating three-dimensional modeling data based on the pose data; and generating a three-dimensional motion image using the three-dimensional modeling data.
[0029] In order to solve the above technical problems, a method for generating an object image based on artificial intelligence can be executed by a computing device and can include the following steps: acquiring a first object image and a UV map (UVmap) corresponding to the first object image; and generating a second object image in which the body part of the object is restored based on the first object image and the UV map, wherein the first object image can be an image in which the body part of the object is cropped, hidden or blocked.
[0030] As an embodiment, in the step of acquiring the UV map, the UV map may be generated from the first object image using a third artificial intelligence module that has been machine-learned.
[0031] As one embodiment, the third artificial intelligence module can be obtained by machine learning through the following method, which includes: a step of converting a three-dimensional mesh object corresponding to a two-dimensional object image to generate a first UV map corresponding to the two-dimensional object image; a step of acquiring a plurality of learning data based on the two-dimensional object image; a step of providing the plurality of learning data to the third artificial intelligence module, and the third artificial intelligence module outputting a second UV map corresponding to the learning data; and a step of updating the parameters of the third artificial intelligence module based on a comparison result of the first UV map and the second UV map.
[0032] As an embodiment, the three-dimensional mesh object may be generated by utilizing an artificial intelligence module that extracts the three-dimensional mesh object from the two-dimensional object image through machine learning.
[0033] As one embodiment, the step of generating the second object image may include: using a fourth artificial intelligence module to respectively generate a foreground image and a background image corresponding to the first object image; and generating the second object image by synthesizing the foreground image and the background image.
[0034] As one embodiment, the fourth artificial intelligence module may include: an encoder, extracting a feature vector based on the first object image and the UV map; a foreground generator, generating the foreground image using the first part of the feature vector; and a background generator, generating the background image using the second part of the feature vector, wherein the encoder, the foreground generator and the background generator may include a convolutional neural network.
[0035] As an embodiment, the fourth artificial intelligence module can be obtained by machine learning based on a generative adversarial network model.
[0036] As one embodiment, the fourth artificial intelligence module can be obtained by machine learning through the following method, which includes: a step of extracting a first feature vector based on a two-dimensional object image generated using the original image and a UV map corresponding to the two-dimensional object image; a step of generating a first foreground image corresponding to the two-dimensional object image using the first part of the first feature vector; a step of generating a first background image corresponding to the two-dimensional object image using the second part of the first feature vector; a step of generating a composite image based on the first background image and the first foreground image; and a step of updating the parameters of the fourth artificial intelligence module based on a result of comparing the first foreground image, the first background image or the composite image with the original image.
[0037] As an embodiment, the two-dimensional object image may be an image in which a portion of the original image is cut out or replaced by random noise.
[0038] In order to solve the above technical problems, an object image generation device based on artificial intelligence may include a processor, a memory loaded with a computer program executed by the processor; and a storage device storing the computer program, wherein the computer program may include instructions for performing the following operations: obtaining a first object image and a UV map corresponding to the first object image, and generating a second object image in which a body part of the object is restored based on the first object image and the UV map, wherein the first object image is an image in which the body part of the object is cut, hidden or blocked.
[0039] In order to solve the above technical problems, a method for calculating the action similarity of objects based on artificial intelligence according to an embodiment of the present invention can be executed by a computing device and can include the following steps: inputting a source image and a target image; extracting a first feature vector from a first object corresponding to the source image, and extracting a second feature vector from a second object corresponding to the target image; calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector, wherein the first object and the second object can be three-dimensional mesh objects, and the step of calculating the action similarity can be: using an artificial intelligence module that performs machine learning based on metric learning to calculate the action similarity.
[0040] As an embodiment, the first object may be a three-dimensional mesh object generated by inferring the volume of the object based on the object in the source image.
[0041] As an embodiment, in the extracting step, a Skinned Multi-Person Linear (SMPL) model may be used to extract the first feature vector and the second feature vector from the first object and the second object.
[0042] As an embodiment, the first feature vector may include: a grid vector, each grid vector representing a plurality of grids constituting the first object; and a trend vector representing a change trend of the grid vector.
[0043] As an embodiment, the closer the distance between the first feature vector and the second feature vector is, the higher the calculated action similarity is.
[0044] As an embodiment, the method may further include: providing the target image as a similar image retrieval result of the source image based on the action similarity.
[0045] As an embodiment, in the step of providing the search results, a plurality of target images including the target image may be provided in a listing based on their respective action similarities.
[0046] As an embodiment, the method may further include: based on the user's selection of the target image provided as a retrieval result, merging the three-dimensional motion data of the second object into the first object to generate a motion image of the first object moving according to the three-dimensional motion data.
[0047] As an embodiment, the method may further include: outputting a similarity score between the source image and the target image based on the action similarity.
[0048] In order to solve the above technical problems, according to an embodiment of the present invention, an object action similarity calculation device based on artificial intelligence may include: a processor, a memory for loading a computer program executed by the processor; and a storage device for storing the computer program, wherein the computer program may include instructions for performing the following operations: inputting a source image and a target image; extracting a first feature vector from a first object corresponding to the source image, and extracting a second feature vector from a second object corresponding to the target image; calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector, wherein the first object and the second object are three-dimensional mesh objects, and the operation of calculating the action similarity may be: calculating the action similarity using an artificial intelligence module that performs machine learning based on metric learning.
[0049] In order to solve the above technical problems, a method for generating a motion image from a sentence according to an embodiment of the present invention can be executed by a computing device and can include the following steps: obtaining a sentence; identifying one or more words in the sentence; retrieving partial images corresponding to the one or more words in a database; and generating a motion image corresponding to the sentence using the partial images.
[0050] As an embodiment, the step of identifying the one or more words may include: a step of identifying the part of speech of the one or more words.
[0051] As one embodiment, the one or more words may include a first word and a second word, and the step of retrieving from the database may include: a step of retrieving a first partial image corresponding to the first word from the database, and a step of retrieving a second partial image corresponding to the second word from the database.
[0052] As one embodiment, the one or more words also include a third word associated with the first word, and the step of retrieving the first partial image from the database may include: a step of retrieving multiple partial images from the database based on the first word, and a step of selecting the first partial image from the multiple partial images based on the third word.
[0053] As an embodiment, the step of generating the motion image may include the step of determining an arrangement order of the first partial image and the second partial image based on a context of the sentence.
[0054] As an embodiment, the first partial image may be an image located immediately before the second partial image, and the step of generating the motion image may include the step of generating an intermediate image to be inserted between the first partial image and the second partial image.
[0055] As an embodiment, the intermediate image may be an image used to connect the posture of the object in the first frame of the first partial image and the posture of the object in the second frame of the second partial image.
[0056] As one embodiment, the step of generating the intermediate image may include: a step of identifying a first position of a key point of the object from a first frame of the first partial image; a step of identifying a second position of the key point of the object from a second frame of the second partial image; and a step of generating the intermediate image so that the position of the key point of the object in the intermediate image changes from the first position to the second position.
[0057] As an embodiment, the step of generating the motion image may include the step of generating the motion image by sequentially connecting the first partial image, the intermediate image, and the second partial image.
[0058] As an embodiment, in the step of obtaining the sentence, the sentence may be obtained based on text input through a user interface.
[0059] As an embodiment, in the step of acquiring the sentence, the sentence may be acquired by converting a voice signal input through a voice input device into text.
[0060] In order to solve the above technical problems, according to an embodiment of the present invention, a device for generating a motion image from a sentence may include: a processor; a memory for loading a computer program executed by the processor; and a storage device for storing the computer program, wherein the computer program may include instructions for performing the following operations: obtaining a sentence; identifying one or more words in the sentence; retrieving partial images corresponding to the one or more words in a database; and generating a motion image corresponding to the sentence using the partial images.
[0061] Beneficial effects
[0062] According to the present invention, a three-dimensional moving image can be generated from a two-dimensional image.
[0063] Furthermore, according to the present invention, even if a body part of an object in an image is occluded or cut off, the occluded or cut off body part can be automatically restored by inferring the entire body of the object based on artificial intelligence.
[0064] Furthermore, according to the present invention, by comparing the postures and actions of the objects with each other using a three-dimensional mesh model, it is also possible to take into account the physical characteristics of the objects, thereby more accurately calculating the action similarity.
[0065] Furthermore, according to the present invention, by analyzing a sentence provided in a text format, a moving image corresponding to the meaning of the sentence can be automatically generated. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 FIG. 1 is a diagram illustrating an apparatus for generating a three-dimensional moving image based on artificial intelligence according to an embodiment of the present invention.
[0067] Figure 2 It shows Figure 1 A block diagram of an exemplary configuration of a three-dimensional moving image generating apparatus.
[0068] Figure 3 FIG. 4 is a flowchart illustrating a method for generating a three-dimensional motion image based on artificial intelligence according to an embodiment of the present invention.
[0069] Figure 4 It shows that Figure 3 The step S120 is a flowchart of the embodiment.
[0070] Figure 5 is a block diagram exemplarily showing a detailed configuration of the first artificial intelligence module 110 .
[0071] Figure 6 is a block diagram exemplarily showing a detailed configuration of the second artificial intelligence module 120 .
[0072] Figure 7 It shows that Figure 3 The step S130 is a flowchart of the embodiment.
[0073] Figure 8 It is shown that further Figure 7 The step S131 is a flowchart of the embodiment.
[0074] Figure 9 It shows that Figure 3 FIG. 1 is a flow chart of another embodiment of the present invention that concretizes step S130.
[0075] Figure 10 FIG. 1 is a diagram illustrating an exemplary configuration of an apparatus for generating an object image based on artificial intelligence inference of a portion of an object according to an embodiment of the present invention.
[0076] Figure 11 is a flowchart illustrating a method for generating an object image based on artificial intelligence reasoning of a portion of an object according to an embodiment of the present invention.
[0077] Figure 12 For further explanation Figure 11 FIG. 10 is a diagram of step S210.
[0078] Figure 13 It shows that Figure 11 The step S220 is a flowchart of the embodiment.
[0079] Figure 14 For further explanation Figure 13 FIG. 1 is an embodiment of the present invention.
[0080] Figure 15 is a flowchart exemplarily illustrating a learning method of the third artificial intelligence module.
[0081] Figure 16 For further explanation Figure 15 FIG. 1 is an embodiment of the present invention.
[0082] Figure 17 is a flowchart exemplarily illustrating a learning method of the fourth artificial intelligence module.
[0083] Figure 18 For further explanation Figure 17 FIG. 1 is an embodiment of the present invention.
[0084] Figure 19 FIG. 1 is a diagram illustrating an exemplary configuration of an apparatus for calculating object motion similarity based on artificial intelligence according to an embodiment of the present invention.
[0085] Figure 20 is a flowchart illustrating a method for calculating motion similarity of objects based on artificial intelligence according to an embodiment of the present invention.
[0086] Figure 21 and Figure 22 This is a diagram for explaining the characteristics of using a three-dimensional mesh model when comparing the similarity of the movements of objects.
[0087] Figure 23 and Figure 24 For further explanation Figure 20 FIG. 10 is a diagram of step S520.
[0088] Figure 25 and Figure 26 For further explanation Figure 20 FIG. 10 is a diagram of step S530.
[0089] Figure 27 is a flowchart of a method for calculating motion similarity of objects based on artificial intelligence according to another embodiment of the present invention.
[0090] Figure 28 For further explanation Figure 27FIG. 5 is a diagram of step S540.
[0091] Figure 29 is a flowchart illustrating a method for calculating motion similarity of objects based on artificial intelligence according to another embodiment of the present invention.
[0092] Figure 30 is a diagram showing an exemplary configuration of a moving image generating apparatus according to an embodiment of the present invention.
[0093] Figure 31 is a flowchart illustrating a method for generating a moving image from a sentence according to an embodiment of the present invention.
[0094] Figure 32 It shows that Figure 31 Step S620 is a flowchart of an embodiment.
[0095] Figure 33 For further explanation Figure 32 FIG. 1 is an embodiment of the present invention.
[0096] Figure 34 It shows that Figure 31 Step S630 is a flowchart of an embodiment.
[0097] Figure 35 For further explanation Figure 34 FIG. 1 is an embodiment of the present invention.
[0098] Figure 36 It shows that Figure 34 The step S631 is a flowchart of the embodiment.
[0099] Figure 37 For further explanation Figure 36 FIG. 1 is an embodiment of the present invention.
[0100] Figure 38 It shows that Figure 31 The step S640 is a flowchart of the embodiment.
[0101] Figure 39 and Figure 40 For further explanation Figure 38 FIG. 1 is an embodiment of the present invention.
[0102] Figure 41 It shows that Figure 38 The step S642 is a flowchart of the embodiment.
[0103] Figure 42 For further explanation Figure 41 FIG. 1 is an embodiment of the present invention.
[0104] Figure 43is a block diagram illustrating an exemplary hardware configuration of a computing device 500 for implementing various embodiments of the present invention. DETAILED DESCRIPTION
[0105] The following embodiments of the present invention will be described in detail with reference to the accompanying drawings. The advantages, features, and implementation methods of the present invention will become more apparent with reference to the accompanying drawings and the embodiments described in detail below. However, the technical concept of the present invention is not limited to the following embodiments, but can be implemented in various forms. The following embodiments are only used to improve the technical concept of the present invention and to fully inform ordinary technicians in the field to which the present invention belongs of the scope of the present invention. The technical concept of the present invention is only defined by the scope of the claims.
[0106] When adding reference symbols to components in the various drawings, it should be noted that the same symbols are used whenever possible even if the same components appear in different drawings. In addition, when describing the present invention, if it is determined that a detailed description of a related known structure or function may obscure the main purpose of the present invention, its detailed description will be omitted.
[0107] Unless otherwise defined, all terms (including technical terms and scientific terms) used in this specification can be used with the meaning commonly understood by those of ordinary skill in the art to which the present invention belongs. In addition, unless otherwise clearly defined, the terms defined in common dictionaries should not be idealized or over-interpreted. The terms used in this specification are intended to describe embodiments and are not intended to limit the present invention. In this specification, unless otherwise stated, singular forms also include plural forms.
[0108] In addition, when describing the components of the present invention, terms such as first, second, A, B, (a), (b), etc. may be used. These terms are only used to distinguish one component from other components, and the nature, order, or sequence of the corresponding components are not limited by these terms. When describing a component as being "connected," "coupled," or "coupled" to another component, it should be understood that the component can not only be directly connected or coupled to another component, but also that each component can be "connected," "coupled," or "coupled" to another component.
[0109] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0110] - Generating 3D moving images from 2D images -
[0111] In this embodiment, a method is proposed to realize the motion of a two-dimensional object as a motion image of a three-dimensional object by analyzing the motion of the object in the two-dimensional image.
[0112] Figure 1 FIG. 1 is a diagram illustrating an apparatus for generating a three-dimensional moving image based on artificial intelligence according to an embodiment of the present invention.
[0113] exist Figure 1 In the embodiment, a three-dimensional motion image generating apparatus 100 receives a two-dimensional image through a user terminal 10 including a camera, etc., and generates a three-dimensional motion image 20 simulating the motion of a two-dimensional object 11 in the two-dimensional image based on the two-dimensional image.
[0114] In the present invention, apparatus 100 may also be referred to as a system or platform. Three-dimensional motion image 20 represents the motion of a three-dimensional mesh object 21 that simulates the motion of a two-dimensional object 11. Three-dimensional mesh object 21 may be created by modeling the shape of two-dimensional object 11 or may be a pre-stored object that visualizes a specific character.
[0115] As an embodiment, the user terminal 10 may be configured to be able to communicate with the three-dimensional motion image generation device 100 via a network (not shown). Depending on the installation environment, the network may be configured as, for example, a wired network (such as Ethernet, a wired home network (power line communication), a telephone line communication device, and RS-serial communication), a wireless network (such as a mobile communication network, a wireless local area network (WLAN), Wi-Fi, Bluetooth, and ZigBee), or a combination thereof.
[0116] The user terminal 10 is a computing device capable of communicating with a network and may be equipped with hardware or applications for capturing, editing, or transmitting two-dimensional images. For example, the user terminal 10 may include a camera, a smartphone, a mobile phone, a navigation device, a computer, a notebook computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a tablet computer, a game console, a wearable device, an Internet of Things (IoT) device, a virtual reality (VR) device, an augmented reality (AR) device, and / or a set-top box.
[0117] Figure 2 It shows Figure 1 A block diagram showing an exemplary configuration of a three-dimensional moving image generating apparatus. Figure 2 The 3D motion image generating apparatus 100 may include a first artificial intelligence module 110 , a second artificial intelligence module 120 and a 3D motion image generating module 130 .
[0118] The first artificial intelligence module 110 may be a three-dimensional motion data generation module that extracts and analyzes the motion of an object in a two-dimensional image and generates three-dimensional motion data that simulates the motion. The first artificial intelligence module 110 includes one or more artificial neural networks that perform machine learning based on multiple learning data, and may include an artificial neural network based on a convolutional neural network (CNN).
[0119] The second artificial intelligence module 120 may be a 3D mesh object generation module that extracts and analyzes the shape of an object in a 2D image and generates a 3D mesh object by estimating the object's texture and volume. The second artificial intelligence module 120 includes one or more artificial neural networks that perform machine learning based on multiple learning data, and may include an artificial neural network based on a convolutional neural network.
[0120] The 3D motion image generation module 130 is a module that generates a 3D motion image based on output results of the first artificial intelligence module 110 and the second artificial intelligence module 120 .
[0121] For example, the 3D motion image generation module 130 may apply the 3D motion data generated by the first artificial intelligence module 110 to the 3D mesh object generated by the second artificial intelligence module 120, thereby generating a 3D motion image that depicts the motion of the 3D mesh object. Alternatively, the 3D motion image generation module 130 may apply the 3D motion data generated by the first artificial intelligence module 110 to a pre-stored 3D character object, thereby generating a 3D motion image that depicts the motion of the 3D character object.
[0122] As an embodiment, the 3D motion image generation module 130 may also be an artificial intelligence module based on a convolutional neural network.
[0123] Hereinafter, the specific operation of the 3D motion image generating apparatus 100 and the method of generating a 3D motion image using the apparatus will be described with reference to various embodiments.
[0124] Figure 3 FIG. 4 is a flowchart illustrating a method for generating a three-dimensional motion image based on artificial intelligence according to an embodiment of the present invention. Figure 3 The method for generating three-dimensional motion images can be obtained by Figure 1 Therefore, when the execution entity is omitted in the following steps, it is assumed that the execution entity is the 3D motion image generating device 100.
[0125] In step S110 , a two-dimensional image is acquired from a user terminal, etc. As an embodiment, the two-dimensional image is an image composed of a plurality of frames, and may be a moving image (or video) expressing the action or motion of an object or subject.
[0126] In step S120 , an artificial intelligence module that has undergone machine learning is used to identify an object from a two-dimensional image and analyze the movement of the identified object to generate three-dimensional modeling data of the object.
[0127] As an example, the three-dimensional modeling data may include three-dimensional motion data representing the motion of the recognized object and three-dimensional mesh data representing the shape and volume of the recognized object.
[0128] The generation of three-dimensional modeling data can be performed using one or more artificial intelligence modules that have been machine-learned, and as an embodiment, the generation of three-dimensional motion data and the generation of three-dimensional mesh data can be performed using different artificial intelligence modules respectively.
[0129] For example, a first artificial intelligence module can be preconfigured to perform machine learning using multiple learning data to extract the motion of a moving object in a two-dimensional image as three-dimensional motion. As one embodiment, the first artificial intelligence module can be a module that estimates the joints of the object in the two-dimensional image and infers the motion of the object in the two-dimensional image as three-dimensional motion based on the estimated joint positions.
[0130] When a two-dimensional image is input to the first artificial intelligence module, the first artificial intelligence module detects an object from the input two-dimensional image, then identifies the portion corresponding to the object through object segmentation, and extracts the object's pose data for each frame of the two-dimensional image. Pose data is data that represents the object's pose using key points and lines, and can be skeleton data. In addition, the first artificial intelligence module can merge the pose data extracted from each frame in chronological order to generate three-dimensional motion data that represents continuous motion.
[0131] Furthermore, a second artificial intelligence module can be pre-configured to perform machine learning using multiple learning data to extract the shape of objects in a two-dimensional image as a three-dimensional mesh. As one embodiment, the second artificial intelligence module can estimate the depth and volume of an object based on its volume, width, etc. in the two-dimensional image.
[0132] When a 2D image is input to the second AI module, it detects objects from the input 2D image, identifies the portion corresponding to the object through object segmentation, and extracts the object's pose data for each frame of the 2D image. Furthermore, the second AI module can collect and merge the pose data extracted from each frame from various angles, and based on the merged data, generate 3D mesh data representing the object's volume and texture.
[0133] As one embodiment, 3D mesh data can be generated based on 3D volume metrics. 3D volume metrics uses the initial position of an object in a 2D image as a reference point. Then, by determining the positional information of the object as it moves to other locations in the 2D image, the object's size is estimated, thereby generating a volume in the form of a 3D mesh.
[0134] In step S130 , a three-dimensional motion image is generated using the previously generated three-dimensional modeling data.
[0135] A three-dimensional motion image is an image that expresses the motion of a three-dimensional object by incorporating previously generated three-dimensional motion data into the three-dimensional object so that the three-dimensional object moves according to the three-dimensional motion data.
[0136] At this time, the three-dimensional object may be a three-dimensional mesh object created based on previously generated three-dimensional mesh data, or may be a pre-stored three-dimensional character object.
[0137] For example, since a three-dimensional mesh object is created based on the shape of an object in a two-dimensional image, if three-dimensional motion data is incorporated into the three-dimensional mesh object, a motion image can be generated in which the motion of the object in the two-dimensional image is reproduced by a three-dimensional object that simulates the shape of the object in the two-dimensional image.
[0138] Alternatively, when a 3D character object simulating a celebrity or virtual character is used instead of a 3D mesh object, if 3D motion data is incorporated into the 3D character object, a motion image can be generated in which the 3D character object reproduces the motion of an object in a 2D image.
[0139] Figure 4 It is further Figure 3 Flowchart of the embodiment of step S120. Figure 4 In the embodiment, each step is described by taking the sequential execution as an example, but the scope of the present invention is not limited thereto. For example, step S122 and step S124 can be executed simultaneously and in parallel with each other.
[0140] In step S121 , a first artificial intelligence module is used to extract posture data of an object from a two-dimensional image.
[0141] In step S122, the first artificial intelligence module generates three-dimensional motion data representing the motion of the object in the two-dimensional image based on the extracted posture data. The three-dimensional motion data may include key point data corresponding to the joints of the object.
[0142] Reference Figure 5 This is further explained.
[0143] Figure 5 1 is a block diagram exemplarily showing a detailed configuration of the first artificial intelligence module 110. Figure 5 , the first artificial intelligence module 110 may include a first sub-module 111 and a second sub-module 112 .
[0144] The first artificial intelligence module 110 receives a two-dimensional image as input and outputs three-dimensional motion data representing the motion of an object in the two-dimensional image. In this case, the input two-dimensional image may include multiple frames representing the motion of the object in a time series.
[0145] The first submodule 111 is a convolutional neural network module that extracts the posture data of the object from each frame of the two-dimensional image.
[0146] The second submodule 112 is a recurrent neural network (or RNN) based module that provides information about previous frames to the first submodule 112 .
[0147] When the nth frame 31 among multiple frames of a two-dimensional image is input to the first artificial intelligence module 110, the first submodule 111 detects an object from the frame 31, then identifies the part corresponding to the object through object segmentation, and extracts posture data 32 from the identified object.
[0148] As described above, pose data 32 is data representing the pose of an object using key points 32a and lines 32b, and may be skeletal data with three-dimensional coordinates (x, y, z). Key points 32a are components representing the joints of the object identified by analyzing the object in frame 31. Lines 32b are components representing the skeleton of the object identified by analyzing the object in frame 31. One or both ends of line 32b may be connected to key points 32a.
[0149] In addition, when the first submodule 111 extracts the gesture data 32 from the frame 31 , it may refer to the information of the previous frame provided by the second submodule 112 , for example, the information Fn-1 of the n-1th frame.
[0150] For example, since 3D motion data can be considered the result of connecting consecutive gesture data, if the gesture data 32 of the nth frame does not match the gesture data of the preceding and following frames, the motion of the generated 3D motion data will be unnatural. Therefore, to achieve smoother motion of the 3D motion data, when extracting the gesture data 32 of the nth frame 31, the first submodule 111 can receive and refer to the information Fn-1 of the previous frame from the second submodule 112.
[0151] Back again Figure 4 In step S123, the second artificial intelligence module is used to extract the shape information of the object from the two-dimensional image.
[0152] In step S123 , a second artificial intelligence module is used to generate a three-dimensional mesh object representing the volume of the object based on the extracted shape information.
[0153] Reference Figure 6 This is further explained.
[0154] Figure 6 1 is a block diagram exemplarily showing a detailed configuration of the second artificial intelligence module 120. Figure 6 The second artificial intelligence module 120 may include a three-dimensional reconstruction module 121 and a mesh modeling module 122 .
[0155] The second artificial intelligence module 120 receives a two-dimensional image as input and outputs a three-dimensional mesh object representing the volume and texture of the object in the two-dimensional image. In this case, the input two-dimensional image may include multiple frames representing the shape of the object at various angles and / or positions.
[0156] The 3D reconstruction module 121 analyzes key points and lines representing an object from a 2D image, and based on this, generates a 3D point cloud as shape information of the object.
[0157] The mesh modeling module 122 performs 3D mesh modeling and texture mapping based on the generated 3D point cloud, thereby generating a 3D mesh object.
[0158] When the 2D image 41 is input to the second artificial intelligence module 120, the 3D reconstruction module 121 tracks the key points of the object in the 2D image 41 and reconstructs 3D shape information based on the key points. Simultaneously, the 3D reconstruction module 121 detects and fits the lines of the object in the 2D image 41 and reconstructs line-based 3D shape information. The 3D reconstruction module 121 can use predetermined reconstruction parameters to reconstruct the 3D shape information.
[0159] In addition, the 3D reconstruction module 121 generates a 3D point cloud PC that estimates the 3D shape of the object in the 2D image using the 3D shape information reconstructed based on key points and / or lines. The generated 3D point cloud PC is provided to the mesh modeling module 122.
[0160] The mesh modeling module 122 performs 3D mesh modeling based on the three-dimensional point cloud PC. The result of the 3D mesh modeling is output as mesh information MI representing the volume of the three-dimensional object. The 3D mesh modeling performed by the mesh modeling module 122 can use predetermined modeling parameters. Simultaneously, the mesh modeling module 122 performs texture mapping based on the color image of the two-dimensional image 41. The result of the texture mapping is output as texture information TI representing the texture of the three-dimensional object.
[0161] In addition, the mesh modeling module 122 constructs three-dimensional mesh data using the mesh information MI and the texture information TI, and then generates a three-dimensional mesh object 42 based on the data.
[0162] Figure 7 It is further Figure 3FIG. 1 is a flow chart of another embodiment of the present invention that concretizes step S130. Figure 7 An embodiment is described in which the three-dimensional motion data generated in step S122 is applied to a pre-stored three-dimensional character object.
[0163] In step S131 , the three-dimensional motion data is merged into a pre-stored three-dimensional character object, and a motion image representing the motion of the three-dimensional character object is generated.
[0164] In addition, since the 3D motion data is created based on the input 2D image, and the 3D character object is a pre-stored object that is not related to the 2D image, there may be a discrepancy between the 3D motion data and the 3D character object. To resolve this discrepancy, it is necessary to adjust the 3D motion data to match the re-targeting operation of the 3D character object. This will refer to Figure 8 Further explanation.
[0165] Figure 8 It is further Figure 7 The step S131 is a flowchart of the embodiment. Figure 8 Retargeting embodiments are described for adjusting joint positions of three-dimensional motion data to match joint positions of a three-dimensional character object.
[0166] In step S131a, the joint positions of the 3D motion data are matched with the joint positions of the 3D character object based on the key point data of the 3D motion data. Here, matching the joint positions of the 3D motion data with the joint positions of the 3D character object may mean, for example, adjusting the joint positions of the 3D motion data so that the joint positions of the 3D motion data and the joint positions of the 3D character object are consistent.
[0167] For example, when applying 3D motion data to a 3D character object, the physical characteristics of the 2D object that forms the basis of the 3D motion data may differ from those of the 3D character object. Therefore, to eliminate the inconsistency caused by the differences in physical characteristics between the objects, the joint positions (i.e., key points) of the 3D motion data must be adjusted to align with those of the 3D character object.
[0168] At this time, the lines of the 3D motion data are also adjusted based on the adjusted key points of the 3D motion data. For example, assuming that the original 3D motion data contains key points K1 (1, 1, 1) and K2 (2, 2, 2), and a line L1 connecting key points K1 and K2, when the key points K1 and K2 are adjusted to K1 (1, 2, 2) and K2 (2, 3, 3) respectively through redirection, line L1 is also adjusted to a line connecting (1, 2, 2) and (2, 3, 3).
[0169] In step S131b, based on the joint positions of the matched three-dimensional motion data, the motion of the three-dimensional character object is realized as the motion simulating the three-dimensional motion data.
[0170] To this end, the motion of the 3D motion data can be adjusted based on the adjusted joint positions (i.e., key points) and lines of the 3D motion data, for example, the position, motion distance, and motion range of each key point and line can be adjusted. In addition, the motion of the 3D character object can be realized by incorporating the adjusted motion of the 3D motion data into the 3D character object.
[0171] The realized motion of the three-dimensional character object can be output and stored as a three-dimensional motion image.
[0172] Figure 9 It is shown that further Figure 3 The step S130 is a flowchart of the embodiment. Figure 9 An embodiment in which the three-dimensional motion data generated in step S122 is applied to the three-dimensional mesh object generated in step S124 is described.
[0173] In step S132, a motion image representing the motion of the three-dimensional mesh object is generated by merging the three-dimensional motion data into the previously generated three-dimensional mesh object.
[0174] A 3D mesh object is a 3D object extracted from an original 2D image. By incorporating 3D motion data into the 3D mesh object, a 3D motion image can be generated using the 3D object that reflects the appearance of the object presented in the 2D image as it is.
[0175] The 3D mesh object and the 3D motion data are both generated from the same 2D object, so the joint positions of the 3D mesh object and the joint positions of the 3D motion data will be consistent with each other. Therefore, in this embodiment, the reorientation operation can be omitted.
[0176] According to the embodiments of the present invention described above, the motion of an object can be extracted from a two-dimensional image based on artificial intelligence and realized as a three-dimensional motion image.
[0177] Furthermore, three-dimensional mesh data of an object may be generated using a two-dimensional image of the object based on artificial intelligence.
[0178] -Object Image Generation Based on Whole Body Reasoning-
[0179] In this embodiment, a method is proposed as follows: even if a body part of an object is hidden or cut out in an image, the hidden or cut out body part can be automatically restored by inferring the whole body of the object based on artificial intelligence, thereby generating an image of the object.
[0180] Generally speaking, techniques for restoring hidden or cropped areas of an image are called "in-painting" if the area to be restored is inside the image, and "out-painting" if it is outside the image. In-painting refers to filling in the area to be restored by referencing the surrounding pixels. However, in this case, it may be difficult to distinguish between the object and the background in the image.
[0181] To effectively address this problem, semantic information about the object is used as conditional information. Traditional image restoration techniques typically use semantic information such as 2D segmentation masks or 2D key points as input. However, 2D segmentation masks struggle to distinguish the various parts that make up an object, making it difficult to recover image details. 2D key points, on the other hand, struggle to predict the object's volume, blurring the distinction between the object and the background, leading to reduced quality in the restored image.
[0182] Therefore, in this embodiment, a method for restoring an image based on 3D information containing more information is proposed, thereby achieving a more detailed and realistic image restoration method.
[0183] Furthermore, currently widely used AI-based image analysis technologies are typically optimized for two-dimensional data. Therefore, in this embodiment, rather than using three-dimensional information such as a 3D mesh model as is, image restoration is performed using two-dimensional data containing 3D information—UV maps. This presents a method that can restore high-quality, realistic images while leveraging various currently used AI module design technologies.
[0184] Figure 10 FIG. 1 is a diagram showing an exemplary configuration of an object image generation device based on artificial intelligence reasoning of a local object according to an embodiment of the present invention. Figure 10 The object image generating apparatus 200 may include a first artificial intelligence module 210 , a second artificial intelligence module 220 , a three-dimensional motion image generating module 230 , a third artificial intelligence module 240 , a fourth artificial intelligence module 250 and a learning data generating module 260 .
[0185] Some configurations of the object image generating device 200 may be Figure 2 For example, the first artificial intelligence module 210, the second artificial intelligence module 220 and the 3D motion image generation module 230 may be respectively Figure 2The first artificial intelligence module 110, the second artificial intelligence module 120, and the 3D motion image generation module 130 are configured identically. Therefore, to avoid duplication, the detailed description of the first artificial intelligence module 210, the second artificial intelligence module 220, and the 3D motion image generation module 230 will be omitted in this embodiment.
[0186] The third artificial intelligence module 240 is a UV map generation module. It uses a CNN-based artificial neural network to generate a UV map corresponding to an image of an object in which a portion of the object's body is clipped, hidden, or obscured. The specific definition and composition of a UV map are well known in the field of image processing technology to which this technology belongs, and will not be further described here.
[0187] The fourth artificial intelligence module 250 generates an object image with the body part restored according to the object image and the UV map using a CNN-based artificial neural network. As an embodiment, the object image with the body part restored may be an image showing the entire body of the object.
[0188] The learning data generation module 260 generates learning data for machine learning of the third artificial intelligence module 240 and / or the fourth artificial intelligence module 250 .
[0189] The specific functions and / or learning methods of the third artificial intelligence module 240, the fourth artificial intelligence module 250 and the learning data generation module 260 will be referred to in detail. Figure 11 and described in more detail below.
[0190] Figure 11 is a flowchart illustrating a method for generating an object image based on artificial intelligence reasoning of a local object according to an embodiment of the present invention. Figure 11 The object image generation method can be Figure 10 Therefore, when the execution entity is omitted in the following steps, it is assumed that the execution entity is the object image generation device 200.
[0191] In step S210, a first object image and a UV map corresponding to the first object image are acquired. In this case, the first object image may be an image in which a body part of the object is cut, hidden or blocked.
[0192] The UV map can be obtained by the third artificial intelligence module 240. Figure 12 This is further explained.
[0193] exist Figure 12In the embodiment, the third artificial intelligence module 240 is an artificial intelligence module that is trained by machine learning to output a UV map corresponding to the entire body of the object when an image in which a body part of the object is missing (for example, a body part of the object is cut, hidden, or blocked) is input, and can be an artificial intelligence module based on a convolutional artificial neural network.
[0194] For example, Figure 12 As shown, when the first object image 41 with one arm of the object cut off is input to the third artificial intelligence module 240, the third artificial intelligence module 240 uses the internal machine-learned artificial neural network to output the UV corresponding to the whole body of the object. Figure 42 In addition, Figure 15 The machine learning method of the third artificial intelligence module 240 is described below.
[0195] Back again Figure 11 In step S220, a second object image in which the body part of the object is restored is generated based on the first object image and the UV map. The second object image may be generated by the fourth artificial intelligence module 250. Figure 13 This is further explained.
[0196] Figure 13 It shows that Figure 11 The step S220 is a flowchart of the embodiment.
[0197] In step S221, a fourth artificial intelligence module is used to generate a foreground image and a background image corresponding to the first object image, wherein the foreground image may refer to an image showing the object, and the background image may refer to an image showing the background of the object.
[0198] In step S222, a second object image in which the body part of the object is restored is generated by synthesizing the previously generated foreground image and background image. Figure 14 This is further explained.
[0199] Reference Figure 14 , shows a specific configuration of the fourth artificial intelligence module 250. The fourth artificial intelligence module 250 may include an encoder 251, a foreground generator 252, a background generator 253 and / or a synthesizer 254. However, this only shows an exemplary configuration of the fourth artificial intelligence module 250, and the scope of the present invention is not limited thereto. For example, the fourth artificial intelligence module 250 may also include Figure 14 Other configurations than the encoder 251, foreground generator 252, background generator 253 and / or synthesizer 254 shown.
[0200] The encoder 251 receives the first object image 41 and UV Figure 42 As input, and extract the feature vector v based on it.
[0201] The foreground generator 252 generates a foreground image 43 corresponding to the first object image 41 using the first part v1 of the feature vector v.
[0202] The background generator 253 generates a background image 44 corresponding to the first object image 41 using the second portion v2 of the feature vector v. As an embodiment, the first portion v1 may be half of the feature vector v, and the second portion v2 may be the other half of the feature vector v.
[0203] The synthesizer 254 synthesizes (or mixes) the generated foreground image 43 and background image 44 to generate a second object image 45 in which the body part of the object is restored.
[0204] As one embodiment, the encoder, foreground generator and background generator can be structures including convolutional neural networks, and can be structures that are respectively subjected to machine learning to extract feature vectors based on object images and UV maps, generate a foreground image using a first part of the feature vector, and generate a background image using a second part of the feature vector.
[0205] According to the configuration described above, when the first object image 41 and UV Figure 42 When input to the fourth artificial intelligence module 250, the fourth artificial intelligence module 250 uses the internal artificial neural network that has been machine-learned to generate a foreground image 43 and a background image 44 corresponding to the first object image 41, and synthesizes (blending) the generated foreground image 43 and background image 44 to generate a second object image 45 in which the cut body part of the object is restored.
[0206] As an embodiment, the fourth artificial intelligence module 250 may be a module for performing machine learning based on a generative adversarial network model. Figure 17 and described below.
[0207] Figure 15 is a flowchart exemplarily illustrating the learning method of the third artificial intelligence module. Figure 15 and Figure 16 Provide explanation.
[0208] First, in step S310 , a first UV map corresponding to the two-dimensional object image is generated by converting the three-dimensional mesh object corresponding to the two-dimensional object image.
[0209] Reference Figure 16 , shows a two-dimensional object image 51, which is an original image including the entire body of a subject. The two-dimensional object image 51 is input to the second artificial intelligence module 220, which outputs a three-dimensional mesh object 52 corresponding to the two-dimensional object image 51. The second artificial intelligence module 220 is a machine-learned artificial intelligence module that extracts and analyzes the shape of the object in the two-dimensional image to generate a three-dimensional mesh object that estimates the texture and volume of the object. The specific functions and operations of the second artificial intelligence module 220 have been described in detail above and will not be repeated here.
[0210] Furthermore, by converting the three-dimensional mesh object 52, a first UV map 53 corresponding to the two-dimensional object image 51, which serves as the original image, is generated. Generally, a three-dimensional mesh object can be easily converted to generate a corresponding UV map through a predetermined mathematical transformation, and this method is well known in the art. Therefore, a detailed description of the method for converting the three-dimensional mesh object to a UV map is omitted here. As one embodiment, the first UV map 53 can be used as a pseudo ground truth for evaluating the reasoning results of the third artificial intelligence module 240.
[0211] In step S320 , a plurality of learning data are acquired based on the two-dimensional object image. The plurality of learning data may be generated by the learning data generating module 260 .
[0212] Specifically, a plurality of learning data 54 may be acquired by randomly cutting, hiding, or blocking a portion of the two-dimensional object image 51 and filling the cut, hidden, or blocked portion with random noise.
[0213] In step S330 , a plurality of learning data are provided to a third artificial intelligence module, and the third artificial intelligence module outputs a second UV map corresponding to the learning data.
[0214] The third artificial intelligence module 240 is an artificial intelligence module based on a convolutional neural network designed to receive an image in which a body part of the subject is cut, hidden, or occluded as input, infer, and output a UV map corresponding to the entire body of the subject.
[0215] When the plurality of learning data 54 is input to the third artificial intelligence module 240 , the third artificial intelligence module 240 outputs a second UV map 55 corresponding to the plurality of learning data 54 through the operation of the internal artificial neural network.
[0216] In step S340 , based on the comparison result of the first UV map and the second UV map, the parameters of the third artificial intelligence module are updated.
[0217] As an embodiment, the MSE loss may be calculated by calculating the difference between the first UV map 53 and the second UV map 55 using a predetermined loss function, and the parameters of the third artificial intelligence module 240 may be updated based on the MSE loss.
[0218] In addition, steps S310 to S340 may be repeated until the third artificial intelligence module 240 is fully learned or reaches a predetermined number of times.
[0219] Figure 17 is a flowchart exemplarily illustrating a learning method of the fourth artificial intelligence module. Figure 17 and Figure 18 Explain this.
[0220] First, in step S410 , a two-dimensional object image is generated using an original image.
[0221] The two-dimensional object image 62 may be an image in which a portion of the original image 61 or a body part is randomly cropped, hidden, or blocked, or an image in which the cropped, hidden, or blocked portion is replaced with random noise. As one embodiment, the two-dimensional object image may be generated using the learning data generation module 260.
[0222] In step S420 , a UV map corresponding to the two-dimensional object image is generated.
[0223] As an example, the UV map 63 may be an output of the third artificial intelligence module 240 , which is generated by inputting the two-dimensional object image 62 into the third artificial intelligence module 240 .
[0224] In step S430 , a first feature vector is extracted based on the two-dimensional object image and the UV map.
[0225] For example, when the two-dimensional object image 62 and the UV map 63 are input to the encoder 251 , the encoder 251 outputs a first feature vector v corresponding thereto based on the two-dimensional object image 62 and the UV map 63 .
[0226] As an embodiment, the two-dimensional object image 62 and the UV map 63 may be input to the encoder 251 as concatenated data.
[0227] In step S440, a first foreground image corresponding to the two-dimensional object image is generated using the first part of the first eigenvector.
[0228] For example, after extracting the first feature vector v from the encoder 251, the foreground generator 252 generates a first foreground image 64 in which a body part of the object is restored using a first part v1 of the first feature vector v.
[0229] In step S450 , a first background image corresponding to the two-dimensional object image is generated using the second part of the first eigenvector.
[0230] For example, the background generator 253 generates a first background image 65 showing the background of the object using the second portion v2 of the first feature vector v. At this time, the second portion v2 may be the remaining portion of the first feature vector v except the first portion v1.
[0231] In step S460 , a composite image is generated based on the first foreground image and the first background image.
[0232] For example, the synthesizer 254 may generate a synthesized image 66 by synthesizing (or mixing) the first foreground image 64 and the first background image 65 .
[0233] In step S470, based on the comparison results of the first foreground image, the first background image and / or the synthesized image with the original image, the parameters of the fourth artificial intelligence module are updated.
[0234] For example, the segmentation loss Lseg may be calculated by calculating the difference between the feature vector vf of the first foreground image 64 and the feature vector vo of the original image 61. The segmentation loss Lseg may be used to update the parameters of the foreground generator 252.
[0235] Alternatively, the patch loss Lpatch can be calculated by calculating the difference between the first background image 65 and the original image 61. Specifically, the distance between the patches (Pb, Po) at corresponding positions in the first background image 65 and the original image 61 can be calculated and used as the patch loss Lpatch. The patch loss Lpatch can be used to update the parameters of the background generator 253.
[0236] Alternatively, a reconstruction loss Lrec may be calculated based on a comparison result between the synthesized image 65 and the original image 61. The reconstruction loss Lrec is a loss value defined to make the synthesized image 65 that is inferred and outputted more similar to the original image 61.
[0237] Alternatively, an adversarial loss Ladv can be calculated based on the result of discriminator 270 determining whether the synthesized image 65 is realistic. Adversarial loss Ladv is a loss value defined to prevent blurring of the synthesized image 65 output by fourth artificial intelligence module 250 and to make the image more realistic. The reconstruction loss Lrec and adversarial loss Ladv can be used to update the parameters of fourth artificial intelligence module 250.
[0238] In addition, steps S410 to S470 may be repeated until the third artificial intelligence module 240 is fully learned or reaches a predetermined number of times.
[0239] According to the embodiments of the present invention described above, even if a body part of an object in an image is hidden or cut, the hidden or cut body part can be automatically restored by inferring the whole body of the object based on artificial intelligence.
[0240] In addition, by using UV maps, which are two-dimensional data containing three-dimensional information, the body parts of the object that are hidden or cut can be restored more carefully and realistically while utilizing existing convolutional neural network-based model technology.
[0241] -Calculating motion similarity of 3D mesh objects based on metric learning-
[0242] In this embodiment, a technology is proposed for identifying the motion of objects in a source image in a three-dimensional grid manner and calculating the similarity between the motions of the objects using a metric learning technique.
[0243] In recent years, with the advancement of imaging equipment and the expansion of the image content market, research has been actively underway to estimate the posture and movement of subjects (particularly human subjects) based on images. In particular, research into comparing subjects' movements or calculating their similarity by analyzing only their appearance in images using artificial intelligence modules, without attaching markers or sensors to the body, is still in its infancy.
[0244] Traditional AI-based object motion estimation techniques are typically based on 2D skeleton models. These models extract the body's joints and the lines connecting them, and estimate the object's motion based on the movement of these joints and lines. While this approach offers the advantage of being able to concisely estimate and compare object motions, it is limited by its difficulty distinguishing facial expressions or gestures, and its difficulty identifying detailed motion details.
[0245] Therefore, in this embodiment, a method and device for calculating the similarity of an object's motion based on artificial intelligence is provided, which can infer the object's motion based on a three-dimensional mesh model and can also consider the object's volume and body characteristics, thereby identifying more detailed motions.
[0246] In addition, when constructing the artificial intelligence module, an artificial intelligence module that performs machine learning based on metric learning is utilized, thereby providing a method that can perform similarity analysis on postures and actions that have not been learned.
[0247] Figure 19 FIG. 1 is a diagram showing an exemplary configuration of an object action similarity calculation device based on artificial intelligence according to an embodiment of the present invention. Figure 19The object action similarity calculation device 300 may include a first artificial intelligence module 310 , a second artificial intelligence module 320 , a three-dimensional motion image generation module 330 , a feature vector extraction module 340 , a fifth artificial intelligence module 350 and a retrieval result providing module 360 .
[0248] Some configurations of the object action similarity calculation device 300 can be Figure 2 For example, the first artificial intelligence module 310, the second artificial intelligence module 320 and the three-dimensional motion image generation module 330 may be respectively the same as the three-dimensional motion image generation device 100. Figure 2 The first artificial intelligence module 110, the second artificial intelligence module 120, and the 3D motion image generation module 130 have the same configuration. Therefore, to avoid duplication, the detailed description of the first artificial intelligence module 310, the second artificial intelligence module 320, and the 3D motion image generation module 330 will be omitted in this embodiment.
[0249] The feature vector extraction module 340 extracts feature vectors corresponding to the 3D mesh object. The feature vectors are values representing the posture and motion of the 3D mesh object and may include values representing the directionality and trend of each mesh data constituting the 3D mesh object.
[0250] The fifth artificial intelligence module 350 is a module that calculates the motion similarity between three-dimensional mesh objects based on feature vectors extracted from multiple three-dimensional mesh objects. The fifth artificial intelligence module 350 is a module that performs machine learning based on learning data and can be a module that performs machine learning based on metric learning.
[0251] The retrieval result providing module 360 is a module that selects a target image having a posture and action similar to the source image as a retrieval result based on the action similarity calculated by the fifth artificial intelligence module 350, or outputs the action similarity between the source image and the target image as a specific value or level.
[0252] For the specific functions and / or learning methods of the feature vector module 340, the fifth artificial intelligence module 350 and the search result providing module 360, please refer to Figure 20 and described in more detail below.
[0253] Figure 20 is a flowchart illustrating a method for calculating motion similarity of objects based on artificial intelligence according to an embodiment of the present invention. Figure 20 The method for calculating the action similarity of the object can be obtained by Figure 19 Therefore, when the execution entity is omitted in the following steps, it is assumed that the execution entity is the object action similarity calculation device 300.
[0254] In step S510 , a source image and a target image are input.
[0255] Here, the source image is the image used as the reference for action comparison, and the target image is the target image used for action comparison. For example, if you want to find a second image that has a similar action to the first image after acquiring the first image and calculate their action similarity, the first image becomes the source image and the second image becomes the target image.
[0256] As an embodiment, the target image may be an image pre-stored in a database within the object motion similarity calculation device 300 along with other target images, or may be an image sequentially read from the database along with other target images for motion comparison with the source image.
[0257] In step S520 , a first feature vector is extracted from a first object corresponding to the source image, and a second feature vector is extracted from a second object corresponding to the target image.
[0258] As an embodiment, the first object and the second object may be three-dimensional mesh objects.
[0259] The first object may be an object in the source image, or an object newly processed or generated based on an object in the source image. Similarly, the second object may be an object in the target image, or an object newly processed or generated based on an object in the target image.
[0260] The first feature vector is a vector representing the posture and motion of the first object. The first feature vector may include values representing directionality and tendency of each of the plurality of mesh data constituting the first object.
[0261] Likewise, the second feature vector is a vector representing the posture and motion of the second object.The second feature vector may include values representing the directionality and trend of each of the plurality of mesh data constituting the second object.
[0262] For the extraction of the first and second eigenvectors, Figure 23 and described in more detail below.
[0263] In step S530 , the action similarity between the first object and the second object is calculated based on the first feature vector and the second feature vector.
[0264] As an embodiment, an artificial intelligence module based on metric learning for machine learning can be used to calculate action similarity. Figure 25 The machine learning method based on metric learning and the method of calculating the action similarity between objects using this method are described in more detail in .
[0265] Figure 21 and Figure 22 This is a diagram for explaining the characteristics of using a three-dimensional mesh model when comparing the similarity of the movements of objects.
[0266] Reference Figure 21 , shows an example of estimating a two-dimensional (or 2D) skeleton model 72 based on an original image 71 and estimating a three-dimensional (or 3D) mesh.
[0267] The two-dimensional skeleton model 72 is constructed by estimating the joint points and the lines connecting the joint points from the original image 71. The three-dimensional mesh model 73 is also constructed by estimating the joint points and the lines connecting the joint points from the original image 71, but in addition, it is constructed into a three-dimensional mesh form having volume after estimating the texture and depth of the object.
[0268] When estimating the subject's motion, using a two-dimensional skeleton model 72 estimates the subject's motion based on the positions and motions of joints and lines. On the other hand, using a three-dimensional mesh model 73 also considers mesh data in addition to joints and lines. Therefore, using the three-dimensional mesh model 73 allows for more detailed motion recognition by taking into account the subject's volume and body features.
[0269] Reference Figure 22 This is further explained.
[0270] Figure 22 The difference in recognition performance between using a 2D skeleton model and a 3D mesh model for arm rotation motion is shown.
[0271] Figure 22 (a) shows the arm rotation motion in the original image. Figure 22 (b) shows the case of recognizing arm rotation motion based on a two-dimensional skeleton model, and Figure 22 (c) shows the case of recognizing arm rotation motion based on a three-dimensional mesh model.
[0272] Reference Figure 22 (b) Since the two-dimensional skeleton model recognizes the movement of the object based on joint points and lines, when there is almost no change in the positions of the joint points and lines between 74a before the movement and 74b after the movement, such as an arm rotation movement, there is a problem in not being able to recognize the corresponding movement well.
[0273] On the contrary, refer to Figure 22 (c) Since the three-dimensional mesh model considers mesh data in addition to joint points and lines to identify the movement of the object, even if the positions of the joint points and lines hardly change, the corresponding action can be well identified by the changes in the mesh data before 75a and after 75b of the action.
[0274] Figure 23 and Figure 24 For further explanation Figure 20 FIG. 10 is a diagram of step S520.
[0275] Reference Figure 23 , showing a feature vector extraction module 340 that extracts feature vectors from source images and feature vectors respectively.
[0276] When a source image is input to the feature vector extraction module 340 , a first feature vector is extracted from a first object corresponding to the source image.
[0277] In this case, the first object can be an object in the source image, or an object newly generated based on an object in the source image. For example, if the source image is a 3D image and contains a 3D mesh object, the first object can be a 3D mesh object extracted from the source image. On the other hand, if the source image is a 2D image and the object in the source image is a 2D object, the first object can be a 3D mesh object generated by extracting the 2D object from the source image and then estimating the texture and volume of the 2D object using an artificial intelligence module. In this case, the second artificial intelligence module 320 can be used to estimate and generate the 3D mesh object.
[0278] Similarly, when the target image is input to the feature vector extraction module 340 , a second feature vector is extracted from the second object corresponding to the target image.
[0279] The second object can be an object in the target image, or a newly generated object based on an object in the target image. For example, if the target image contains a 3D mesh object, the second object can be a 3D mesh object extracted from the target image. Alternatively, if the object in the target image is a 2D object, the second object can be a 3D mesh object generated by extracting the 2D object from the target image and then estimating the texture and volume of the 2D object using an artificial intelligence module. In this case, the second artificial intelligence module 320 can also be used to estimate and generate the 3D mesh object.
[0280] As one embodiment, feature vector extraction module 340 may use a Skinned Multi-Person Linear (SMPL) model to extract first and second feature vectors from the first and second objects, respectively. The SMPL model is a solution for estimating object poses based on a three-dimensional mesh model. Since the specific technical content related to the SMPL model is well known in the art, its description is omitted here.
[0281] In addition, the first feature vector may include a first mesh vector representing each of the first plurality of meshes constituting the first object, and a first trend vector representing a change trend of the first mesh vector.
[0282] Similarly, the second feature vector may include a second mesh vector representing each of the second plurality of meshes constituting the second object, and a second trend vector representing a change trend of the second mesh vector.
[0283] Reference Figure 24 This is further explained.
[0284] Figure 24 (a) shows a human arm part as a three-dimensional mesh object. Figure 24 (b) shows mesh vectors representing the directionality of each mesh data constituting the three-dimensional mesh object.
[0285] For the sake of convenience, only the four mesh data 81a, 81b, 81c, and 81d in the mesh data constituting the three-dimensional mesh object are taken as an example. The feature vector extraction module 340 extracts corresponding mesh vectors 82a, 82b, 82c, and 82d for each mesh data 81a, 81b, 81c, and 81d.
[0286] As one embodiment, each of mesh vectors 82a, 82b, 82c, and 82d may be a normal vector of the corresponding mesh data 81a, 81b, 81c, and 81d. For example, mesh vector 82a may be obtained by extracting the normal vector of mesh data 81a, mesh vector 82b may be obtained by extracting the normal vector of mesh data 81b, mesh vector 82c may be obtained by extracting the normal vector of mesh data 81c, and mesh vector 82d may be obtained by extracting the normal vector of mesh data 81d.
[0287] In addition, although the mesh vectors 82a, 82b, 82c, and 82d are described herein as normal vectors of the mesh data 81a, 81b, 81c, and 81d, the scope of the present invention is not limited thereto. For example, any vector that can express the directionality of the mesh data 81a, 81b, 81c, and 81d and that represents the mesh data 81a, 81b, 81c, and 81d can be used as a mesh vector.
[0288] In addition, the feature vector extraction module 340 may also extract a trend vector representing a changing trend of the grid vectors 82a, 82b, 82c, and 82d. As one embodiment, the trend vector is a vector representing a changing trend of each grid vector 82a, 82b, 82c, and 82d in consecutive frames of the image, and may be a vector representing the trend in terms of magnitude and direction.
[0289] The grid vector and the trend vector extracted by the feature vector extraction module 340 may be extracted and provided as a feature vector of the object.
[0290] Figure 25 and Figure 26For further explanation Figure 20 FIG. 10 is a diagram of step S530.
[0291] Reference Figure 25 , showing a fifth artificial intelligence module 350 for calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector.
[0292] The fifth artificial intelligence module 350 is a module that undergoes machine learning to receive a first feature vector extracted from a first object and a second feature vector extracted from a second object as input, calculates a distance between the first feature vector and the second feature vector based thereon, and then outputs an action similarity between the first object and the second object based on the calculated distance.
[0293] As an embodiment, the closer the distance between the first feature vector and the second feature vector is, the higher the calculated action similarity may be.
[0294] As described above, the fifth artificial intelligence module 350 may be a module that performs machine learning in a metric learning manner.
[0295] The cosine similarity method is a representative existing method for calculating object motion similarity. This method determines motion similarity based on absolute coordinates. Therefore, when comparing the dance movements of an adult and a child, even if the overall posture and motion are similar, the coordinates identified from the adult and the child are significantly different, and therefore judged as dissimilar.
[0296] In contrast, when determining the motion similarity of objects using a model learned using a metric learning approach, the motion similarity is determined based on the directionality and trend of each object, so as long as the overall posture and motion are similar, the motion of an adult and the motion of a child are also judged to be similar.
[0297] Reference Figure 26 The learning method of the fifth artificial intelligence module 350 based on metric learning is described.
[0298] Metric learning-based machine learning focuses on whether the features of two objects are similar or different, and determines the degree of similarity. In other words, when learning an AI module, it doesn't perform "classification" but instead uses a metric to calculate a distance that indicates similarity or dissimilarity. The specific technical details of metric learning are well known in the art and will not be elaborated here.
[0299] exist Figure 26In the example, the fifth artificial intelligence module 350 receives as input anchor data, positive data, and negative data as learning data. The anchor data is a value corresponding to reference data a, which may be a feature vector extracted from the reference data a. The positive data is a value corresponding to data b similar to the reference data a, which may be a feature vector extracted from the similar data b. The negative data is a value corresponding to data c different from the reference data a, which may be a feature vector extracted from the different data c.
[0300] After receiving anchor data, positive data, and negative data, the fifth artificial intelligence module 350 learns to reduce the distance to similar positive data relative to the anchor data and to increase the distance to different negative data relative to the anchor data. In other words, the fifth artificial intelligence module 350 learns features to identify that the anchor data and positive data are similar, while the anchor data and negative data are different.
[0301] The fifth artificial intelligence module 350 that has been learned in this way can distinguish between similar postures and movements and dissimilar postures and movements, and can also perform similarity analysis on postures and movements that have not been learned.
[0302] Figure 27 is a flow chart illustrating a method for calculating the similarity of actions of objects based on artificial intelligence according to another embodiment of the present invention. Figure 27 In the embodiment of the present invention, steps S510 to S530 are Figure 20 Steps S510 to S530 are the same as those in the embodiment. The only difference is that steps S540 and S550 are also included. In this embodiment, the description of steps S510 to S530 is omitted to avoid repeated description.
[0303] In step S540 , based on the action similarity, the target image is provided as a similar image retrieval result of the source image. The retrieval result providing module 360 may perform the provision of the retrieval result.
[0304] As one embodiment, if the calculated action similarity exceeds a predetermined threshold, the target image is determined to be a similar image to the source image, and the target image is provided as a search result corresponding to the source image. On the other hand, if the calculated action similarity is less than the predetermined threshold, the target image is determined not to be a similar image to the source image, and the target image is not provided as a search result.
[0305] In addition, the retrieval results can be provided in the form of listing multiple target images whose action similarity with the source image exceeds a predetermined threshold, such as Figure 28 shown.
[0306] For example, suppose there are multiple target images stored in the database, and Figures 20 to 26The method described in ,computes the action similarity of each target image with the source image 90.
[0307] Among them, the target images 91 , 92 , 93 , 94 , 95 , and 96 whose action similarities exceed a predetermined threshold may be sorted and listed in order of their action similarities, and provided as retrieval results of similar images of the source image 90 .
[0308] In step S550 , based on the user's selection of the target image provided as the retrieval result, the three-dimensional motion data of the second object is merged into the first object to generate a motion image in which the first object moves according to the three-dimensional motion data.
[0309] For example, if a user films their own movements and inputs them as a source image, and searches for an image with movements similar to the source image from target images stored in a database, the method for calculating the movement similarity of an object according to the present invention can provide a target image with a movement similarity greater than or equal to a threshold value to the source image as a similar image retrieval result. At this time, if a target image is selected through a provided user interface, the movement data of the sample object (corresponding to the second object) in the target image can be combined with the user's appearance (corresponding to the first object) extracted from the source image, thereby generating a movement image whose appearance corresponds to the user but whose movements are consistent with those of the sample object.
[0310] As an embodiment, the second artificial intelligence module 320 can be used to extract the user's appearance, the first artificial intelligence module 310 can be used to extract the motion data of the sample object, and the three-dimensional motion image generation module 330 can be used to combine the user's appearance and motion data.
[0311] Figure 29 Flowchart of a method for calculating the similarity of actions of objects based on artificial intelligence according to another embodiment of the present invention. Figure 29 In the embodiment of the present invention, steps S510 to S530 are Figure 20 Steps S510 to S530 are the same as those in the embodiment. The only difference is that step S560 and step S550 are also included. In this embodiment, the description of steps S510 to S530 is omitted to avoid repeated description.
[0312] In step S560 , a similarity score between the source image and the target image is output based on the action similarity.
[0313] As an embodiment, the similarity score may be output as a score or a level according to a score interval.
[0314] For example, the similarity score can be output as a numerical score, such as 76, 86, 96, or as a grade based on the score range, such as greater than or equal to 70 points and less than 80 points as C grade, greater than or equal to 80 points and less than 90 points as B grade, and greater than or equal to 90 points as A grade.
[0315] The output of similarity scores can be used in applications such as games. For example, in a game designed to mimic actions in a specific video, the user can be filmed imitating the actions in the specific video, the actions in the filmed video can be compared with the actions in the specific video, the similarity between the two actions can be calculated, and the similarity score based on this can be output as the user's game result.
[0316] According to the embodiments of the present invention described above, by comparing the postures and actions of objects using a three-dimensional mesh model, the body characteristics of the objects can be taken into account, thereby more accurately calculating the action similarity.
[0317] In addition, by utilizing an artificial intelligence module that performs machine learning based on metric learning, similarity analysis can also be performed on postures and movements that have not been learned.
[0318] Generate 3D moving images from input sentences
[0319] In this embodiment, a technology is proposed for automatically generating a three-dimensional moving image corresponding to the meaning of an input sentence by analyzing an input sentence in a text format.
[0320] For example, a technology is proposed that retrieves a motion image that expresses the meaning of each phrase of an input sentence from a database, and automatically generates a three-dimensional motion image that expresses the overall meaning of the input sentence by combining the retrieved motion images. For example, various motion images are generated in advance and stored in a database. Then, when a sentence such as "a person is walking" is input, a motion image that expresses the way a person walks is retrieved from the database and output.
[0321] Figure 30 FIG. 1 is a diagram showing an exemplary configuration of a moving image generating apparatus according to an embodiment of the present invention. Figure 30 The motion image generation device 400 is a device for automatically generating a three-dimensional motion image from a sentence, which may include a sentence acquisition module 410, a morpheme analysis module 420, an image retrieval module 430 and a motion image generation module 440.
[0322] The sentence acquisition module 410 is a module for acquiring sentences for generating 3D moving images. The sentence acquisition module 410 can directly receive sentences in text form as input through a user interface, or convert a voice signal input through a voice input device into text to acquire sentences.
[0323] The morpheme analysis module 420 is a module that analyzes the acquired sentence, distinguishes words in the sentence, and recognizes the part of speech of each word.
[0324] The image retrieval module 430 is a module that searches a database for a partial image having a meaning corresponding to a word whose part of speech has been identified.
[0325] The moving image generation module 440 generates a moving image corresponding to the overall meaning of the acquired sentence by combining the retrieved partial images. In this case, the moving image generation module 440 may generate an intermediate image and insert it between the partial images so that the partial images connect naturally when combined.
[0326] Will refer to Figure 31 The specific operation method of the motion image generation module 440 is described in detail.
[0327] As an embodiment, the sentence acquisition module 410, the morpheme analysis module 420, the image retrieval module 430, and / or the motion image generation module 440 may include an artificial intelligence module that has undergone machine learning. For example, the sentence acquisition module 410 may include an artificial intelligence module that has undergone machine learning to recognize speech and automatically convert the recognized speech into text; the morpheme analysis module 420 may include an artificial intelligence module that has undergone machine learning to distinguish each word in a text-formatted sentence and identify the part of speech of the distinguished words; the image retrieval module 430 may include an artificial intelligence module that has undergone machine learning to selectively retrieve partial images corresponding to the meaning of specific words from a database; and the motion image generation module 440 may include an artificial intelligence module that has undergone machine learning to determine the arrangement order of each partial image according to the context when generating an entire motion image by combining multiple partial images, and to generate an intermediate image to be inserted between the partial images so that the beginning and end of each partial image are naturally connected.
[0328] Figure 31 is a flowchart illustrating a method for generating a moving image from a sentence according to an embodiment of the present invention. Figure 31 The method of generating moving images from sentences can be obtained by Figure 30 Therefore, when the execution entity is omitted in the following steps, it is assumed that the execution entity is the motion image generation device 400.
[0329] In step S310 , a sentence for generating a three-dimensional moving image is acquired.
[0330] As one embodiment, a sentence can be obtained based on text input through a user interface. Alternatively, a sentence can be obtained by converting a voice signal input through a voice input device into text using an artificial intelligence module. Since the specific technical content of the method for converting a voice signal into text using an artificial intelligence module is well known in the art, its description is omitted here.
[0331] In step S320, one or more words in the sentence are identified. Figure 32 This is further explained.
[0332] Figure 32 It shows that Figure 31 Step S620 is a flowchart of an embodiment.
[0333] In step S621, one or more words constituting a sentence are identified.
[0334] In step S622 , the parts of speech of the identified one or more words are identified.
[0335] As one embodiment, the recognition of one or more words and the recognition of parts of speech may be performed by an artificial intelligence module that performs machine learning for natural language processing.
[0336] Reference Figure 33 This is further explained.
[0337] Figure 33 An example of inputting a sentence 'A Person sits down with crossed legs, beforeegetting up' to generate a three-dimensional motion image is shown.
[0338] At this time, each word and its part of speech in the sentence is identified by the morpheme analysis module 420, and the words can be distinguished and classified according to each part of speech, such as Figure 33 shown.
[0339] Back again Figure 31 In step S330, the partial image corresponding to one or more words is retrieved from the database. As an embodiment, the database may be a database storing Figures 1 to 9 The structure of the plurality of three-dimensional moving images generated by the method for generating a three-dimensional moving image described in the invention. The plurality of three-dimensional moving images stored in the database can be used as partial images for generating a moving image corresponding to the meaning of a sentence.
[0340] In addition, about Figure 33 A more detailed description of the illustrated embodiment will be given with reference to Figure 34 .
[0341] Figure 34 It shows that Figure 31 Step S630 is a flowchart of an embodiment.
[0342] In this embodiment, a case is exemplified where at least two words are recognized from a sentence, and a corresponding partial image is retrieved for each of the at least two words.
[0343] First, in step S631 , a first partial image corresponding to a first word is retrieved from a database.
[0344] Next, in step S632 , a second partial image corresponding to the second word is retrieved from the database.
[0345] Reference Figure 35 This is further explained.
[0346] Reference Figure 35 , illustrating that a first partial image 1011 corresponding to a first word “sits down” and a second partial image 1021 corresponding to a second word “getting up” among previously distinguished and classified words are respectively retrieved from a database (not shown).
[0347] The first partial image 1011 corresponds to the meaning of the first word "sits down" and is an image showing a person sitting down, and the second partial image 1021 corresponds to the meaning of the second word "getting up" and is an image showing a person standing up.
[0348] In addition, when searching for partial images corresponding to the meaning of each word, you can further refer to other words associated with the word to find partial images that better match the context. Figure 36 and Figure 37 This is further explained.
[0349] First, in step S631a, multiple partial images are retrieved from the database based on a first word. For example, if the first word is "sits down," then when searching for partial images matching the meaning of the first word, multiple partial images showing various ways of sitting down can be retrieved.
[0350] For example, Figure 37As shown, for "sits down", multiple partial images 1010 can be retrieved, such as a partial image 1011 of sitting cross-legged, a partial image 1012 of squatting, and a partial image 1013 of sitting sideways. Since all of the retrieved partial images 1011, 1012, and 1013 show an action corresponding to the meaning of "sits down", it is difficult to determine which of the partial images 1011, 1012, and 1013 is the most appropriate partial image based solely on "sits down".
[0351] In step S632b, a first partial image is selected from the plurality of partial images based on the third word. In this case, the third word may be a word associated with the first word. Figure 37 This is further explained.
[0352] Reference Figure 37 The third word "with crossed legs" is identified as a related word of the first word "sits down." Previously, it was difficult to determine the most appropriate partial image among partial images 1011, 1012, and 1013 based solely on the word "sits down." However, by also referring to "with crossed legs," the first partial image 1011 can be determined and selected as the partial image that best fits the context.
[0353] In addition, although the example of using the reference to the related word (ie, the third word) when retrieving the first partial image is described, the related word may be similarly referenced when retrieving the second partial image.
[0354] For example, if there is a fourth word associated with the second word in the sentence and multiple partial images are retrieved for the second word, one of the multiple partial images that best suits the context may be determined and selected with reference to the fourth word.
[0355] Back again Figure 31 In step S640, a motion image corresponding to the sentence is generated using the partial image. Figure 38 This is further explained.
[0356] Figure 38 It shows that Figure 31 The step S640 is a flowchart of the embodiment.
[0357] In step S641, the arrangement order of the first partial image and the second partial image is determined based on the context of the sentence. In this embodiment, based on the context, it is assumed that the first partial image is located immediately before the second partial image.
[0358] In step S642, an intermediate image is generated to be inserted between the first partial image and the second partial image. In this case, the intermediate image may be an image generated to naturally connect the posture of the subject in the first frame of the first partial image and the posture of the subject in the second frame of the second partial image.
[0359] In step S643, a three-dimensional moving image corresponding to the overall meaning of the sentence is generated by sequentially connecting the first partial image, the intermediate image, and the second partial image.
[0360] Reference Figure 39 and Figure 40 This is further explained.
[0361] Reference Figure 39 , shows an example of determining the arrangement order of each partial image based on the context and inserting an intermediate image between the partial images.
[0362] exist Figure 39 In the example, considering the context of the sentence, it can be seen that due to the presence of the preposition "before", the action "sits down" precedes the action "getting up" in time. Therefore, the first partial image 1011 corresponding to "sits down" is placed in front, while the second partial image 1021 corresponding to "getting up" is placed in the back.
[0363] At this time, when viewing the last frame of the first partial image 1011 ( Figure 38 ) and the first frame of the second partial image 1021 ( Figure 38 Therefore, if the first partial image 1011 and the second partial image 1021 are directly connected, an unnatural situation will occur in which the posture of the object is discontinuous at the connection point of the partial images 1011 and 1021.
[0364] Therefore, in order to naturally connect the posture of the object in the last frame of the first partial image 1011 and the posture of the object in the first frame of the second partial image 1021, the intermediate image 1031 is generated and inserted between the first partial image 1011 and the second partial image 1021.
[0365] Reference Figure 40 , shows an example of generating a three-dimensional moving image by sequentially connecting a first partial image, an intermediate image, and a second partial image.
[0366] The generated three-dimensional motion image shows a person sitting down cross-legged according to the meaning of the sentence and then standing up again.
[0367] At this time, it can be seen that by inserting the intermediate image 1031 between the first partial image 1011 and the second partial image 1021 , the posture of the object continues continuously and naturally from the beginning of the first partial image 1011 to the end of the second partial image 1021 .
[0368] Figure 41 It shows that Figure 38 The step S642 is a flowchart of the embodiment. Figure 41 A specific method for generating an intermediate image is illustrated.
[0369] In step S642a, first positions of key points of the object are identified from the first frame of the first partial image.
[0370] In step S642b, second positions of key points of the object are identified from a second frame of the second partial image.
[0371] In step S642c, an intermediate image is generated so that the position of the key point of the object in the intermediate image changes from the first position to the second position.
[0372] Reference Figure 42 This is further explained.
[0373] In order to naturally connect the first partial image 1011 and the second partial image 1021, the key point positions (first positions, P1) of the object in the last frame 1011a of the first partial image 1011 and the key point positions (second positions, P2) of the object in the first frame 1021a of the second partial image 1021 are identified. Here, the key points may be points representing joint parts of the object.
[0374] At this time, intermediate image 1031 is generated such that the key point position of the object in intermediate image 1031 changes from the first position P1 to the second position P2. Generation of intermediate image 1031 can be performed by an artificial intelligence module that, given the first and last frames, has been trained to generate a motion image that connects the pose of the object in the first frame with the pose of the object in the last frame.
[0375] According to this method, even if the posture of the object in the first partial image 1011 is discontinuous with the posture of the object in the second partial image 1021 , the first partial image 1011 and the second partial image 1021 can be naturally connected through the intermediate image.
[0376] According to the embodiments of the present invention described above, when a sentence containing content to be generated as a three-dimensional moving image is input in the form of text or voice, a moving image matching its meaning can be automatically generated by analyzing the sentence.
[0377] Furthermore, the entire moving image can be generated by combining partial images corresponding to the meanings of words in a sentence, wherein appropriate intermediate images can be inserted to allow a natural connection between the partial images.
[0378] Below, we will refer to Figure 43 An exemplary computing device 500 for implementing the methods described in various embodiments of the present invention is described. For example, Figure 43 The computing device 500 in the embodiment may be Figure 4 The three-dimensional moving image generating device 100, Figure 10 The object image generating device 200 shown, Figure 19 The object action similarity calculation device 300, Figure 30 The motion image generating device 400 in.
[0379] Figure 43 is a diagram illustrating an exemplary hardware configuration of the computing device 500 .
[0380] like Figure 43 As shown, the computing device 500 may include one or more processors 510, a bus 550, a communication interface 570, a memory 530 for loading a computer program 591 executed by the processor 510, and a storage device 590 for storing the computer program 591. However, Figure 43 Only components related to the embodiment of the present invention are shown. Therefore, those skilled in the art will understand that, except for Figure 43 In addition to the components shown, other common components may also be included.
[0381] The processor 510 controls the overall operation of the various components of the computing device 500. The processor 510 may be configured to include at least one of a central processing unit (CPU), a microprocessing unit (MPU), a microcontroller unit (MCU), a graphics processing unit (GPU), or any other type of processor known in the art. Furthermore, the processor 510 may execute operations for at least one application or program for performing methods / operations according to various embodiments of the present invention. The computing device 500 may have one or more processors.
[0382] The memory 530 stores various data, commands, and / or information. The memory 530 may load one or more programs 591 from the storage device 590 to execute the methods / operations according to various embodiments of the present invention. An example of the memory 530 may be RAM, but is not limited thereto.
[0383] The bus 550 provides a communication function between the various components of the computing device 500. The bus 550 may be implemented as various types of buses, such as an address bus, a data bus, and a control bus.
[0384] The communication interface 570 supports wired and wireless Internet communications of the computing device 500. The communication interface 570 may also support various communication methods other than Internet communications. To this end, the communication interface 570 may be configured to include a communication module known in the technical field of the present invention.
[0385] The storage device 590 can non-temporarily store one or more computer programs 591. The storage device 590 can be configured to include a non-volatile memory such as a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a hard disk, a removable disk, or any other form of computer-readable recording medium known in the art to which the present invention pertains.
[0386] The computer program 591 may include one or more instructions for implementing methods / operations according to various embodiments of the present invention.
[0387] For example, the computer program 591 may include instructions for performing the following operations: acquiring a two-dimensional image; extracting pose data of an object from the two-dimensional image using an artificial intelligence model that has been machine-learned; generating three-dimensional modeling data based on the pose data; and generating a three-dimensional motion image using the three-dimensional modeling data.
[0388] Alternatively, the computer program 591 may include instructions for performing the following operations: acquiring a first object image and a UV map corresponding to the first object image; and generating a second object image in which the body part of the object is restored based on the first object image and the UV map, wherein the first object image may be an image in which the body part of the object is cut, hidden, or blocked.
[0389] Alternatively, the computer program 591 may include instructions for performing the following operations: inputting a source image and a target image; extracting a first feature vector from a first object corresponding to the source image, and extracting a second feature vector from a second object corresponding to the target image; and calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector, wherein the first object and the second object are three-dimensional mesh objects, and the operation of calculating the action similarity may be an operation of calculating the action similarity using an artificial intelligence module that performs machine learning based on metric learning.
[0390] Alternatively, the computer program 591 may include instructions for: acquiring a sentence; identifying one or more words in the sentence; retrieving partial images corresponding to the one or more words from a database; and generating a motion image corresponding to the sentence using the partial images.
[0391] When the computer program 591 is loaded into the memory 530 , the processor 510 may perform methods / operations according to various embodiments of the present invention by executing one or more instructions.
[0392] The technical concept of the present invention described above can be implemented as computer-readable code on a computer-readable medium. The computer-readable recording medium can be, for example, a removable recording medium (CD, DVD, Blu-ray disc, USB storage device, removable hard disk) or a fixed recording medium (ROM, RAM, computer built-in hard disk). The computer program recorded on the computer-readable recording medium can be transmitted to another computing device via a network such as the Internet and installed in the other computing device, thereby being used in the other computing device.
[0393] Although the embodiments of the present invention have been described with reference to the accompanying drawings, it will be understood by those skilled in the art that the present invention can be implemented in other specific forms without changing its technical concept or essential features. Therefore, it should be understood that the embodiments described above are exemplary in all aspects and not restrictive. The scope of protection of the present invention should be interpreted by the following claims, and all technical concepts within the scope of equivalence should be interpreted as included in the scope of rights of the technical concepts defined in the present invention.
Claims
1. A method for generating a three-dimensional moving image based on artificial intelligence, performed by a computing device, comprising the following steps: Acquire a two-dimensional image; Recognizing an object from the two-dimensional image using an artificial intelligence module that has undergone machine learning, and analyzing movement of the recognized object to generate three-dimensional modeling data of the object; as well as A three-dimensional motion image is generated using the three-dimensional modeling data.
2. The method for generating three-dimensional motion images based on artificial intelligence according to claim 1, wherein: The steps of generating the three-dimensional modeling data include: A first artificial intelligence module is used to extract posture data of the object from the two-dimensional image, and to generate three-dimensional motion data representing the motion of the object based on the posture data.
3. The method for generating three-dimensional motion images based on artificial intelligence according to claim 2, wherein The two-dimensional image includes a plurality of frames representing the motion of the object in a time sequence, The first artificial intelligence module includes: a first submodule based on a convolutional neural network, which extracts posture data of the object from a first frame of the plurality of frames; as well as a second submodule based on a recurrent neural network, which provides information about a second frame of the plurality of frames to the first submodule, The second frame is a frame that precedes the first frame in the time sequence.
4. The method for generating three-dimensional motion images based on artificial intelligence according to claim 2, wherein: The step of generating the three-dimensional motion image comprises: By integrating the three-dimensional motion data into a pre-stored three-dimensional character object, a motion image representing the motion of the three-dimensional character object is generated.
5. The method for generating three-dimensional motion images based on artificial intelligence according to claim 4, wherein: The three-dimensional motion data includes: Key point data corresponding to the joints of the object.
6. The method for generating three-dimensional motion images based on artificial intelligence according to claim 5, wherein: The step of generating a motion image representing the motion of the three-dimensional character object comprises: A redirection step is performed to match the joint positions of the three-dimensional motion data with the joint positions of the three-dimensional character object based on the key point data.
7. The method for generating three-dimensional motion images based on artificial intelligence according to claim 2, wherein: The step of generating the three-dimensional modeling data further includes: A second artificial intelligence module is used to extract shape information of the object from the two-dimensional image, and a three-dimensional mesh object representing the volume of the object is generated based on the shape information.
8. The method for generating three-dimensional motion images based on artificial intelligence according to claim 7, wherein: The step of generating the three-dimensional motion image comprises: The three-dimensional mesh object is incorporated into the three-dimensional motion data to generate a motion image representing the motion of the three-dimensional mesh object.
9. The method for generating three-dimensional motion images based on artificial intelligence according to claim 7, wherein: The second artificial intelligence module includes: a 3D reconstruction module for generating a 3D point cloud representing the object from the 2D image; and A mesh modeling module performs three-dimensional mesh modeling and texture mapping based on the three-dimensional point cloud to generate the three-dimensional mesh object.
10. An artificial intelligence-based three-dimensional motion image generation device, comprising: processor; a memory for loading a computer program for execution by said processor; as well as a storage device storing the computer program, The computer program includes instructions for: Acquire a two-dimensional image; extracting posture data of the object from the two-dimensional image using an artificial intelligence module that has undergone machine learning, and generating three-dimensional modeling data based on the posture data; as well as A three-dimensional motion image is generated using the three-dimensional modeling data.
11. A method for generating an image of an object based on artificial intelligence, performed by a computing device, comprising: Acquire a first object image and a UV map corresponding to the first object image; as well as generating a second object image in which a body part of the object is restored based on the first object image and the UV map, The first object image is an image in which the body part of the object is cut, hidden or blocked.
12. An object image generation device based on artificial intelligence, comprising: processor; a memory for loading a computer program for execution by said processor; as well as a storage device storing the computer program, The computer program includes instructions for performing the following operations: Acquire a first object image and a UV map corresponding to the first object image, and generating a second object image in which a body part of the object is restored based on the first object image and the UV map; The first object image is an image in which the body part of the object is cut, hidden or blocked.
13. A method for calculating the similarity of actions of objects based on artificial intelligence, performed by a computing device, comprising the following steps: Input source image and target image; extracting a first feature vector from a first object corresponding to the source image, and extracting a second feature vector from a second object corresponding to the target image; Calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector, wherein the first object and the second object are three-dimensional mesh objects, The step of calculating the action similarity is: using an artificial intelligence module that performs machine learning based on metric learning to calculate the action similarity.
14. An artificial intelligence object motion similarity calculation device, comprising: processor; a memory for loading a computer program for execution by said processor; as well as a storage device storing the computer program, The computer program includes instructions for performing the following operations: Input source image and target image; extracting a first feature vector from a first object corresponding to the source image, and extracting a second feature vector from a second object corresponding to the target image; Calculating the action similarity between the first object and the second object based on the first feature vector and the second feature vector, wherein the first object and the second object are three-dimensional mesh objects, The operation of calculating the action similarity is: using an artificial intelligence module that performs machine learning based on metric learning to calculate the action similarity.
15. A method for generating a moving image from a sentence, performed by a computing device, comprising the steps of: Get the sentence; identifying one or more words in the sentence; Retrieving a partial image corresponding to the one or more words from a database; as well as A moving image corresponding to the sentence is generated using the partial image.
16. A motion image generating apparatus, comprising: processor; a memory for loading a computer program for execution by said processor; as well as a storage device storing the computer program, The computer program includes instructions for performing the following operations: Get the sentence; identifying one or more words in the sentence; Retrieving a partial image corresponding to the one or more words from a database; and A moving image corresponding to the sentence is generated using the partial image.