Apparatus for generating three-dimensional human model on basis of photograph by using artificial intelligence, and method thereof

The AI-based 3D modeling device preprocesses images to estimate dimensions and features, generating a complete 3D human model efficiently from a single photograph, addressing cost and space issues in traditional methods.

WO2025159266A1PCT designated stage expired Publication Date: 2025-07-31CREADTO INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/014277
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2024-09-23
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing 3D modeling techniques require expensive equipment to capture images from multiple angles, increasing cost and space requirements, making them inefficient for generating photo-based 3D human models.

Method used

A device and method using artificial intelligence to preprocess images, estimate human body and facial dimensions, extract features, and perform AI-based learning to generate a 3D model, fitting it with joint feature information for a complete 3D human model from a single photograph.

Benefits of technology

Enables easy generation of a full-body 3D human model with high-resolution texture and no shadows or diffuse reflections, applicable in a 1:1 virtual environment, reducing costs and space requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024014277_31072025_PF_FP_ABST
    Figure KR2024014277_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention performs preprocessing on image information including a user, estimates a human body size and a face size for the user which is an object included in the preprocessed image information, extracts a human body feature and a face feature for the user which is the object included in the preprocessed image information, performs learning based on artificial intelligence on the basis of the estimated human body size and face size so as to generate a preliminary three-dimensional model related to the user, performs different learning based on different artificial intelligence on the basis of the extracted human body feature and face feature and the generated preliminary three-dimensional model so as to generate a modified three-dimensional model related to the user, and finally generates a three-dimensional human model related to the user by fitting the generated modified three-dimensional model and preconfigured joint feature information, whereby a full-body three-dimensional human model can be simply generated by using a single photograph, the generated three-dimensional human model can be directly applied to a one-to-one virtual environment as a human body size is reflected, there is no shadow or diffused reflection by an image processing module in a process of generating a three-dimensional human model, and a texture having high resolution can be generated by using an upscaling technique.
Need to check novelty before this filing date? Find Prior Art

Description

Device and method for generating a 3D human model based on a photo using artificial intelligence

[0001] The present invention relates to a device and method for generating a photo-based 3D human model using artificial intelligence, and more particularly, to a device and method for generating a photo-based 3D human model using artificial intelligence, which performs preprocessing on image information including a user, estimates human body dimensions and facial dimensions for the user as an object included in the preprocessed image information, extracts human body features and facial features for the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, performs other artificial intelligence-based learning based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and fits the generated modified 3D model with preset joint feature information to finally generate a 3D human model related to the user.

[0002] 3D modeling (three-dimensional modeling) is the process of creating the basic skeleton of a three-dimensional image, and is the process of creating a three-dimensional object using various methods such as using geometric objects or curves.

[0003] For this type of 3D modeling, expensive equipment is used to capture images of the object from various angles and positions, and multiple images corresponding to the captured angles and positions are used to create a 3D image of the object, which increases the cost and requires a place to install the equipment.

[0004] An object of the present invention is to provide a device and method for generating a photo-based 3D human model using artificial intelligence, which performs preprocessing on image information including a user, estimates human body dimensions and facial dimensions for the user as an object included in the preprocessed image information, extracts human body features and facial features for the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, performs other artificial intelligence-based learning based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and fits the generated modified 3D model with preset joint feature information to finally generate a 3D human model related to the user.

[0005] A device for generating a photo-based 3D human model using artificial intelligence according to an embodiment of the present invention may include: a camera unit for acquiring image information including a user; a control unit for performing preprocessing on the acquired image information, estimating human body dimensions and facial dimensions for the user as an object included in the preprocessed image information, extracting human body features and facial features for the user as an object included in the preprocessed image information, performing artificial intelligence-based learning based on the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information, thereby generating a preliminary 3D model related to the user based on the learning results; performing other artificial intelligence-based learning based on the extracted human body features, the extracted facial features, and the generated preliminary 3D model, thereby generating a modified 3D model related to the user based on the other learning results, and fitting the generated modified 3D model and preset joint feature information to generate a 3D human model related to the user; and a display unit for displaying the generated 3D human model.

[0006] As an example related to the present invention, the control unit may perform preprocessing including at least one of noise removal, background removal, glare removal, face image quality enhancement through upscaling, and depth estimation on the acquired image information.

[0007] As an example related to the present invention, the control unit may generate a body-related model based on the estimated human body dimensions and the preprocessed image information, generate a face-related model based on the estimated face dimensions and the preprocessed image information, and synthesize the generated body-related model and the generated face-related model to generate a preliminary 3D model related to the user.

[0008] A method for generating a photo-based 3D human model using artificial intelligence according to an embodiment of the present invention comprises the steps of: acquiring image information including a user by a camera unit; performing preprocessing by a control unit on the acquired image information; estimating, by the control unit, human body dimensions and facial dimensions of a user as an object included in the preprocessed image information; extracting, by the control unit, human body features and facial features of a user as an object included in the preprocessed image information; performing, by the control unit, artificial intelligence-based learning based on the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information, and generating a preliminary 3D model related to the user based on the learning results; performing, by the control unit, other artificial intelligence-based learning based on the extracted human body features, the extracted facial features, and the generated preliminary 3D model, and generating a modified 3D model related to the user based on the other learning results; fitting, by the control unit, the generated modified 3D model and preset joint feature information to generate a 3D human model related to the user; And, by the control unit, it may include a step of displaying the generated three-dimensional human model on the display unit.

[0009] As an example related to the present invention, the step of estimating the human body size and facial size may include a process of recognizing an object included in the preprocessed image information; a process of estimating the human body size for the recognized object; and a process of estimating the facial size for the user's face within the recognized object.

[0010] As an example related to the present invention, the step of extracting human body features and facial features may include a process of recognizing an object included in the preprocessed image information; a process of extracting human body features for the recognized object; and a process of extracting facial features for the user's face within the recognized object.

[0011] As an example related to the present invention, the step of generating a preliminary 3D model related to the user may include performing learning using the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information as input values ​​of a preset 3D generation model, and generating a preliminary 3D model related to the user based on the learning result.

[0012] The present invention performs preprocessing on image information including a user, estimates human body dimensions and facial dimensions for the user as an object included in the preprocessed image information, extracts human body features and facial features for the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, performs other artificial intelligence-based learning based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and fits the generated modified 3D model with preset joint feature information to finally generate a 3D human model related to the user, thereby enabling the easy generation of a full-body 3D human model with a single photograph, and the 3D human model generated by reflecting human body dimensions can be immediately applied in a 1:1 virtual environment, and in the process of generating the 3D human model, there are no shadows or diffuse reflections due to the image processing module, and there is an effect of being able to generate a high-resolution texture using an upscaling technique.

[0013] FIG. 1 is a block diagram showing the configuration of a device for generating a photo-based 3D human model using artificial intelligence according to an embodiment of the present invention.

[0014] FIG. 2 is a flowchart illustrating a method for creating a photo-based 3D human model using artificial intelligence according to an embodiment of the present invention.

[0015] FIG. 3 is a diagram showing an example of image information according to an embodiment of the present invention.

[0016] Figures 4 and 5 are diagrams showing examples of three-dimensional human models according to embodiments of the present invention.

[0017] It should be noted that the technical terms used in the present invention are used merely to describe specific embodiments and are not intended to limit the present invention. Furthermore, unless specifically defined otherwise herein, the technical terms used herein should be interpreted as having a meaning generally understood by those skilled in the art to which the present invention pertains, and should not be interpreted in an overly comprehensive or overly narrow sense. Furthermore, if a technical term used herein is incorrect and fails to accurately express the spirit of the present invention, it should be replaced with a technical term that can be correctly understood by those skilled in the art. Furthermore, general terms used herein should be interpreted according to their dictionary definitions or according to the context, and should not be interpreted in an overly narrow sense.

[0018] Additionally, singular expressions used in the present invention include plural expressions unless the context clearly dictates otherwise. Terms such as "consist of" or "include" in the present invention should not necessarily be construed to include all of the components or steps described in the invention, and should be construed to mean that some of the components or steps may not be included, or that additional components or steps may be included.

[0019] Additionally, terms including ordinal numbers, such as "first" and "second," used in the present invention may be used to describe components, but the components should not be limited by these terms. The terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component could be referred to as a second component, and similarly, a second component could also be referred to as a first component.

[0020] Hereinafter, a preferred embodiment of the present invention will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components are given the same reference numbers and redundant descriptions thereof will be omitted.

[0021] Furthermore, when describing the present invention, detailed descriptions of related known technologies will be omitted if they are deemed to obscure the gist of the present invention. Furthermore, it should be noted that the attached drawings are intended solely to facilitate understanding of the spirit of the present invention and should not be construed as limiting the spirit of the present invention.

[0022] FIG. 1 is a block diagram showing the configuration of a photo-based 3D human model generation device (100) using artificial intelligence according to an embodiment of the present invention.

[0023] As illustrated in FIG. 1, the device (100) for generating a photo-based 3D human model using artificial intelligence is composed of a camera unit (110), a communication unit (120), a storage unit (130), a display unit (140), a voice output unit (150), and a control unit (160). Not all of the components of the device (100) for generating a photo-based 3D human model using artificial intelligence shown in FIG. 1 are essential components, and the device (100) for generating a photo-based 3D human model using artificial intelligence may be implemented with more components than the components illustrated in FIG. 1, or the device (100) for generating a photo-based 3D human model using artificial intelligence may be implemented with fewer components.

[0024] The above artificial intelligence-based 3D human model generation device (100) is a smart phone, a portable terminal, a mobile terminal, a foldable terminal, a personal digital assistant (PDA), a portable multimedia player (PMP) terminal, a telematics terminal, a navigation terminal, a personal computer, a notebook computer, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, a head mounted display (HMD), etc.), a Wibro terminal, an IPTV (Internet Protocol Television) terminal, a smart TV, a digital broadcasting terminal, an AVN (Audio Video Navigation) terminal, an A / V (Audio / Video) system, a flexible terminal, a digital signage device, It can be applied to various terminals such as home theater systems, information centers, and call centers.

[0025] The above camera unit (or photographing unit) (110) is configured in an independent form or is configured as a part of the 3D human model creation device (100). In this case, when the camera unit (110) is configured outside the 3D human model creation device (100), the camera unit (110) may further include a communication unit for communicating with the 3D human model creation device (100) (or the communication unit (120)).

[0026] In addition, the camera unit (110) is configured with one or more image sensors (camera modules or cameras) so as to be able to capture the front (or front / front full body) including the front, back, side, bottom, and top of the user (or photographer). At this time, the camera unit (110) may also be configured with a stereo camera, depth camera, etc. capable of acquiring image information in all directions of 360 degrees.

[0027] In addition, the camera unit (110) acquires (or photographs) image information (or images / photos / still images / videos) including any user.

[0028] In addition, the one or more camera units (110) process image frames such as still images or moving images obtained by an image sensor (camera module or camera) in video call mode, shooting mode, video conference mode, etc. That is, the corresponding image data obtained by the image sensor is encoded / decoded according to each standard according to the CODEC. For example, the camera unit (110) photographs an object (or subject) and outputs a video signal corresponding to the photographed image (subject image).

[0029] In addition, the image frame (or image information) processed in the camera unit (110) may be stored in a digital video recorder (DVR), stored in the storage unit (130), or transmitted to an external server, etc. via the communication unit (120).

[0030] Additionally, the camera unit (110) acquires (or captures) an image (or image information) displayed in a preview item (or viewfinder item).

[0031] The above communication unit (120) communicates with any internal component or at least one external terminal via a wired / wireless communication network. At this time, the external terminal may include a server (not shown), a terminal (not shown), etc. Here, wireless Internet technologies include Wireless LAN (WLAN), Digital Living Network Alliance (DLNA), Wireless Broadband (Wibro), World Interoperability for Microwave Access (Wimax), High Speed ​​Downlink Packet Access (HSUPA), High Speed ​​Uplink Packet Access (HSUPA), IEEE 802.16, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), LTE-M (LTE based maritime wireless communication), Wireless Mobile Broadband Service (WMBS), 5G network / 5G communication network, 6G network / 6G communication network, Wireless Smart Utility Network (Wi-SUN), Narrowband Internet of Things (NB-IoT), etc., and the communication unit (120) includes at least one Internet technology not listed above. Data is transmitted and received based on wireless Internet technology.In addition, short-range communication technologies may include Bluetooth, Bluetooth Low Energy (BLE), ANT, ANT+, LoRa (Long Range), RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra Wideband), ZigBee, Near Field Communication (NFC), Ultra Sound Communication (USC), Visible Light Communication (VLC), Wi-Fi, Wi-Fi Direct, Magnetic Secure Transmission (MST), Beacon, EnOcean, Near Field Magnetic Induction (NFMI), Z-WAVE, and SIGFOX. In addition, wired communication technologies may include Power Line Communication (PLC), USB communication, This may include Ethernet, serial communication, optical / coaxial cables, etc.

[0032] In addition, the communication unit (120) can mutually transmit information to any terminal via a universal serial bus (USB).

[0033] In addition, the communication unit (120) transmits and receives wireless signals with a base station, the server, the terminal, etc. on a mobile communication network constructed according to technical standards or communication methods for mobile communication (e.g., GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), CDMA2000 (Code Division Multi Access 2000), EV-DO (Enhanced Voice-Data Optimized or Enhanced Voice-Data Only), WCDMA (Wideband CDMA), HSDPA (High Speed ​​Downlink Packet Access), HSUPA (High Speed ​​Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G network / 5G communication network, 6G network / 6G communication network, etc.).

[0034] In addition, the communication unit (120) transmits image information including the corresponding user obtained through the camera unit (110) to the server, the terminal, etc., under the control of the control unit (160).

[0035] In addition, the communication unit (120) receives image information including any user transmitted from another terminal (not shown) under the control of the control unit (160).

[0036] The above storage unit (130) stores various user interfaces (UI), graphical user interfaces (GUI), etc.

[0037] Additionally, the storage unit (130) stores data and programs necessary for the operation of the three-dimensional human model creation device (100).

[0038] That is, the storage unit (130) can store a plurality of application programs (or applications) run in the 3D human model generation device (100), data for the operation of the 3D human model generation device (100), and commands. At least some of these application programs can be downloaded from an external server via wireless communication. In addition, at least some of these application programs can exist on the 3D human model generation device (100) from the time of shipment for the basic functions of the 3D human model generation device (100). Meanwhile, the application programs can be stored in the storage unit (130), installed in the 3D human model generation device (100), and driven by the control unit (160) to perform the operation (or function) of the 3D human model generation device (100).

[0039] In addition, the storage unit (130) may include at least one storage medium among a Flash Memory Type, a Hard Disk Type, a Multimedia Card Micro Type, a card type memory (e.g., SD memory, XD memory, CF (compact flash) memory, etc.), a stick type memory stick, a magnetic memory, a magnetic disk, an optical disk, a Random Access Memory (RAM), a Static Random Access Memory (SRAM), a Read-Only Memory (ROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Programmable Read-Only Memory (PROM), a one time programmable ROM (OTPROM), a mask ROM, and a flash ROM. In addition, the 3D human model creation device (100) may operate a web storage that performs the storage function of the storage unit (130) on the Internet, or may operate in relation to the web storage.

[0040] In addition, the storage unit (130) stores image information including the user acquired through the camera unit (110) under the control of the control unit (160).

[0041] The above display unit (or display unit) (140) can display various contents, such as various menu screens, using the user interface and / or graphical user interface stored in the storage unit (130) under the control of the control unit (160). Here, the contents displayed on the display unit (140) include various text or image data (including various information data) and menu screens including data such as icons, list menus, and combo boxes. In addition, the display unit (140) may be a touch screen.

[0042] In addition, the display unit (140) may include at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, a 3D display, an e-ink display, a light emitting diode (LED), a beam projector, a goggle-type VR, a hologram, and a HUD (Head Up Display). Here, when the display unit (140) is implemented as a HUD, the display unit (140) may include a projection module to output information through an image projected onto a windshield or a window.

[0043] In addition, the display unit (140) may be implemented as a touch screen by forming a mutual layer structure with the touch input unit (not shown) or forming it as an integral part.

[0044] Additionally, the display unit (140) may include a transparent display. The transparent display may be attached to a windshield or window.

[0045] In addition, the transparent display can display a predetermined screen while having a predetermined transparency. In order to have transparency, the transparent display can include at least one of a transparent TFEL (Thin Film Electroluminescent), a transparent OLED (Organic Light-Emitting Diode), a transparent LCD (Liquid Crystal Display), a transparent display, and a transparent LED (Light Emitting Diode) display. The transparency of the transparent display can be adjusted.

[0046] In addition, the display unit (140) displays image information acquired by the camera unit (110) through the viewfinder screen under the control of the control unit (160).

[0047] In addition, the display unit (140) displays image information including the user obtained through the camera unit (110) under the control of the control unit (160).

[0048] The above voice output unit (150) outputs voice information included in a signal processed by the control unit (160). Here, the voice output unit (150) may include a receiver, a speaker, a buzzer, etc.

[0049] Additionally, the voice output unit (150) outputs the guidance voice generated by the control unit (160).

[0050] In addition, the voice output unit (150) outputs voice information (or sound effects) corresponding to image information including the user acquired through the camera unit (110) under the control of the control unit (160).

[0051] The above control unit (controller, or MCU (microcontroller unit)) (160) executes the overall control function of the artificial intelligence-based photo-based 3D human model creation device (100).

[0052] In addition, the control unit (160) executes the overall control function of the 3D human model creation device (100) using the program and data stored in the storage unit (130). The control unit (160) may include a RAM, a ROM, a CPU, a GPU, and a bus, and the RAM, ROM, CPU, GPU, etc. may be connected to each other through a bus. The CPU may access the storage unit (130) and perform booting using the O / S stored in the storage unit (130), and may perform various operations using various programs, contents, data, etc. stored in the storage unit (130).

[0053] In addition, the control unit (160) utilizes a plurality of image information (or a plurality of preprocessed image information) including a user collected in advance, an estimated human body size (or a user-specific human body size) related to the user, an estimated facial size (or a user-specific facial size) related to the user, an extracted human body feature (or a user-specific human body feature) related to the user, an extracted facial feature (or a user-specific facial feature) related to the user, a preliminary 3D model related to the user, preset joint feature information, a modified 3D model related to the user, etc., as data for continuous learning (or machine learning / deep learning). Here, the input dataset for learning may be divided into a training set and a test set, including the plurality of image information (or the plurality of preprocessed image information), the estimated human body size (or the user-specific human body size) related to the user, the estimated facial size (or the user-specific facial size) related to the user, the extracted human body features (or the user-specific human body features) related to the user, the extracted facial features (or the user-specific facial features) related to the user, the preliminary 3D model related to the user, the preset joint feature information, the modified 3D model related to the user, etc., at a preset ratio (e.g., including 7:3, 8:2, etc.), so as to perform training and testing functions. In addition, the input dataset for the above learning includes multiple image information (or multiple preprocessed image information) collected later, human body dimensions estimated in relation to the user (or human body dimensions by user), facial dimensions estimated in relation to the user (or facial dimensions by user), human body features extracted in relation to the user (or human body features by user), facial features extracted in relation to the user (or facial features by user), a preliminary 3D model in relation to the user, a modified 3D model in relation to the user, etc.In addition, the output dataset for the above learning includes a portion to be predicted, and includes a plurality of image information (or a plurality of preprocessed image information), estimated human body dimensions (or user-specific human body dimensions) related to the user, estimated facial dimensions (or user-specific facial dimensions) related to the user, and a preliminary 3D model related to the user, which is learned and later predicted based on the extracted human body features (or user-specific human body features) related to the user, extracted facial features (or user-specific facial features) related to the user, and a preliminary 3D model related to the user, which is learned and later predicted based on the extracted human body features (or user-specific human body features), and a preliminary 3D model related to the user, and a 3D human model related to the user, which is learned and later predicted based on preset joint feature information, the modified 3D model related to the user, and a 3D model related to the user, etc.

[0054] That is, the control unit (160) performs a learning function to generate (or predict / classify) a preliminary 3D model related to a specific user, for example, for specific image information (or a plurality of preprocessed image information) related to specific raw data, for a 3D generation model, for estimated human body dimensions (or human body dimensions by user), for estimated facial dimensions (or facial dimensions by user) related to a specific user, through preset learning data.

[0055] In addition, the control unit (160) performs a learning function to generate (or predict / classify) a modified 3D model (or modified 3D model / 3D modified model) related to a specific user, for example, human body features (or user-specific human body features) extracted related to a specific user, facial features (or user-specific facial features) extracted related to a specific user, and a preliminary 3D model related to a specific user previously generated, in relation to a shape modification model through the preset learning data.

[0056] In addition, the control unit (160) performs a learning function to create (or predict / classify) a final 3D human model related to a specific user, based on the preset joint characteristic information related to specific raw data, the modified 3D model related to a specific user previously created, etc., for the joint control model through the preset learning data. At this time, the control unit (160) stores raw data (or specific image information of a component (or a plurality of preprocessed image information), estimated human body dimensions (or human body dimensions per user) related to a specific user, estimated facial dimensions (or facial dimensions per user) related to a specific user, extracted human body features (or user-specific human features) related to a specific user, extracted facial features (or user-specific facial features) related to a specific user, a preliminary 3D model related to a specific user generated previously, preset joint feature information, a modified 3D model related to a specific user generated previously, etc.) in parallel and in a distributed manner, refines unstructured data, structured data, and semi-structured data included in the stored raw data (or including data for learning, etc.), performs preprocessing including classification as metadata, performs analysis including data mining on the preprocessed data, and performs learning, training, and testing based on at least one type of machine learning to build big data. At this time, at least one type of machine learning may be one or a combination of at least one of supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and deep reinforcement learning.And data mining can include performing classification, which predicts the class of new data by learning a training data set whose classes are known by exploring the inherent relationships between preprocessed data, or clustering, which groups data based on similarity without class information.

[0057] In this way, the control unit (160) performs a learning function for the 3D creation model, the shape modification model, the joint control model, etc. in the form of a neural network or an artificial neural network through the learning data, etc.

[0058] In addition, the control unit (160) controls the camera unit (110) to obtain (or photograph) image information including the user (or the frontal full body of the user) through the camera unit (100).

[0059] In addition, the control unit (160) performs preprocessing (or preprocessing function) on the acquired (or photographed / received) image information.

[0060] That is, the control unit (160) performs preprocessing, such as noise removal, background removal, glare removal, face image quality improvement through upscaling, and depth estimation, on the acquired image information.

[0061] At this time, the control unit (160) recognizes an object (or user) included in the acquired image information.

[0062] Additionally, the control unit (160) performs upscaling on the face (or head) of the recognized object.

[0063] Additionally, the control unit (160) removes the diffuse reflection of the corresponding image information.

[0064] Additionally, the control unit (160) removes the background excluding the object from the image information in order to focus only on the object in the image information.

[0065] In this way, the control unit (160) can perform various preprocessing functions on the acquired image information.

[0066] In addition, the control unit (160) estimates the human body dimensions and facial dimensions of the user, which is an object included in the preprocessed image information.

[0067] That is, the control unit (160) recognizes an object included in the preprocessed image information and estimates (or measures / confirms) the human body dimensions of the recognized object. Here, the above human body measurements are the height of the fist extended above the head, height, height of the back of the neck, shoulder height, armpit height, waist reference line height (female), waist height, upper front iliac spine height, groin height, lateral malleolus height, fist height, elbow height (bent arm), gas width, waist width, hip width, width below the calf (ankle width), armpit thickness, chest thickness, waist thickness, neck circumference, back length under the neck, neck circumference, armpit circumference, uneven arm circumference, uneven elbow circumference, wrist circumference, upper arm circumference (bent arm), chest circumference, breast circumference, breast circumference (female), waist circumference, waist circumference at navel level, hip circumference, thigh circumference, knee circumference, calf circumference, back waist circumference line length at the side of the neck, side of the neck nipple length, It includes the length of the nipple-waistline at the side of the neck, the length of the front and back of the instep, the length of the back of the shoulder neck (left), the length between the back of the neck and the shoulders, the length of the upper arm, the length of the arm, the sitting height, the thickness of the sitting stomach, etc. At this time, the instep height refers to the vertical distance from the bottom of the instep to the floor in a standing position, the lateral malleolus height refers to the height of the malleolus on the outside of the ankle, and the thigh circumference refers to the width of the part of the leg above the knee joint.

[0068] In addition, the control unit (160) estimates (or measures / confirms) the facial dimensions of the user's face (or head) within the recognized object. Here, the facial dimensions include head circumference, head side arc length, head front arc length, head thickness, head width, ear bead width, head vertical length, face vertical length, left ear height position, right ear height position, left ear depth position, right ear depth position, left ear length, right ear depth, left chin length, right chin length, etc. At this time, the ear bead width represents the width of the area where the cartilage protrudes in front of the ear canal. At this time, the control unit (160) may estimate the human body size, face size, etc. in the form of text, such as 'short arms and long face', and when the human body size, face size, etc. estimated in the form of text in this way are used in the learning process, the text may be applied (or reflected) to the values ​​of each item included in the human body size, face size, etc. according to a preset ratio corresponding to the text (for example, adjusting the human body size items, face size items, etc. according to the state of having short arms and a long face).

[0069] In addition, the control unit (160) extracts (or measures) human body features and facial features for the user, which is an object included in the preprocessed image information.

[0070] That is, the control unit (160) recognizes an object included in the preprocessed image information and extracts (or verifies) human body features for the recognized object. Here, the human body features include metaphysical properties, such as the shape and texture of muscles related to the object.

[0071] In addition, the control unit (160) extracts (or verifies) facial features for the user's face within the recognized object. Here, the facial features include metaphysical properties, such as the appearance of the face, muscle shape, and muscle texture, related to the object.

[0072] In addition, the control unit (160) performs artificial intelligence-based learning (artificial neural network / machine learning / deep learning) based on the estimated human body dimensions, the estimated facial dimensions, the preprocessed image information, etc., and creates a preliminary 3D model related to the user based on the learning results.

[0073] That is, the control unit (160) performs learning (or artificial intelligence / machine learning / deep learning) using the estimated human body dimensions, the estimated facial dimensions, the preprocessed image information, etc. as input values ​​of a preset 3D generation model, and generates a preliminary 3D model related to the user based on the learning result (or artificial intelligence result / machine learning result / deep learning result). At this time, the control unit (160) generates a body-related model based on the estimated human body dimensions, the preprocessed image information, etc., and generates a face-related model based on the estimated facial dimensions, the preprocessed image information, etc., and then synthesizes the generated body-related model and the generated face-related model to generate a preliminary 3D model related to the user.

[0074] In addition, when generating a face related to the user included in the preliminary 3D model, the control unit (160) generates the user's facial features, including the shape of the face, preset characteristics of Asians (e.g., Sizekorea data reflecting the body dimensions of Koreans), data of AI HUB, etc.) by reflecting the estimated facial dimensions.

[0075] In addition, the control unit (160) extracts a plurality of preset facial landmarks (e.g., 468 facial landmarks) (or coordinate / position information corresponding to the plurality of facial landmarks) from the user's face included in the preprocessed image information, separates a plurality of facial landmarks left and right from the plurality of extracted facial landmarks centered on the preset nose center point, and generates two images that are flipped left and right through one or more of the facial landmarks separated left and right, and then performs an alignment process for the two images respectively generated so that they face the front, and performs learning using the two images aligned according to the alignment process, the estimated facial dimensions, etc. as input values ​​of the 3D generation model, and based on the learning result, generates a model related to the face implemented in a left / right asymmetrical or symmetrical state.

[0076] In addition, through an analysis function for the above-mentioned preprocessed image information, it is possible to estimate the male / female gender and create a model related to the face that reflects the male or female characteristics according to the male / female gender estimation.

[0077] In the embodiment of the present invention, for high accuracy, it is mainly described that the preliminary 3D model is generated according to the performance of the learning function by using image information of the user's full body state looking straight ahead, but it is not limited thereto, and the control unit (160) may also generate the preliminary 3D model according to the performance of the learning function by using image information including a part of the user's body, the user's oblique state (or side view of the face), etc.

[0078] In this way, when using image information that includes a part of the user's body, a slanted state of the user, etc., rather than a full frontal view, the control unit (160) may set each item included in the unestimated human body size, face size, etc., to a preset reference value, or may create the preliminary 3D model using a value corrected based on the remaining estimated items.

[0079] In addition, the control unit (160) performs other artificial intelligence-based learning (other artificial neural network / other machine learning / other deep learning) based on the extracted human body features, the extracted facial features, the generated preliminary 3D model, etc., and generates a modified 3D model related to the user based on other learning results.

[0080] That is, the control unit (160) performs other learning (or other artificial intelligence / other machine learning / other deep learning) using the extracted human body features, the extracted facial features, the generated preliminary 3D model, etc. as input values ​​of a preset shape modification model, and generates a modified 3D model related to the user based on the other learning results (or other artificial intelligence results / other machine learning results / other deep learning results). At this time, the generated modified 3D model may be composed of points and surfaces.

[0081] In addition, the control unit (160) fits (or maps / matches / links) preset joint feature information to the generated modified 3D model, thereby ultimately generating a 3D human model related to the user. Here, the preset joint feature information includes information on features for each joint set as default for a plurality of joints that make up the human body (e.g., including part, size, function, etc.).

[0082] That is, the control unit (160) fits the preset joint characteristic information to the generated modified 3D model, thereby generating a 3D human model related to the user that includes joint characteristic information corresponding to the generated modified 3D model. Here, the generated 3D human model may be in the form of a mesh, which is a solid model.

[0083] At this time, the control unit (160) performs another artificial intelligence-based learning (another artificial neural network / another machine learning / another deep learning) based on the preset joint characteristic information, the generated modified 3D model, etc., and finally generates a 3D human model related to the user based on the results of another learning.

[0084] That is, in order to move the modified 3D model composed of the points and surfaces, the control unit (160) performs another learning (or another artificial intelligence / another machine learning / another deep learning) using the preset joint feature information, the generated modified 3D model, etc. as input values ​​of the preset joint control model, and creates a 3D human model related to the user based on another learning result (or another artificial intelligence result / another machine learning result / another deep learning result).

[0085] In addition, the control unit (160) displays (or outputs) the generated three-dimensional human model in the form of a mesh on the display unit (140).

[0086] In the embodiment of the present invention, for high accuracy and precision, the sequential creation of a preliminary 3D model, a modified 3D model, and a 3D human model is mainly described, but is not limited thereto, and the control unit (160) may perform learning through artificial intelligence based on human body dimensions and facial dimensions estimated from the image information, human body features and facial features extracted from the corresponding image information, preset joint information, etc., and ultimately create a 3D human model at once.

[0087] In addition, although the embodiment of the present invention mainly describes generating a 3D human model related to a user (or person) for a single piece of image information including the user, it is not limited thereto, and the control unit (160) may generate a 3D human model related to all users included in a plurality of pieces of image information, each piece of image information including one or more users, and may display the 3D human models related to all users individually or as a whole on the display unit (140).

[0088] In addition, although the embodiment of the present invention mainly describes displaying the generated 3D human model through the display unit (140), it is not limited thereto, and the control unit (160) can be configured to upload the generated 3D human model to a cloud server (not shown), and allow any user to access (or connect / communicate) the cloud server from a terminal (not shown) possessed by the user, and check (or display / output) the 3D human model provided by the cloud server.

[0089] In addition, in the embodiment of the present invention, the device (100) for generating a photo-based 3D human model using artificial intelligence can perform various functions (e.g., a preprocessing function for image information, a function for estimating human dimensions and facial dimensions for image information, a function for extracting human features and facial features for image information, a function for generating a preliminary 3D model, a function for generating a modified 3D model, a function for generating a 3D human model, etc.) provided by the server (not shown) in the form of a dedicated app or website.

[0090] In this way, preprocessing is performed on image information including a user, human body dimensions and facial dimensions are estimated for the user as an object included in the preprocessed image information, human body features and facial features are extracted for the user as an object included in the preprocessed image information, artificial intelligence-based learning is performed based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, and other artificial intelligence-based learning is performed based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and the generated modified 3D model is fitted with preset joint feature information to finally generate a 3D human model related to the user.

[0091] Hereinafter, a method for generating a photo-based 3D human model using artificial intelligence according to the present invention will be described in detail with reference to FIGS. 1 to 5.

[0092] FIG. 2 is a flowchart illustrating a method for creating a photo-based 3D human model using artificial intelligence according to an embodiment of the present invention.

[0093] First, the camera unit (110) acquires (or captures) image information (or images / photos / still images / videos) including the user. Here, the camera unit (110) can capture image information including the user's full frontal view. At this time, the communication unit (120) can also receive image information including any user transmitted from another terminal (not shown).

[0094] For example, as illustrated in FIG. 3, the first camera unit (110) acquires first image information including the first user (S210).

[0095] Thereafter, the control unit (160) performs preprocessing (or preprocessing function) on the acquired (or photographed / received) image information.

[0096] That is, the control unit (160) performs preprocessing such as noise removal, background removal, diffuse reflection removal, face image quality improvement through upscaling, and depth estimation on the acquired image information. At this time, the control unit (160) recognizes an object (or user) included in the captured image information, performs upscaling on the face (or head) of the recognized object, removes diffuse reflection of the corresponding image information, and removes the background excluding the corresponding object from the corresponding image information in order to focus only on the object in the corresponding image information.

[0097] For example, the first control unit (160) performs preprocessing, such as noise removal, background removal, glare removal, face image quality improvement through upscaling, and depth estimation, on the acquired first image information (S220).

[0098] Thereafter, the control unit (160) estimates the human body dimensions and facial dimensions of the user, which is an object included in the preprocessed image information.

[0099] That is, the control unit (160) recognizes an object included in the preprocessed image information and estimates (or measures / confirms) the human body dimensions of the recognized object. Here, the above human body measurements are the height of the fist extended above the head, height, height of the back of the neck, shoulder height, armpit height, waist reference line height (female), waist height, upper front iliac spine height, groin height, lateral malleolus height, fist height, elbow height (bent arm), gas width, waist width, hip width, width below the calf (ankle width), armpit thickness, chest thickness, waist thickness, neck circumference, back length under the neck, neck circumference, armpit circumference, uneven arm circumference, uneven elbow circumference, wrist circumference, upper arm circumference (bent arm), chest circumference, breast circumference, breast circumference (female), waist circumference, waist circumference at navel level, hip circumference, thigh circumference, knee circumference, calf circumference, back waist circumference line length at the side of the neck, side of the neck nipple length, Includes the length of the neck-nipple waistline, the length of the front and back of the crotch, the length of the back of the neck (left), the length between the back of the neck and the shoulders, the length of the upper arm, the length of the arm, the sitting height, and the thickness of the sitting stomach.

[0100] In addition, the control unit (160) estimates (or measures / confirms) the facial dimensions of the user's face (or head) within the recognized object. Here, the facial dimensions include head circumference, side head length, front head length, head thickness, head width, ear bead width, head vertical length, face vertical length, left ear height position, right ear height position, left ear depth position, right ear depth position, left ear length, right ear depth, left chin length, right chin length, etc.

[0101] For example, the first control unit recognizes the first user, which is an object included in the preprocessed first image information.

[0102] Additionally, the first control unit estimates a first human body dimension related to the recognized first user.

[0103] Additionally, the first control unit estimates a first facial dimension related to the recognized face of the first user (S230).

[0104] Thereafter, the control unit (160) extracts human body features and facial features for the user, which is an object included in the preprocessed image information.

[0105] That is, the control unit (160) recognizes an object included in the preprocessed image information and extracts (or verifies) human body features for the recognized object. Here, the human body features include metaphysical properties, such as the shape and texture of muscles related to the object.

[0106] In addition, the control unit (160) extracts (or verifies) facial features for the user's face within the recognized object. Here, the facial features include metaphysical properties, such as the appearance of the face, muscle shape, and muscle texture, related to the object.

[0107] For example, the first control unit extracts a first human body feature related to the recognized first user.

[0108] Additionally, the first control unit extracts a first facial feature related to the recognized face of the first user (S240).

[0109] Thereafter, the control unit (160) performs artificial intelligence-based learning (artificial neural network / machine learning / deep learning) based on the estimated human body dimensions, the estimated facial dimensions, the preprocessed image information, etc., and creates a preliminary 3D model related to the user based on the learning results.

[0110] That is, the control unit (160) performs learning (or artificial intelligence / machine learning / deep learning) using the estimated human body dimensions, the estimated facial dimensions, the preprocessed image information, etc. as input values ​​of a preset 3D generation model, and generates a preliminary 3D model related to the user based on the learning result (or artificial intelligence result / machine learning result / deep learning result). At this time, the control unit (160) generates a body-related model based on the estimated human body dimensions, the preprocessed image information, etc., and generates a face-related model based on the estimated facial dimensions, the preprocessed image information, etc., and then synthesizes the generated body-related model and the generated face-related model to generate a preliminary 3D model related to the user.

[0111] In addition, when generating a face related to the user included in the preliminary 3D model, the control unit (160) generates the shape of the user's face, including the user's facial features, by reflecting the estimated facial dimensions.

[0112] In addition, the control unit (160) extracts a plurality of preset facial landmarks (e.g., 468 facial landmarks) (or coordinate / position information corresponding to the plurality of facial landmarks) from the user's face included in the preprocessed image information, separates a plurality of facial landmarks left and right from the plurality of extracted facial landmarks centered on the preset nose center point, and generates two images that are flipped left and right through one or more of the facial landmarks separated left and right, and then performs an alignment process for the two images respectively generated so that they face the front, and performs learning using the two images aligned according to the alignment process, the estimated facial dimensions, etc. as input values ​​of the 3D generation model, and based on the learning result, generates a model related to the face implemented in a left / right asymmetrical or symmetrical state.

[0113] In addition, through an analysis function for the above-mentioned preprocessed image information, it is possible to estimate the male / female gender and create a model related to the face that reflects the male or female characteristics according to the male / female gender estimation.

[0114] For example, the first control unit performs learning using the first human body size, the estimated first facial size, the preprocessed first image information, etc. related to the estimated first user as input values ​​of the 3D generation model, and generates a first preliminary 3D model related to the first user based on the learning result (S250).

[0115] Thereafter, the control unit (160) performs other artificial intelligence-based learning (other artificial neural network / other machine learning / other deep learning) based on the extracted human body features, the extracted facial features, the generated preliminary 3D model, etc., and generates a modified 3D model related to the user based on the other learning results.

[0116] That is, the control unit (160) performs other learning (or other artificial intelligence / other machine learning / other deep learning) using the extracted human body features, the extracted facial features, the generated preliminary 3D model, etc. as input values ​​of a preset shape modification model, and generates a modified 3D model related to the user based on the other learning results (or other artificial intelligence results / other machine learning results / other deep learning results). At this time, the generated modified 3D model may be composed of points and surfaces.

[0117] For example, the first control unit performs another learning using the first human body feature related to the extracted first user, the extracted first facial feature, the generated first preliminary 3D model, etc. as input values ​​of the shape modification model, and generates a first modified 3D model related to the first user based on the other learning results (S260).

[0118] Thereafter, the control unit (160) fits (or maps / matches / links) the preset joint feature information to the generated modified 3D model, thereby ultimately generating a 3D human model related to the user. Here, the preset joint feature information includes information on the features of each joint set as default for a plurality of joints that make up the human body (e.g., including part, size, function, etc.).

[0119] That is, the control unit (160) fits the preset joint characteristic information to the generated modified 3D model, thereby generating a 3D human model related to the user that includes joint characteristic information corresponding to the generated modified 3D model. Here, the generated 3D human model may be in the form of a mesh, which is a solid model.

[0120] For example, the first control unit fits the preset joint feature information to the generated first modified three-dimensional model to ultimately generate a first three-dimensional human model in mesh form in relation to the first user.

[0121] In addition, as shown in FIGS. 4 and 5, the first control unit displays the generated first three-dimensional human model on the first display unit (140) (S270).

[0122] As described above, an embodiment of the present invention performs preprocessing on image information including a user, estimates human body dimensions and facial dimensions of the user as an object included in the preprocessed image information, extracts human body features and facial features of the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, performs other artificial intelligence-based learning based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and fits the generated modified 3D model with preset joint feature information to finally generate a 3D human model related to the user, thereby enabling the easy generation of a full-body 3D human model with a single photograph, and the 3D human model generated by reflecting human body dimensions can be immediately applied in a 1:1 virtual environment, and in the process of generating the 3D human model, there are no shadows or diffuse reflections by the image processing module, and a high-resolution texture can be generated using an upscaling technique.

[0123] Those skilled in the art will appreciate that modifications and variations of the above-described content may be made without departing from the essential characteristics of the present invention. Therefore, the embodiments disclosed herein are intended to illustrate, rather than limit, the technical spirit of the present invention, and the scope of the technical spirit of the present invention is not limited by these embodiments. The scope of protection of the present invention should be construed according to the following claims, and all technical ideas within the scope equivalent thereto should be construed as being included within the scope of the present invention.

[0124] The mode for carrying out the invention has been described together with the best mode for carrying out the invention above.

[0125] The present invention performs preprocessing on image information including a user, estimates human body dimensions and facial dimensions for the user as an object included in the preprocessed image information, extracts human body features and facial features for the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions and facial dimensions to generate a preliminary 3D model related to the user, performs other artificial intelligence-based learning based on the extracted human body features and facial features and the generated preliminary 3D model to generate a modified 3D model related to the user, and fits the generated modified 3D model with preset joint feature information to finally generate a 3D human model related to the user, thereby enabling the easy generation of a full-body 3D human model with a single photograph, and the 3D human model generated by reflecting human body dimensions can be immediately applied in a 1:1 virtual environment, and in the process of generating the 3D human model, there is no shadow or diffuse reflection by the image processing module, and a high-resolution texture can be generated using an upscaling technique, so that the present invention has industrial applicability.

Claims

1. A camera unit that acquires image information including a user; A control unit that performs preprocessing on the acquired image information, estimates human body dimensions and facial dimensions for a user as an object included in the preprocessed image information, extracts human body features and facial features for the user as an object included in the preprocessed image information, performs artificial intelligence-based learning based on the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information, generates a preliminary 3D model related to the user based on the learning results, performs other artificial intelligence-based learning based on the extracted human body features, the extracted facial features, and the generated preliminary 3D model, generates a modified 3D model related to the user based on the other learning results, and fits the generated modified 3D model and preset joint feature information to generate a 3D human model related to the user; and A device for generating a photo-based 3D human model using artificial intelligence, including a display unit for displaying the generated 3D human model.

2. In paragraph 1, The above control unit, A device for generating a 3D human model based on a photo using artificial intelligence, characterized in that it performs preprocessing including at least one of noise removal, background removal, glare removal, face image quality improvement through upscaling, and depth estimation on the acquired image information.

3. In paragraph 1, The above control unit, A device for generating a photo-based 3D human model using artificial intelligence, characterized in that it generates a body-related model based on the estimated human body dimensions and the preprocessed image information, generates a face-related model based on the estimated face dimensions and the preprocessed image information, and synthesizes the generated body-related model and the generated face-related model to generate a preliminary 3D model related to the user.

4. A step of acquiring image information including a user by a camera unit; A step of performing preprocessing on the acquired image information by the control unit; A step of estimating human body dimensions and facial dimensions for a user, which is an object included in the preprocessed image information, by the control unit; A step of extracting human body features and facial features for a user, who is an object included in the preprocessed image information, by the control unit; A step of performing artificial intelligence-based learning based on the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information by the control unit, and generating a preliminary 3D model related to the user based on the learning results; A step of performing another artificial intelligence-based learning based on the extracted human body features, the extracted facial features, and the generated preliminary three-dimensional model by the control unit, thereby generating a modified three-dimensional model related to the user based on the other learning results; A step of creating a 3D human model related to the user by fitting the generated modified 3D model and preset joint feature information by the control unit; and A method for generating a photo-based 3D human model using artificial intelligence, comprising a step of displaying the generated 3D human model on a display unit by the above control unit.

5. In paragraph 4, The step of estimating the above human body dimensions and facial dimensions is: A process of recognizing an object included in the above preprocessed image information; A process of estimating human body dimensions for the above recognized object; and A method for creating a photo-based 3D human model using artificial intelligence, characterized by including a process of estimating facial dimensions for a user face within the above-described recognized object.

6. In paragraph 4, The step of extracting the above human features and facial features is: A process of recognizing an object included in the above preprocessed image information; A process of extracting human features from the above recognized object; and A method for creating a photo-based 3D human model using artificial intelligence, characterized in that it includes a process of extracting facial features for a user's face within the recognized object.

7. In paragraph 4, The step of creating a preliminary 3D model related to the above user is: A method for generating a photo-based 3D human model using artificial intelligence, characterized in that learning is performed using the estimated human body dimensions, the estimated facial dimensions, and the preprocessed image information as input values of a preset 3D generation model, and a preliminary 3D model related to the user is generated based on the learning results.

Citation Information

Patent Citations

  • Semiconductor device and method for fabricating of the same

    KR1020250080274A

  • Hole location updating device and operating method of hole location updating device

    KR102279165B1

  • Method for generating a facial image of a virtual character through deep learning-based optimized operation, and computer readable recording medium and system therefor

    KR102393702B1

  • Device and method for generating user avatar based on attributes detected in user image and controlling natural motion of the user avatar

    KR102627035B1

  • Gate post for cargo box of truck

    KR102819935B1