System

The system addresses the challenge of online shopping by using a smartphone to create accurate virtual try-on videos, enhancing consumer satisfaction and reducing returns through precise 3D scanning and simulation.

JP2026021192APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122874
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Consumers face challenges in online shopping as they cannot try on clothes, leading to issues with size and fit, resulting in high return rates, reduced satisfaction, and increased logistics and environmental impact, while current virtual try-on systems lack accuracy.

Method used

A system that uses a smartphone camera to 3D scan a user's body, generating accurate body shape data, combines it with product data to create a virtual try-on video, simulating light reflection and clothing stretch, and delivers this experience to the user's device.

Benefits of technology

Enables highly accurate virtual try-on experiences, allowing consumers to select suitable clothing, reducing returns and associated burdens on logistics and the environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021192000001_ABST
    Figure 2026021192000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving image data from a plurality of directions photographed by a user and generating a three dimensional model of the user from the image data; means for receiving product data and generating a three dimensional model of a product from the product data; means for generating a virtual fitting video by combining the three dimensional model of the user and the three dimensional model of the product; and means for transmitting the virtual fitting video to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Traditional online shopping has the problem that consumers cannot try on clothes, making it difficult to confirm the size and fit of the clothes. As a result, many consumers return the clothes they purchased because they did not meet their expectations. This reduces consumer satisfaction and increases the logistics burden and environmental impact of returns. In addition, current virtual try-on systems have low accuracy, making it difficult to provide a realistic try-on experience. [Means for solving the problem]

[0005] The present invention provides a means for 3D scanning a user's entire body using a camera such as a smartphone and generating accurate body shape data from the obtained image data. It also includes a means for generating a 3D model of a product from received product data and combining this 3D model with the user's body shape data to generate a virtual try-on video. Specifically, a 3D model of the user is created based on image data from multiple angles, and a 3D model of the product is generated based on product dimensional information. These are then combined in real time to create a virtual try-on video with the same accuracy as if the user were actually trying on the product. The system further provides a means for transmitting this virtual try-on video to a user's device and simulating the simultaneous fitting of multiple products, as well as simulating light reflection and clothing stretch. This allows consumers to accurately select clothing that suits them when shopping online, potentially reducing product returns.

[0006] "User" refers to an individual or corporation that uses the system to virtually try on clothing.

[0007] An "imaging device" is a device that takes an image of a user's entire body and generates 3D image data, and specifically refers to a device that includes a camera function, such as a smartphone or tablet.

[0008] "Image data" refers to photographic data taken of the user from multiple angles, and is the basic information used to generate a 3D model.

[0009] A "3D model" is data that reproduces a user's body or a product in three dimensions, and includes visual shape and dimensional information.

[0010] "Product data" refers to information about the clothing to be tried on, and specifically includes photographic data, size information, and the like.

[0011] "Virtual try-on video" refers to a video that visually displays a virtual try-on experience, generated by combining a 3D model of the user and a 3D model of the product.

[0012] "Dimensional information" refers to specific numerical data such as the length and width of each part of a product, and is the basis for accurately reproducing a 3D model.

[0013] "Real-time" refers to a state in which processing is performed immediately in response to a user's operation and results are provided without delay.

[0014] "Light reflection simulation" is a technology that realistically reproduces how light reflects on a generated 3D model.

[0015] "Clothing stretch simulation" is a technology that reproduces how clothing stretches and shrinks in virtual fitting footage in a manner that is close to the actual experience of trying it on.

[0016] "Server" refers to the central processing unit that receives and analyzes user image data and product data, generates 3D models and virtual try-on videos, and transmits them to the user's terminal.

[0017] A "terminal" refers to a device such as a smartphone or tablet held by the user, which is connected to the system and allows the user to try on clothes. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] This invention relates to a system that 3D scans a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the image data obtained. The system consists of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0040] Program processing overview

[0041] 1. 3D scanning on your device

[0042] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[0043] 2. Uploading data to the server and processing it

[0044] The generated 3D scan data is uploaded from the user's device to a central server, which analyzes the received data and generates accurate body shape data for the user. This analysis process involves extracting the user's dimensions from each image, generating 3D point cloud data, and then constructing a mesh.

[0045] 3. Product data capture and generation

[0046] The server receives product data provided by the manufacturer, including photos of the garment and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the specific shape of the product from the measurements.

[0047] 4. Creation and distribution of virtual try-on videos

[0048] The server combines the generated user's body data with a 3D product model. This allows the scale of the product to be adjusted to fit the user's body data, and the product is positioned to fit. It also performs stretching and simulation of the clothing to recreate a more realistic fitting experience. The generated virtual fitting video is sent to the user's device.

[0049] Specific examples

[0050] 3D scanning on your device

[0051] The user opens the app and selects the photo mode. The app instructs the user to "stand facing forward," and the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, and then the app displays the instruction "Turn to your back," and the camera takes a rear-side photo. At this point, a 3D scan of the user is generated.

[0052] Data upload to server and processing

[0053] The 3D scan data taken by the user's device is uploaded to the server. The server analyzes the received data and generates data on the user's body shape. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[0054] Product data ingestion and generation

[0055] The server generates a 3D model of the product based on product photos and measurement information provided by the manufacturer. For example, a 3D model of a jacket is generated using a provided photo and detailed measurement information. This 3D model includes details such as sleeve length and collar shape.

[0056] Creation and distribution of virtual try-on videos

[0057] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. For example, the video of the user trying on the jacket they selected accurately shows how the jacket will fit the user's shoulders and waist. The server also simulates light reflection and the stretchiness of the clothing to recreate a more realistic fitting experience. The final virtual try-on video is sent to the user's device, where they can view it.

[0058] This invention allows consumers to choose the clothes that best suit them through highly accurate virtual try-on sessions when shopping online. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0059] The processing flow will be explained below.

[0060] Step 1:

[0061] The user launches the dedicated app on their device, selects 3D scan mode, and is prompted to take photos from three directions in order: the front, right side, and back.

[0062] Step 2:

[0063] The user stands facing forward and the device camera takes a full-body frontal photo of the user. After the frontal photo is taken, the app automatically saves the image temporarily on the device.

[0064] Step 3:

[0065] The user stands facing right, and the device camera takes a photo of the user's whole body from the right side. After the right side image is taken, the app automatically saves this image temporarily on the device.

[0066] Step 4:

[0067] The user stands with their back to the camera and takes a rear-view photo of the user's entire body. After the rear-view photo is taken, the app automatically saves the image temporarily on the device.

[0068] Step 5:

[0069] The device collects the image data of the front, side, and back taken by the device into a single file and sends a request to upload it to the server.

[0070] Step 6:

[0071] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts the user's body dimensions from each image and generates a 3D mesh model based on this.

[0072] Step 7:

[0073] The server receives product data provided by the manufacturer, including photos of the garment and measurements, and uses this data to generate a 3D model of the product.

[0074] Step 8:

[0075] The server combines the user's body data generated by the server with a 3D model of the product to generate a virtual fitting video, which simulates how the product will fit the user's body shape and measurements, and also realistically reproduces the reflection of light and the stretchiness of the clothing.

[0076] Step 9:

[0077] The server sends the generated virtual try-on video to the user's device, which plays the video and displays the try-on results to the user.

[0078] Step 10:

[0079] The user can view the virtual try-on video on the app, select different sizes or designs of clothing as needed, and request the generation of another try-on video. The server will repeat the same process as soon as it receives a new request.

[0080] Example 1

[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0082] With traditional online shopping, customers are unable to try on items, so the size and fit of the purchased item often differ from what they actually are, leading to an increase in returns, which can result in lower consumer satisfaction, increased logistics costs, and an increased environmental impact.

[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0084] In this invention, the server includes means for receiving image data captured by a user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to a user terminal, means for extracting the user's size information from each image and generating 3D point cloud data, means for constructing a mesh from the point cloud data, and means for simulating light reflection and clothing stretch. This allows consumers to experience a highly accurate virtual try-on experience, allowing them to select the most suitable clothing, and reducing logistics and environmental burdens by reducing returns.

[0085] "Image data taken by a user from multiple directions" refers to multiple sets of image data taken by a user from different directions using an imaging device.

[0086] "Means for generating a three-dimensional model" refers to software and computational processes for digitally reconstructing the three-dimensional shape of a user or product based on received image data.

[0087] "Product data" refers to information including various data such as detailed product descriptions, dimensional information, and photographs.

[0088] "Means for generating virtual try-on footage" refers to software and computational processes that combine a 3D model of the user with a 3D model of the product to visually recreate the state of the user virtually trying on the product.

[0089] "Means for transmitting the virtual try-on video to the user terminal" refers to a communication means for transmitting the generated virtual try-on video to the user's device via a network such as the Internet.

[0090] "Means for extracting dimensional information and generating 3D point cloud data" refers to software and computational processes for measuring the user's body dimensions from the received image data and generating points in 3D space based on that information.

[0091] "Means for constructing a mesh from point cloud data" refers to software and a calculation process for connecting the surface of a three-dimensional shape with triangles or polygons based on the generated point cloud data to generate a continuous mesh.

[0092] "Means for simulating light reflection and clothing stretch" refers to software and computational processes for calculating and reproducing in real time the reflection of light and the dynamic changes that occur as clothing fits the user's body in a virtual environment in which the user is trying on the clothing.

[0093] This invention relates to a system that 3D scans a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the obtained image data. This system is composed of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0094] 1. (3D scanning on device)

[0095] The user launches the dedicated smartphone app and selects the shooting mode. The app instructs the user to take three poses: front, right side, and back, and the camera takes a photo for each pose. The captured image data is temporarily stored on the device.

[0096] For example, when a user launches the app and selects the photo mode, the instruction "Please stand facing forward" is displayed. When the user stands facing forward, the camera automatically takes a photo, and then the instruction "Please turn to the right" is displayed. This process is repeated to generate 3D scan data.

[0097] 2. (Uploading data to the server and analyzing it)

[0098] The 3D scan data generated by the device is uploaded to a server. The server extracts the user's dimensional information from each image and generates 3D point cloud data. A mesh is then constructed from the point cloud data. This analysis process uses software such as Python's OpenCV and Point Cloud Library (PCL).

[0099] For example, when a user uploads image data they have taken to a server, the server analyzes the pixel information of each photo and calculates the user's height, shoulder width, waist size, etc. Then, it automatically generates 3D point cloud data and meshes based on this information.

[0100] 3. (Importing product data and generating 3D models)

[0101] The server receives product data (photos and dimensional information) provided by the manufacturer and generates a 3D model of the product based on this data. An image processing library (e.g., OpenCV) is used to use the photo as a texture and to recreate the specific shape from the dimensional information.

[0102] Example: The server receives a photo of a jacket provided by a manufacturer and detailed measurement information, and based on this information, generates a 3D model of the jacket, including details such as sleeve length and collar shape.

[0103] 4. (Generation and distribution of virtual try-on videos)

[0104] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. The scale of the product is adjusted and positioned to fit the user's body. 3D rendering software such as Blender or Unity is also used to simulate light reflection and the stretchiness of the clothing. The final virtual try-on video is sent to the user's device.

[0105] Example: A 3D model of a jacket is adjusted to fit the user's body shape data, and a virtual try-on video is generated. This video shows how the jacket fits perfectly to the user's shoulders and waist. The generated video is sent to the user's device in real time, and the user can view it on their smartphone.

[0106] Examples of prompts:

[0107] "Please generate a video of user A trying on jacket B based on the 3D scan data."

[0108] This invention allows users to perform highly accurate virtual try-on sessions and select clothing that best suits them. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0110] Step 1:

[0111] The user launches a dedicated app on their smartphone and selects a shooting mode.

[0112] Input: The user operates the app to select a shooting mode.

[0113] Specific operation: The user selects "3D scan mode" and the app launches.

[0114] Output: The app enters shooting mode.

[0115] Step 2:

[0116] The app asks the user to pose from the front, right side, and back.

[0117] Input: The user acts according to the app's instructions.

[0118] Specific behavior: The app displays the instruction "Please stand facing forward," and the user faces forward.

[0119] Output: User facing forward.

[0120] Step 3:

[0121] The device's camera takes a photo for each pose.

[0122] Input: The user strikes a pose and the device camera activates.

[0123] Specific operation: The camera automatically takes photos of the front, right side, and back in sequence.

[0124] Output: Image data from three directions.

[0125] Step 4:

[0126] The captured image data is temporarily stored on the device.

[0127] Input: Image data captured by a camera.

[0128] Specific operation: The captured image data is saved in the smartphone's temporary memory.

[0129] Output: Saved image data.

[0130] Step 5:

[0131] The device uploads the generated 3D scan data to the server.

[0132] Input: Saved image data.

[0133] Specific operation: The device sends image data to the server via an Internet connection.

[0134] Output: Image data uploaded to the server.

[0135] Step 6:

[0136] The server extracts the user's dimensional information from each image and generates 3D point cloud data.

[0137] Input: Image data uploaded to the server.

[0138] How it works: The server uses Python's OpenCV and SciPy libraries to analyze pixel information from each photo and extract measurements such as the user's height, shoulder width, and waist.

[0139] Output: Extracted dimensional information and 3D point cloud data.

[0140] Step 7:

[0141] The server constructs a mesh from the point cloud data.

[0142] Input: Extracted dimensional information and 3D point cloud data.

[0143] Specific operation: The server generates a 3D mesh model using Blender or Point Cloud Library (PCL).

[0144] Output: Reconstructed 3D mesh data.

[0145] Step 8:

[0146] The server receives product data provided by the manufacturer.

[0147] Input: Product data (photos and dimensions) provided by the manufacturer.

[0148] Specific operation: The server receives and stores the product data.

[0149] Output: Received product data.

[0150] Step 9:

[0151] The server generates a 3D model of the product based on the product data.

[0152] Input: Received product data.

[0153] Specific operation: The server uses an image processing library (e.g., OpenCV) to use the photo as a texture and reproduce the specific shape of the product from the dimensional information.

[0154] Output: A 3D model of the generated product.

[0155] Step 10:

[0156] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video.

[0157] Input: User's 3D mesh data and 3D model of the product.

[0158] How it works: The server adjusts the scale of the product and positions it to fit the user's body shape. It also uses Blender or Unity's physics engine to simulate light reflection and clothing stretching.

[0159] Output: Generated virtual try-on video.

[0160] Step 11:

[0161] The server transmits the generated virtual try-on video to the user's terminal.

[0162] Input: Generated virtual try-on footage.

[0163] Specific operation: The server sends the virtual try-on video to the user's device via the Internet.

[0164] Output: A virtual try-on video displayed on the user's device.

[0165] (Application example 1)

[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0167] With traditional online shopping, consumers are unable to actually try on products, resulting in frequent problems such as products not fitting properly or looking different after purchase. This not only increases the return rate, burdens on logistics and the environment, but also reduces consumer satisfaction. Furthermore, existing virtual try-on systems have low accuracy in user body shape data, making it difficult to reproduce an actual fit. To solve these issues, a new system that combines high-precision 3D scanning technology and realistic try-on simulations is needed.

[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0169] In this invention, the server includes: means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data; means for receiving product data and generating a 3D product model from the product data; means for combining the 3D user model and the 3D product model to generate a virtual try-on video; means for transmitting the virtual try-on video to the user device; means for instructing the user to pose when capturing images using the user device's camera or sensor; means for uploading the user's image data to the server and analyzing it to extract detailed dimensional information about the user; means for constructing a 3D model of the user based on the specific dimensional information; means for specifically reproducing the 3D product model based on the product dimensional information; and means for simulating the stretch and shadow of clothing when generating the virtual try-on video. This allows consumers to experience a highly accurate virtual try-on experience and choose the clothing that best suits them. Furthermore, reducing returns can reduce logistics and environmental burdens.

[0170] "Image data taken by a user from multiple directions" refers to a collection of image data taken by a user using a photographing device from different directions such as the front, side, and rear.

[0171] A "3D model" is three-dimensional digital data that reproduces the shape of a user or product from two-dimensional image data.

[0172] "Virtual try-on video" is a video in which a product model is matched to a 3D model of the user, making it appear as if the user is actually trying on the product.

[0173] A "user terminal" is an electronic device that is directly operated by a user, such as a smartphone, tablet, or computer.

[0174] An "imaging device" is a device used to capture images, such as a camera or sensor mounted on a user terminal.

[0175] A "server" is a computer system used for analyzing, storing, and communicating data.

[0176] "Dimensional information" is detailed information about the size of the product and dimensional data of each part of the user's body.

[0177] "Analysis processing" refers to the calculations and processing required to extract necessary information based on received data and generate a model.

[0178] "Detailed dimensional information" refers to the precise dimensions of parts of a user or product, as well as detailed measurement data.

[0179] "Stretching and shadow simulation" is a computer graphics process that recreates the realistic feeling of trying on clothes, including the stretching and shrinking of clothing and the reflection of light.

[0180] The system of the present invention generates a 3D model from image data captured by the user from multiple angles using a camera, and then combines the generated model with a 3D model of the product to provide a virtual try-on video. This system allows users to experience highly accurate virtual try-on from the comfort of their own home, enabling them to choose the clothing that best suits them.

[0181] Hardware and software used

[0182] Device:

[0183] Smartphone (with high-resolution camera and standard IMU (Inertial Measurement Unit))

[0184] server:

[0185] Data analysis server (using TensorFlow and OpenCV)

[0186] 3D rendering server (using Blender)

[0187] Communication (using REST API)

[0188] Data processing and calculation

[0189] Generate a 3D model of the user

[0190] 1. Image capture:

[0191] The user launches the app and takes photos of the front, right side, and back using the device's camera. During this process, the app displays instructions to help the user stand in the correct position and angle.

[0192] 2. Upload image data:

[0193] The captured image data is temporarily stored on the device and then uploaded to the server.

[0194] 3. Analysis process:

[0195] The server analyzes the received image data and extracts detailed dimensional information about the user. Specifically, it uses TensorFlow for image recognition and dimension extraction, OpenCV to generate point cloud data, and Blender to construct the final 3D mesh.

[0196] 3D product model generation

[0197] 1. Product data acquisition:

[0198] The server receives product photos and detailed dimensional information provided by the manufacturer.

[0199] 2. Product model generation:

[0200] Using the required dimensions and photographs, Blender is used to generate a 3D model of the product, including the product's specific shape and texture.

[0201] Virtual try-on video generation and distribution

[0202] 1. Model combination:

[0203] The server combines a 3D model of the user with a 3D model of the product to realistically simulate the fit, simulating the stretch and shadow of the clothing to recreate the feeling of trying it on.

[0204] 2. Image generation:

[0205] The virtual try-on video generated by the above process is sent to the user's device, where the user can view the video and choose the clothing that best suits them.

[0206] Specific examples

[0207] Image capture and upload

[0208] The user opens the app and takes a frontal image following the instruction "Please stand facing forward," then takes a right-side image following the instruction "Please stand facing right," and then takes a backside image following the instruction "Please stand with your back to the camera." This series of images is then uploaded to the server.

[0209] Server-side processing

[0210] The server analyzes the received image data and extracts detailed dimensional information such as the user's height, shoulder width, waist, etc. Based on the extracted dimensional information, point cloud data is generated using TensorFlow and OpenCV, and a 3D mesh is constructed using Blender.

[0211] 3D product model generation

[0212] The server receives jacket photos and measurements provided by the manufacturer and uses Blender to generate a 3D model of the jacket, including sleeve length, collar shape, texture information, and more.

[0213] Virtual try-on video generation

[0214] The generated user model is combined with the jacket model to simulate the fit, light reflection, and stretch of the clothing, generating a realistic video of the user trying on the garment. This video is then sent to the user's device.

[0215] Specific prompt examples

[0216] "Please explain the overview of the system that enables highly accurate virtual try-on based on a full-body 3D scan."

[0217] The details of this system depend on the user's environment, providing a convenient experience that can be experienced individually by each user. This will improve the online shopping experience for consumers, reduce the return rate, and ease the burden on logistics.

[0218] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0219] Step 1:

[0220] Image capture

[0221] Input: Image data from multiple angles according to user instructions

[0222] How it works: The user launches the dedicated app and takes photos of three poses using the device's camera: the front, right side, and back.

[0223] Output: Three images taken by the user (front, right, and back images)

[0224] Step 2:

[0225] Image data upload

[0226] Input: User image data captured in step 1

[0227] Operation: The device uploads the image data it has taken to the server. The app uses the REST API to send the image data to the server.

[0228] Output: User image data transferred to the server

[0229] Step 3:

[0230] Analysis processing

[0231] Input: User image data uploaded to the server

[0232] How it works: The server uses TensorFlow and OpenCV to extract the user's dimensions from each image, generates point cloud data from each image, and builds a 3D mesh using Blender.

[0233] Output: Detailed dimensional information and your 3D model

[0234] Step 4:

[0235] Get product data

[0236] Input: Product photos and dimensions provided by the manufacturer

[0237] How it works: The server receives product data, generates a 3D model based on the necessary dimensions and photos, and uses Blender to incorporate product shape and texture information into the 3D model.

[0238] Output: 3D model of the product

[0239] Step 5:

[0240] Virtual try-on video generation

[0241] Input: 3D model of the user and 3D model of the product

[0242] How it works: The server combines the user model with the product model, simulates fit, stretching, and shadows, and uses Blender to generate a virtual try-on video.

[0243] Output: Virtual try-on video data

[0244] Step 6:

[0245] Distribution of virtual try-on videos

[0246] Input: Generated virtual try-on video data

[0247] Operation: The server generates a virtual fitting video and sends it to the user's device. The user can view the video on their device and check how the clothes feel when they are tried on.

[0248] Output: Virtual try-on video that can be viewed by the user

[0249] In this way, each processing step of the system is combined to provide the user with a highly accurate virtual try-on experience.

[0250] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0251] This invention relates to a system that provides a virtual try-on experience by 3D scanning a user's entire body with a camera such as a smartphone, generating accurate body shape data from the obtained image data, and combining this with an emotion engine. The system consists of a user's device, a server that analyzes the data, a server that generates 3D product models, the emotion engine, and a server that generates virtual try-on videos and sends them to the user's device.

[0252] Program processing overview

[0253] 1. 3D scanning on your device

[0254] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[0255] 2. Uploading data to the server and processing it

[0256] The generated 3D scan data is uploaded from the user's device to a server, which analyzes the received data and generates accurate data on the user's body shape. The analysis process involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[0257] 3. Product data capture and generation

[0258] The server receives product data provided by the manufacturer, including photos of the garments and detailed measurement information, and generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the shape of the product based on the measurement information.

[0259] 4. Incorporating an Emotional Engine

[0260] While the user is trying on the clothes, the device's built-in camera and microphone analyze the user's facial expressions and tone of voice in real time, and the emotion engine recognizes the user's emotions. For example, if the system detects that the user is smiling, it will determine that the user has a positive feeling toward the clothing.

[0261] 5. Creation and distribution of virtual try-on videos

[0262] The server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, it uses the user's emotional data analyzed by the emotion engine to adjust the content of the virtual try-on video. For example, if the user's emotion is unfavorable, it can generate a scenario that suggests a different product that fits better.

[0263] 6. Video distribution and user feedback

[0264] The generated virtual try-on video is sent to the user's device for viewing. The user's try-on experience is then analyzed for emotion, and feedback is provided based on the user's emotions. For example, if the user is dissatisfied, the system automatically suggests other products.

[0265] Specific examples

[0266] 3D scanning on your device

[0267] The user launches the dedicated app and selects the shooting mode. When the app instructs them to "stand facing forward," the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, then the app displays "Turn your back," and the app takes a rear-side photo. At this point, 3D scan data of the user is generated.

[0268] Data upload to server and processing

[0269] The 3D scan data taken by the user's device is uploaded to the server, which then analyzes the data and generates accurate body shape data for the user. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[0270] Product data ingestion and generation

[0271] The server generates a 3D model of the product based on the product photos and measurements provided by the manufacturer. The server uses the provided jacket photos and measurements to generate a 3D model of the jacket, including details such as sleeve length and collar shape.

[0272] Incorporating an emotion engine

[0273] When a user tries on a jacket during the fitting experience, the device's camera analyzes the user's facial expressions and the emotion engine recognizes the user's emotions. For example, if the user is smiling, the emotion engine determines that the user likes the jacket.

[0274] Creation and distribution of virtual try-on videos

[0275] The server combines the user's body shape data with a 3D model of the product to generate a virtual try-on video. The emotion engine adjusts the video content based on the user's emotional data. For example, if the user is smiling, different colors and styles of jackets will be suggested in the video.

[0276] Video distribution and user feedback

[0277] The generated virtual try-on video is sent to the user's device, where the user reviews it. While reviewing, the emotion engine analyzes the user's facial expressions again to determine their level of satisfaction. If their satisfaction is low, the system suggests a different product. If their satisfaction is high, they can proceed with the purchase process.

[0278] This not only allows users to have a highly accurate virtual try-on experience, but also allows them to receive more personalized product recommendations based on their emotions, resulting in increased consumer satisfaction and reduced returns.

[0279] The processing flow will be explained below.

[0280] Step 1:

[0281] The user launches the dedicated app and selects 3D scan mode. The app then prompts the user to take photos from the front, right side, and back in that order.

[0282] Step 2:

[0283] The user stands facing forward, and the device camera takes a full-body frontal photograph of the user. The captured image is immediately saved temporarily on the device.

[0284] Step 3:

[0285] The user stands facing right, and the device camera takes a photo of the user's entire body from the right side. As with the front-facing photo, this image is also temporarily saved on the device.

[0286] Step 4:

[0287] The user stands with their back to the camera and takes a photo of the user's entire body from the back. The image of the user's back is also temporarily stored on the device.

[0288] Step 5:

[0289] The device compiles the 3D scan data of the front, side, and back and sends a request to upload it to the server.

[0290] Step 6:

[0291] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts dimensional information such as the user's height, shoulder width, and waist from each image and generates a 3D mesh model based on this information.

[0292] Step 7:

[0293] The server receives product data provided by the manufacturer, including photos of the garment and dimensions of each part, and generates a 3D model of the product based on this data.

[0294] Step 8:

[0295] The device's camera and microphone capture the user's facial expressions and tone of voice in real time and send the data to the emotion engine, which analyzes this data and recognizes the user's emotions.

[0296] Step 9:

[0297] The emotion engine sends the user's emotion data to the server, which then takes this emotion data into account when generating the virtual try-on video. For example, if the user expresses positive emotion, the video can include content suggesting different variations of the product.

[0298] Step 10:

[0299] The server combines the user's body data with a 3D model of the product to generate a virtual fitting video. The simulation includes the stretching and light reflection of the clothing, providing an experience as close as possible to a real try-on.

[0300] Step 11:

[0301] The generated virtual fitting video is sent to the user's device, where the user can view it. The device continues to send the user's facial expressions to the emotion engine while the video is playing, and analyzes their emotions in real time.

[0302] Step 12:

[0303] The server provides feedback based on the user's emotional data. If the user is not satisfied, the server generates content suggesting alternative products and adjusts it to attract the user's interest.

[0304] Step 13:

[0305] If the user is satisfied, the app will either proceed with the purchase or offer further suggestions, and the entire system will adjust product selection accordingly.

[0306] This not only allows users to have a highly accurate virtual try-on experience, but also allows for more personalized product recommendations based on their emotions, which in turn increases consumer satisfaction and reduces returns.

[0307] Example 2

[0308] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0309] Conventional virtual try-on systems have low accuracy in generating 3D shape models of users and have difficulty in proposing products that appropriately reflect the user's emotions. As a result, user satisfaction declines and ultimately, product returns increase. Furthermore, there is a lack of technology to analyze the user's emotional state in real time and provide feedback to the virtual try-on video.

[0310] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a three-dimensional shape model of the user from this image data, means for receiving product information and generating a three-dimensional shape model of the product from this product information, means for combining the user's three-dimensional shape model and the product's three-dimensional shape model to generate a virtual try-on video, means for analyzing the user's emotional state using an emotion analysis engine and adjusting the content of the virtual try-on video, and means for transmitting the virtual try-on video to the user terminal. This improves the accuracy of the user's three-dimensional shape model, making it possible to recommend products that reflect the user's emotions, which is expected to improve user satisfaction and reduce returned products.

[0311] "Image data taken by a user from multiple directions" refers to image data of the user taken from different directions using a photographing device owned by the user.

[0312] A "three-dimensional shape model" is a model that represents a shape in three-dimensional space and is generated from image data of a user or a product.

[0313] "Product information" refers to data provided by manufacturers and providers, including product photos and detailed dimensional information.

[0314] An "emotion analysis engine" is software or hardware that can analyze a user's facial expressions and tone of voice to recognize and determine the user's emotional state.

[0315] A "virtual try-on video" is a video that simulates the experience of trying on clothes, generated by combining a three-dimensional shape model of the user and a three-dimensional shape model of the product.

[0316] "User terminal" refers to an electronic device used by a user to display and operate information, such as a smartphone, tablet, or PC.

[0317] This system provides a virtual try-on experience by 3D scanning a user's entire body with a camera, generating accurate body shape data from the image data, and combining this with an emotion analysis engine. This system is comprised of a user's device, a server that analyzes the data, a server that generates a 3D product model, an emotion analysis engine, and a server that generates a virtual try-on video and sends it to the user's device.

[0318] The process begins when the user launches a dedicated smartphone app and takes a full-body photo. The app displays prompts such as "Please stand facing forward," "Please turn to your right," and "Please stand with your back to the camera," and takes images from each direction based on the user's current posture. After the photos are taken, the image data is temporarily stored on the device.

[0319] The user's device then uploads the stored 3D scan data to a server, which analyzes the data to generate accurate body shape data for the user. The analysis process uses computer vision techniques and machine learning algorithms to extract the user's dimensions from each image and create a 3D mesh model based on that information.

[0320] The server also receives product information from manufacturers, including product photos and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses clothing photos as textures and reproduces the product shape using measurements.

[0321] During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time. An emotion analysis engine recognizes the user's emotions and adjusts the content of the fitting video based on that information. For example, if the user is smiling, the system will determine that they have a positive attitude toward the product and reflect this in the video. If the user's emotions are not positive, the system can also generate a scenario that suggests alternative products.

[0322] The generated virtual try-on video is sent to the user's device, where the user reviews it. During the review, sentiment analysis is performed again and the user's feedback is provided. If the user's satisfaction level is low, the system automatically suggests other products. This allows the user to have a highly accurate virtual try-on experience and also receive more personalized product suggestions based on their emotions.

[0323] Example prompt sentence:

[0324] "Please stand facing forward."

[0325] "Please turn to your right."

[0326] "Please stand with your back to me."

[0327] The above is a specific embodiment based on the present invention for providing a highly accurate virtual try-on experience for the user.

[0328] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0329] Processing flow

[0330] Step 1: 3D scan on your device

[0331] 1.1 The user launches the dedicated app and selects the shooting mode.

[0332] Input: User actions

[0333] Output: Start camera, display prompt

[0334] Specific operation: The user opens the dedicated smartphone app and selects the shooting mode. The app then displays the message, "Please stand facing forward."

[0335] 1.2 The app asks the user to pose and the camera takes the photo.

[0336] Input: current user posture, device camera input

[0337] Output: Image data of the front, right side, and back of the user

[0338] Specific operation: The user follows the instructions of the app to face the front, right side, and back, and the device camera takes a photo of each. The captured image data is temporarily stored on the device.

[0339] Step 2: Upload data to the server and process it

[0340] 2.1 The device uploads data to the server.

[0341] Input: Image data stored on the device

[0342] Output: Uploaded image data

[0343] How it works: Once the user has taken a photo, the app sends the image data to a server, where it is uploaded using a secure protocol.

[0344] 2.2 The server receives and analyzes the data.

[0345] Input: Uploaded image data

[0346] Output: 3D model of the user

[0347] How it works: The server reviews the received image data and analyzes it using computer vision techniques and machine learning algorithms. It extracts the user's dimensions from each image and creates a 3D mesh model based on them.

[0348] Step 3: Import and generate product data

[0349] 3.1 The server receives product information from the manufacturer.

[0350] Input: Product photos and dimensions provided by the manufacturer

[0351] Output: Received product information

[0352] Specific operation: The server receives product information provided by the manufacturer and stores it in a database.

[0353] 3.2 The server generates a 3D model of the product.

[0354] Input: Product information

[0355] Output: 3D shape model of the product

[0356] Specific operation: The server generates a 3D model of the product based on the product information. Specifically, it uses a photo of the clothing as texture and recreates the shape of the product using measurement information.

[0357] Step 4: Incorporating the Emotion Engine

[0358] 4.1 The device collects the user's emotional data.

[0359] Input: User's facial expression data, voice data

[0360] Output: Collected emotion data

[0361] Specific operation: During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time to collect emotional data.

[0362] 4.2 The server analyzes the user's emotions using the emotion engine.

[0363] Input: Collected emotion data

[0364] Output: Parsed emotion information

[0365] How it works: The emotion analysis engine recognizes and analyzes the user's emotions based on the collected data. For example, it detects if the user is smiling and provides feedback based on that situation.

[0366] Step 5: Generate and distribute virtual try-on footage

[0367] 5.1 The server generates the virtual try-on video.

[0368] Input: 3D model of the user, 3D model of the product, analyzed emotion information

[0369] Output: Virtual try-on video

[0370] Specific operation: The server combines the 3D model of the user and the 3D model of the product to generate a virtual try-on video that reflects the user's emotional data.

[0371] 5.2 The server transmits the generated video to the user terminal.

[0372] Input: Virtual try-on video

[0373] Output: Video distribution to user devices

[0374] Specific operation: The generated virtual try-on video is sent to the user's device, allowing the user to view the video on the device.

[0375] Step 6: Video distribution and user feedback

[0376] 6.1 The device re-evaluates the user.

[0377] Input: facial expression data and voice data of the user watching the virtual try-on video

[0378] Output: Re-collected emotion data

[0379] Specific operation: While the user is viewing the virtual fitting video, the device's camera and microphone again capture the user's facial expressions and reactions.

[0380] 6.2 The server makes suggestions based on the feedback.

[0381] Input: Recollected emotion data

[0382] Output: New product proposals for users

[0383] Specific operation: The server analyzes the collected user emotion data again and generates a scenario to suggest a different product if the user's satisfaction is low. If the user's satisfaction is high, the system proceeds to the purchase procedure.

[0384] (Application example 2)

[0385] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0386] Conventional virtual try-on systems can generate try-on videos by combining a user's body data with a 3D product model, but they face challenges in providing highly personalized product recommendations that take into account the user's emotions and feedback, as well as improving the user experience. Furthermore, there is a need for an effective method to reduce product return rates while improving user satisfaction.

[0387] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0388] In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to the user terminal, means for analyzing the user's facial expression and tone of voice and extracting emotion data, means for dynamically adjusting the virtual try-on video based on the extracted emotion data, and means for making product suggestions based on the generated virtual try-on video in accordance with the user's emotions. This enables real-time analysis of user emotions and feedback and personalized product suggestions based on the analysis.

[0389] "User" refers to a person who operates a system, such as a customer or a user.

[0390] "Photography device" refers to a device used to acquire image data, such as a camera or smartphone.

[0391] "Image data" refers to data of photographs or videos captured by a photographing device.

[0392] "3D model" refers to a geometric model of digital data expressed in three dimensions.

[0393] "Product Data" refers to data containing product details, such as product dimensions and photos.

[0394] "Virtual try-on video" refers to a try-on simulation video generated by combining a 3D model of the user and a 3D model of the product.

[0395] "User terminal" refers to an information terminal operated by a user, such as a smartphone or tablet.

[0396] "Facial expression" refers to information that indicates emotions expressed through the movement of facial muscles.

[0397] "Tone of voice" refers to audio information that indicates the pitch, strength, and emotional nuances of a speaking voice.

[0398] "Emotion data" refers to data that indicates the user's emotions analyzed from facial expressions, tone of voice, etc.

[0399] "Dynamic adjustment" refers to the process of changing the video content in response to changing conditions in real time.

[0400] "Personalized product suggestions" refer to product suggestions recommended based on the individual preferences and feelings of each user.

[0401] The system for realizing this invention comprises a user terminal, a server for analyzing data, and a server for managing product data. It also includes hardware and software that are equipped with an emotion engine and can analyze user emotions. A specific embodiment of the system will be described below.

[0402] The process begins when the user launches a dedicated smartphone application and takes a full-body photo. The application displays instructions to the user, such as "Please face forward," "Please face your right side," and "Please turn your back," and prompts the user to take photos in each pose. This process generates 3D scan data of the user.

[0403] The captured 3D scan data is uploaded from the user's device to an analysis server, which then uses the data to generate accurate body shape data for the user. This analysis involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[0404] Next, the product data management server receives the product data, which includes product photos and detailed measurement information. The product data server uses this data to generate a 3D model of the product. In this process, the photos are used as textures, and the shape of the product is reproduced based on the measurement information. The 3D model is structured so that the movement of each part can be flexibly reproduced.

[0405] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice in real time. During the user's fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice, and the emotion engine extracts the user's emotional data. For example, if the user is smiling, the system determines that the user has a positive feeling toward the clothing.

[0406] The analysis server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, the content of the virtual try-on video can be dynamically adjusted based on the user's emotional data analyzed by the emotion engine. For example, if the user's emotion is unfavorable, the system can generate a scenario that suggests a different product that fits better.

[0407] The generated virtual try-on video is sent to the user's device, where the user can review the results. During this review process, the device's camera again analyzes the user's facial expressions and extracts emotional data. If the user is not satisfied, the system suggests other products, but if satisfied, the user can proceed with the purchase.

[0408] Specific examples

[0409] For example, a user can launch a dedicated smartphone app and follow the instructions to take a photo facing forward, then facing right, and finally facing away from the user, generating 3D scan data. The captured data is then uploaded to a server, where it is analyzed and detailed body shape data is generated.

[0410] While the user is trying on a jacket, the device's camera analyzes the user's facial expressions, and the emotion engine dynamically suggests products based on prompts such as, "If the user is smiling, suggest different color and style variations of that jacket." For example, when a user tries on a jacket, the emotion engine can suggest other products or customization options based on prompts such as, "If the user is not satisfied, suggest other products or customization options."

[0411] This provides users with a highly accurate and personalized virtual try-on experience, reducing return rates while increasing consumer satisfaction.

[0412] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0413] Step 1:

[0414] The user launches a dedicated application on their smartphone and takes a full-body photo. The user follows the application's instructions to pose and take photos from the front, right side, and back. Image data from multiple angles of the user is obtained as input. As output, this data is generated as 3D scan data. In this step, each image is captured and temporarily saved on the device.

[0415] Step 2:

[0416] The user terminal uploads the generated 3D scan data to the server. As input, the 3D scan data stored on the terminal is used. As output, the server stores the received image data. In this step, data is transferred over the network.

[0417] Step 3:

[0418] The server analyzes the received 3D scan data and generates accurate body shape data for the user. The 3D scan data is used as input. As output, a 3D mesh model of the user's body shape is generated. In this step, image processing techniques are used to extract dimensional information and create a 3D model based on that data.

[0419] Step 4:

[0420] The product data management server receives the product data. As input, it receives a product photo and detailed measurement information. As output, it generates a 3D model of the product. In this step, it uses the photo as a texture and processes the product shape based on the measurement information.

[0421] Step 5:

[0422] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice. When the user tries on the clothes, the device's camera and microphone capture the user's facial expressions and voice to extract emotion data. Real-time video and audio data are used as input. The output is the user's emotion data. In this step, the emotion analysis algorithm is executed.

[0423] Step 6:

[0424] The analysis server combines the generated user's body shape data with the 3D model of the product to generate a virtual try-on video. The 3D mesh model of the user and the 3D model of the product are used as input. The output is a virtual try-on video. In this step, 3D rendering technology is used to create a visually realistic try-on video.

[0425] Step 7:

[0426] The content of the virtual try-on video is dynamically adjusted based on the user's emotional data analyzed by the emotion engine. The user's emotional data is used as input. The output is an adjusted try-on video based on the user's emotions. In this step, the system changes the video content taking the user's reactions into account.

[0427] Step 8:

[0428] The generated virtual try-on video is sent to the user's device. The adjusted virtual try-on video is used as input. The video is played on the user's device as output. In this step, the data transfer and video playback processes are executed.

[0429] Step 9:

[0430] When the user checks the fitting results, the device's camera again analyzes the user's facial expressions and extracts emotional data. The input is a video of the user watching the fitting video. The output is again emotional data. In this step, emotional analysis is performed to evaluate the user's satisfaction.

[0431] Step 10:

[0432] If the satisfaction level is low, the system will suggest another product. If the satisfaction level is high, the user can proceed to the purchase process. As input, the user's emotional data is used. As output, a product suggestion and purchase process scenario is generated. In this step, the personalized product suggestion and purchase process are executed.

[0433] Through these steps, users can enjoy a highly accurate and personalized virtual try-on experience, resulting in increased satisfaction and reduced return rates.

[0434] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0435] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0436] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0437] [Second embodiment]

[0438] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0439] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0440] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0441] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0442] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0443] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0444] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0445] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0446] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0447] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0448] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0449] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0450] This invention relates to a system that 3D scans a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the image data obtained. The system consists of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0451] Program processing overview

[0452] 1. 3D scanning on your device

[0453] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[0454] 2. Uploading data to the server and processing it

[0455] The generated 3D scan data is uploaded from the user's device to a central server, which analyzes the received data and generates accurate body shape data for the user. This analysis process involves extracting the user's dimensions from each image, generating 3D point cloud data, and then constructing a mesh.

[0456] 3. Product data capture and generation

[0457] The server receives product data provided by the manufacturer, including photos of the garment and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the specific shape of the product from the measurements.

[0458] 4. Creation and distribution of virtual try-on videos

[0459] The server combines the generated user's body data with a 3D product model. This allows the scale of the product to be adjusted to fit the user's body data, and the product is positioned to fit. It also performs stretching and simulation of the clothing to recreate a more realistic fitting experience. The generated virtual fitting video is sent to the user's device.

[0460] Specific examples

[0461] 3D scanning on your device

[0462] The user opens the app and selects the photo mode. The app instructs the user to "stand facing forward," and the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, and then the app displays the instruction "Turn to your back," and the camera takes a rear-side photo. At this point, a 3D scan of the user is generated.

[0463] Data upload to server and processing

[0464] The 3D scan data taken by the user's device is uploaded to the server. The server analyzes the received data and generates data on the user's body shape. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[0465] Product data ingestion and generation

[0466] The server generates a 3D model of the product based on product photos and measurement information provided by the manufacturer. For example, a 3D model of a jacket is generated using a provided photo and detailed measurement information. This 3D model includes details such as sleeve length and collar shape.

[0467] Creation and distribution of virtual try-on videos

[0468] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. For example, the video of the user trying on the jacket they selected accurately shows how the jacket will fit the user's shoulders and waist. The server also simulates light reflection and the stretchiness of the clothing to recreate a more realistic fitting experience. The final virtual try-on video is sent to the user's device, where they can view it.

[0469] This invention allows consumers to choose the clothes that best suit them through highly accurate virtual try-on sessions when shopping online. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0470] The processing flow will be explained below.

[0471] Step 1:

[0472] The user launches the dedicated app on their device, selects 3D scan mode, and is prompted to take photos from three directions in order: the front, right side, and back.

[0473] Step 2:

[0474] The user stands facing forward and the device camera takes a full-body frontal photo of the user. After the frontal photo is taken, the app automatically saves the image temporarily on the device.

[0475] Step 3:

[0476] The user stands facing right, and the device camera takes a photo of the user's whole body from the right side. After the right side image is taken, the app automatically saves this image temporarily on the device.

[0477] Step 4:

[0478] The user stands with their back to the camera and takes a rear-view photo of the user's entire body. After the rear-view photo is taken, the app automatically saves the image temporarily on the device.

[0479] Step 5:

[0480] The device collects the image data of the front, side, and back taken by the device into a single file and sends a request to upload it to the server.

[0481] Step 6:

[0482] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts the user's body dimensions from each image and generates a 3D mesh model based on this.

[0483] Step 7:

[0484] The server receives product data provided by the manufacturer, including photos of the garment and measurements, and uses this data to generate a 3D model of the product.

[0485] Step 8:

[0486] The server combines the user's body data generated by the server with a 3D model of the product to generate a virtual fitting video, which simulates how the product will fit the user's body shape and measurements, and also realistically reproduces the reflection of light and the stretchiness of the clothing.

[0487] Step 9:

[0488] The server sends the generated virtual try-on video to the user's device, which plays the video and displays the try-on results to the user.

[0489] Step 10:

[0490] The user can view the virtual try-on video on the app, select different sizes or designs of clothing as needed, and request the generation of another try-on video. The server will repeat the same process as soon as it receives a new request.

[0491] Example 1

[0492] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0493] With traditional online shopping, customers are unable to try on items, so the size and fit of the purchased item often differ from what they actually are, leading to an increase in returns, which can result in lower consumer satisfaction, increased logistics costs, and an increased environmental impact.

[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0495] In this invention, the server includes means for receiving image data captured by a user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to a user terminal, means for extracting the user's size information from each image and generating 3D point cloud data, means for constructing a mesh from the point cloud data, and means for simulating light reflection and clothing stretch. This allows consumers to experience a highly accurate virtual try-on experience, allowing them to select the most suitable clothing, and also reduces logistics and environmental burdens by reducing returns.

[0496] "Image data taken by a user from multiple directions" refers to multiple sets of image data taken by a user from different directions using an imaging device.

[0497] "Means for generating a three-dimensional model" refers to software and computational processes for digitally reconstructing the three-dimensional shape of a user or product based on received image data.

[0498] "Product data" refers to information including various data such as detailed product descriptions, dimensional information, and photographs.

[0499] "Means for generating virtual try-on footage" refers to software and computational processes that combine a 3D model of the user with a 3D model of the product to visually recreate the state of the user virtually trying on the product.

[0500] "Means for transmitting the virtual try-on video to the user terminal" refers to a communication means for transmitting the generated virtual try-on video to the user's device via a network such as the Internet.

[0501] "Means for extracting dimensional information and generating 3D point cloud data" refers to software and computational processes for measuring the user's body dimensions from the received image data and generating points in 3D space based on that information.

[0502] "Means for constructing a mesh from point cloud data" refers to software and a calculation process for connecting the surface of a three-dimensional shape with triangles or polygons based on the generated point cloud data to generate a continuous mesh.

[0503] "Means for simulating light reflection and clothing stretch" refers to software and computational processes for calculating and reproducing in real time the reflection of light and the dynamic changes that occur as clothing fits the user's body in a virtual environment in which the user is trying on the clothing.

[0504] This invention relates to a system that performs a 3D scan of a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the obtained image data. This system is composed of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0505] 1. (3D scanning on device)

[0506] The user launches the dedicated smartphone app and selects the shooting mode. The app instructs the user to take three poses: front, right side, and back, and the camera takes a photo for each pose. The captured image data is temporarily stored on the device.

[0507] For example, when a user launches the app and selects the photo mode, the instruction "Please stand facing forward" is displayed. When the user stands facing forward, the camera automatically takes a photo, and then the instruction "Please turn to the right" is displayed. This process is repeated to generate 3D scan data.

[0508] 2. (Uploading data to the server and analyzing it)

[0509] The 3D scan data generated by the device is uploaded to a server. The server extracts the user's dimensional information from each image and generates 3D point cloud data. A mesh is then constructed from the point cloud data. This analysis process uses software such as Python's OpenCV and Point Cloud Library (PCL).

[0510] For example, when a user uploads image data they have taken to a server, the server analyzes the pixel information of each photo and calculates the user's height, shoulder width, waist size, etc. Then, it automatically generates 3D point cloud data and meshes based on this information.

[0511] 3. (Importing product data and generating 3D models)

[0512] The server receives product data (photos and dimensional information) provided by the manufacturer and generates a 3D model of the product based on this data. An image processing library (e.g., OpenCV) is used to use the photo as a texture and to recreate the specific shape from the dimensional information.

[0513] Example: The server receives a photo of a jacket provided by a manufacturer and detailed measurement information, and based on this information, generates a 3D model of the jacket, including details such as sleeve length and collar shape.

[0514] 4. (Generation and distribution of virtual try-on videos)

[0515] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. The scale of the product is adjusted and positioned to fit the user's body. 3D rendering software such as Blender or Unity is also used to simulate light reflection and the stretchiness of the clothing. The final virtual try-on video is sent to the user's device.

[0516] Example: A 3D model of a jacket is adjusted to fit the user's body shape data, and a virtual try-on video is generated. This video shows how the jacket fits perfectly to the user's shoulders and waist. The generated video is sent to the user's device in real time, and the user can view it on their smartphone.

[0517] Examples of prompts:

[0518] "Please generate a video of user A trying on jacket B based on the 3D scan data."

[0519] This invention allows users to perform highly accurate virtual try-on sessions and select clothing that best suits them. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0520] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0521] Step 1:

[0522] The user launches a dedicated app on their smartphone and selects a shooting mode.

[0523] Input: The user operates the app to select a shooting mode.

[0524] Specific operation: The user selects "3D scan mode" and the app launches.

[0525] Output: The app enters shooting mode.

[0526] Step 2:

[0527] The app asks the user to pose from the front, right side, and back.

[0528] Input: The user acts according to the app's instructions.

[0529] Specific behavior: The app displays the instruction "Please stand facing forward," and the user faces forward.

[0530] Output: User facing forward.

[0531] Step 3:

[0532] The device's camera takes a photo for each pose.

[0533] Input: The user strikes a pose and the device camera activates.

[0534] Specific operation: The camera automatically takes photos of the front, right side, and back in sequence.

[0535] Output: Image data from three directions.

[0536] Step 4:

[0537] The captured image data is temporarily stored on the device.

[0538] Input: Image data captured by a camera.

[0539] Specific operation: The captured image data is saved in the smartphone's temporary memory.

[0540] Output: Saved image data.

[0541] Step 5:

[0542] The device uploads the generated 3D scan data to the server.

[0543] Input: Saved image data.

[0544] Specific operation: The device sends image data to the server via an Internet connection.

[0545] Output: Image data uploaded to the server.

[0546] Step 6:

[0547] The server extracts the user's dimensional information from each image and generates 3D point cloud data.

[0548] Input: Image data uploaded to the server.

[0549] How it works: The server uses Python's OpenCV and SciPy libraries to analyze pixel information from each photo and extract measurements such as the user's height, shoulder width, and waist.

[0550] Output: Extracted dimensional information and 3D point cloud data.

[0551] Step 7:

[0552] The server constructs a mesh from the point cloud data.

[0553] Input: Extracted dimensional information and 3D point cloud data.

[0554] Specific operation: The server generates a 3D mesh model using Blender or Point Cloud Library (PCL).

[0555] Output: Reconstructed 3D mesh data.

[0556] Step 8:

[0557] The server receives product data provided by the manufacturer.

[0558] Input: Product data (photos and dimensions) provided by the manufacturer.

[0559] Specific operation: The server receives and stores the product data.

[0560] Output: Received product data.

[0561] Step 9:

[0562] The server generates a 3D model of the product based on the product data.

[0563] Input: Received product data.

[0564] Specific operation: The server uses an image processing library (e.g., OpenCV) to use the photo as a texture and reproduce the specific shape of the product from the dimensional information.

[0565] Output: A 3D model of the generated product.

[0566] Step 10:

[0567] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video.

[0568] Input: User's 3D mesh data and 3D model of the product.

[0569] How it works: The server adjusts the scale of the product and positions it to fit the user's body shape. It also uses Blender or Unity's physics engine to simulate light reflection and clothing stretching.

[0570] Output: Generated virtual try-on video.

[0571] Step 11:

[0572] The server transmits the generated virtual try-on video to the user's terminal.

[0573] Input: Generated virtual try-on footage.

[0574] Specific operation: The server sends the virtual try-on video to the user's device via the Internet.

[0575] Output: A virtual try-on video displayed on the user's device.

[0576] (Application example 1)

[0577] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0578] With traditional online shopping, consumers are unable to actually try on products, resulting in frequent problems such as products not fitting properly or looking different after purchase. This not only increases the return rate, burdens on logistics and the environment, but also reduces consumer satisfaction. Furthermore, existing virtual try-on systems have low accuracy in user body shape data, making it difficult to reproduce an actual fit. To solve these issues, a new system that combines high-precision 3D scanning technology and realistic try-on simulations is needed.

[0579] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0580] In this invention, the server includes: means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data; means for receiving product data and generating a 3D product model from the product data; means for combining the 3D user model and the 3D product model to generate a virtual try-on video; means for transmitting the virtual try-on video to the user device; means for instructing the user to pose when capturing images using the user device's camera or sensor; means for uploading the user's image data to the server and analyzing it to extract detailed dimensional information about the user; means for constructing a 3D model of the user based on the specific dimensional information; means for specifically reproducing the 3D product model based on the product dimensional information; and means for simulating the stretch and shadow of clothing when generating the virtual try-on video. This allows consumers to experience a highly accurate virtual try-on experience and choose the clothing that best suits them. Furthermore, reducing returns can reduce logistics and environmental burdens.

[0581] "Image data taken by a user from multiple directions" refers to a collection of image data taken by a user using a photographing device from different directions such as the front, side, and rear.

[0582] A "3D model" is three-dimensional digital data that reproduces the shape of a user or product from two-dimensional image data.

[0583] "Virtual try-on video" is a video in which a product model is matched to a 3D model of the user, making it appear as if the user is actually trying on the product.

[0584] A "user terminal" is an electronic device that is directly operated by a user, such as a smartphone, tablet, or computer.

[0585] An "imaging device" is a device used to capture images, such as a camera or sensor mounted on a user terminal.

[0586] A "server" is a computer system used for analyzing, storing, and communicating data.

[0587] "Dimensional information" is detailed information about the size of the product and dimensional data of each part of the user's body.

[0588] "Analysis processing" refers to the calculations and processing required to extract necessary information based on received data and generate a model.

[0589] "Detailed dimensional information" refers to the precise dimensions of parts of a user or product, as well as detailed measurement data.

[0590] "Stretching and shadow simulation" is a computer graphics process that recreates the realistic feeling of trying on clothes, including the stretching and shrinking of clothing and the reflection of light.

[0591] The system of the present invention generates a 3D model from image data captured by the user from multiple angles using a camera, and then combines the generated model with a 3D model of the product to provide a virtual try-on video. This system allows users to experience highly accurate virtual try-on from the comfort of their own home, enabling them to choose the clothing that best suits them.

[0592] Hardware and software used

[0593] Device:

[0594] Smartphone (with high-resolution camera and standard IMU (Inertial Measurement Unit))

[0595] server:

[0596] Data analysis server (using TensorFlow and OpenCV)

[0597] 3D rendering server (using Blender)

[0598] Communication (using REST API)

[0599] Data processing and calculation

[0600] Generate a 3D model of the user

[0601] 1. Image capture:

[0602] The user launches the app and takes photos of the front, right side, and back using the device's camera. During this process, the app displays instructions to help the user stand in the correct position and angle.

[0603] 2. Upload image data:

[0604] The captured image data is temporarily stored on the device and then uploaded to the server.

[0605] 3. Analysis process:

[0606] The server analyzes the received image data and extracts detailed dimensional information about the user. Specifically, it uses TensorFlow for image recognition and dimension extraction, OpenCV to generate point cloud data, and Blender to construct the final 3D mesh.

[0607] 3D product model generation

[0608] 1. Product data acquisition:

[0609] The server receives product photos and detailed dimensional information provided by the manufacturer.

[0610] 2. Product model generation:

[0611] Using the required dimensions and photographs, Blender is used to generate a 3D model of the product, including the product's specific shape and texture.

[0612] Virtual try-on video generation and distribution

[0613] 1. Model combination:

[0614] The server combines a 3D model of the user with a 3D model of the product to realistically simulate the fit, simulating the stretch and shadow of the clothing to recreate the feeling of trying it on.

[0615] 2. Image generation:

[0616] The virtual try-on video generated by the above process is sent to the user's device, where the user can view the video and choose the clothing that best suits them.

[0617] Specific examples

[0618] Image capture and upload

[0619] The user opens the app and takes a frontal image following the instruction "Please stand facing forward," then takes a right-side image following the instruction "Please stand facing right," and then takes a backside image following the instruction "Please stand with your back to the camera." This series of images is then uploaded to the server.

[0620] Server-side processing

[0621] The server analyzes the received image data and extracts detailed dimensional information such as the user's height, shoulder width, waist, etc. Based on the extracted dimensional information, point cloud data is generated using TensorFlow and OpenCV, and a 3D mesh is constructed using Blender.

[0622] 3D product model generation

[0623] The server receives jacket photos and measurements provided by the manufacturer and uses Blender to generate a 3D model of the jacket, including sleeve length, collar shape, texture information, and more.

[0624] Virtual try-on video generation

[0625] The generated user model is combined with the jacket model to simulate the fit, light reflection, and stretch of the clothing, generating a realistic video of the user trying on the garment. This video is then sent to the user's device.

[0626] Specific prompt examples

[0627] "Please explain the overview of the system that enables highly accurate virtual try-on based on a full-body 3D scan."

[0628] The details of this system depend on the user's environment, providing a convenient experience that can be experienced individually by each user. This will improve the online shopping experience for consumers, reduce the return rate, and ease the burden on logistics.

[0629] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0630] Step 1:

[0631] Image capture

[0632] Input: Image data from multiple angles according to user instructions

[0633] How it works: The user launches the dedicated app and takes photos of three poses using the device's camera: the front, right side, and back.

[0634] Output: Three images taken by the user (front, right, and back images)

[0635] Step 2:

[0636] Image data upload

[0637] Input: User image data captured in step 1

[0638] Operation: The device uploads the image data it has taken to the server. The app uses the REST API to send the image data to the server.

[0639] Output: User image data transferred to the server

[0640] Step 3:

[0641] Analysis processing

[0642] Input: User image data uploaded to the server

[0643] How it works: The server uses TensorFlow and OpenCV to extract the user's dimensions from each image, generates point cloud data from each image, and builds a 3D mesh using Blender.

[0644] Output: Detailed dimensional information and your 3D model

[0645] Step 4:

[0646] Get product data

[0647] Input: Product photos and dimensions provided by the manufacturer

[0648] How it works: The server receives product data, generates a 3D model based on the necessary dimensions and photos, and uses Blender to incorporate product shape and texture information into the 3D model.

[0649] Output: 3D model of the product

[0650] Step 5:

[0651] Virtual try-on video generation

[0652] Input: 3D model of the user and 3D model of the product

[0653] How it works: The server combines the user model with the product model, simulates fit, stretching, and shadows, and uses Blender to generate a virtual try-on video.

[0654] Output: Virtual try-on video data

[0655] Step 6:

[0656] Distribution of virtual try-on videos

[0657] Input: Generated virtual try-on video data

[0658] Operation: The server generates a virtual fitting video and sends it to the user's device. The user can view the video on their device and check how the clothes feel when they are tried on.

[0659] Output: Virtual try-on video that can be viewed by the user

[0660] In this way, each processing step of the system is combined to provide the user with a highly accurate virtual try-on experience.

[0661] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0662] This invention relates to a system that provides a virtual try-on experience by 3D scanning a user's entire body with a camera such as a smartphone, generating accurate body shape data from the obtained image data, and combining this with an emotion engine. The system consists of a user's device, a server that analyzes the data, a server that generates 3D product models, the emotion engine, and a server that generates virtual try-on videos and sends them to the user's device.

[0663] Program processing overview

[0664] 1. 3D scanning on your device

[0665] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[0666] 2. Uploading data to the server and processing it

[0667] The generated 3D scan data is uploaded from the user's device to a server, which analyzes the received data and generates accurate data on the user's body shape. The analysis process involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[0668] 3. Product data capture and generation

[0669] The server receives product data provided by the manufacturer, including photos of the garments and detailed measurement information, and generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the shape of the product based on the measurement information.

[0670] 4. Incorporating an Emotional Engine

[0671] While the user is trying on the clothes, the device's built-in camera and microphone analyze the user's facial expressions and tone of voice in real time, and the emotion engine recognizes the user's emotions. For example, if the system detects that the user is smiling, it will determine that the user has a positive feeling toward the clothing.

[0672] 5. Creation and distribution of virtual try-on videos

[0673] The server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, it uses the user's emotional data analyzed by the emotion engine to adjust the content of the virtual try-on video. For example, if the user's emotion is unfavorable, it can generate a scenario that suggests a different product that fits better.

[0674] 6. Video distribution and user feedback

[0675] The generated virtual try-on video is sent to the user's device for viewing. The user's try-on experience is then analyzed for emotion, and feedback is provided based on the user's emotions. For example, if the user is dissatisfied, the system automatically suggests other products.

[0676] Specific examples

[0677] 3D scanning on your device

[0678] The user launches the dedicated app and selects the shooting mode. When the app instructs them to "stand facing forward," the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, then the app displays "Turn your back," and the app takes a rear-side photo. At this point, 3D scan data of the user is generated.

[0679] Data upload to server and processing

[0680] The 3D scan data taken by the user's device is uploaded to the server, which then analyzes the data and generates accurate body shape data for the user. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[0681] Product data ingestion and generation

[0682] The server generates a 3D model of the product based on the product photos and measurements provided by the manufacturer. The server uses the provided jacket photos and measurements to generate a 3D model of the jacket, including details such as sleeve length and collar shape.

[0683] Incorporating an emotion engine

[0684] When a user tries on a jacket during the fitting experience, the device's camera analyzes the user's facial expressions and the emotion engine recognizes the user's emotions. For example, if the user is smiling, the emotion engine determines that the user likes the jacket.

[0685] Creation and distribution of virtual try-on videos

[0686] The server combines the user's body shape data with a 3D model of the product to generate a virtual try-on video. The emotion engine adjusts the video content based on the user's emotional data. For example, if the user is smiling, different colors and styles of jackets will be suggested in the video.

[0687] Video distribution and user feedback

[0688] The generated virtual try-on video is sent to the user's device, where the user reviews it. While reviewing, the emotion engine analyzes the user's facial expressions again to determine their level of satisfaction. If their satisfaction is low, the system suggests a different product. If their satisfaction is high, they can proceed with the purchase process.

[0689] This not only allows users to have a highly accurate virtual try-on experience, but also allows them to receive more personalized product recommendations based on their emotions, resulting in increased consumer satisfaction and reduced returns.

[0690] The processing flow will be explained below.

[0691] Step 1:

[0692] The user launches the dedicated app and selects 3D scan mode. The app then prompts the user to take photos from the front, right side, and back in that order.

[0693] Step 2:

[0694] The user stands facing forward, and the device camera takes a full-body frontal photograph of the user. The captured image is immediately saved temporarily on the device.

[0695] Step 3:

[0696] The user stands facing right, and the device camera takes a photo of the user's entire body from the right side. As with the front-facing photo, this image is also temporarily saved on the device.

[0697] Step 4:

[0698] The user stands with their back to the camera and takes a photo of the user's entire body from the back. The image of the user's back is also temporarily stored on the device.

[0699] Step 5:

[0700] The device compiles the 3D scan data of the front, side, and back and sends a request to upload it to the server.

[0701] Step 6:

[0702] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts dimensional information such as the user's height, shoulder width, and waist from each image and generates a 3D mesh model based on this information.

[0703] Step 7:

[0704] The server receives product data provided by the manufacturer, including photos of the garment and dimensions of each part, and generates a 3D model of the product based on this data.

[0705] Step 8:

[0706] The device's camera and microphone capture the user's facial expressions and tone of voice in real time and send the data to the emotion engine, which analyzes this data and recognizes the user's emotions.

[0707] Step 9:

[0708] The emotion engine sends the user's emotion data to the server, which then takes this emotion data into account when generating the virtual try-on video. For example, if the user expresses positive emotion, the video can include content suggesting different variations of the product.

[0709] Step 10:

[0710] The server combines the user's body data with a 3D model of the product to generate a virtual fitting video. The simulation includes the stretching and light reflection of the clothing, providing an experience as close as possible to a real try-on.

[0711] Step 11:

[0712] The generated virtual fitting video is sent to the user's device, where the user can view it. The device continues to send the user's facial expressions to the emotion engine while the video is playing, and analyzes their emotions in real time.

[0713] Step 12:

[0714] The server provides feedback based on the user's emotional data. If the user is not satisfied, the server generates content suggesting alternative products and adjusts it to attract the user's interest.

[0715] Step 13:

[0716] If the user is satisfied, the app will either proceed with the purchase or offer further suggestions, and the entire system will adjust product selection accordingly.

[0717] This not only allows users to have a highly accurate virtual try-on experience, but also allows for more personalized product recommendations based on their emotions, which in turn increases consumer satisfaction and reduces returns.

[0718] Example 2

[0719] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0720] Conventional virtual try-on systems have low accuracy in generating 3D shape models of users and have difficulty in proposing products that appropriately reflect the user's emotions. As a result, user satisfaction declines and ultimately, product returns increase. Furthermore, there is a lack of technology to analyze the user's emotional state in real time and provide feedback to the virtual try-on video.

[0721] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a three-dimensional shape model of the user from this image data, means for receiving product information and generating a three-dimensional shape model of the product from this product information, means for combining the user's three-dimensional shape model and the product's three-dimensional shape model to generate a virtual try-on video, means for analyzing the user's emotional state using an emotion analysis engine and adjusting the content of the virtual try-on video, and means for transmitting the virtual try-on video to the user terminal. This improves the accuracy of the user's three-dimensional shape model, making it possible to recommend products that reflect the user's emotions, which is expected to improve user satisfaction and reduce returned products.

[0722] "Image data taken by a user from multiple directions" refers to image data of the user taken from different directions using a photographing device owned by the user.

[0723] A "three-dimensional shape model" is a model that represents a shape in three-dimensional space and is generated from image data of a user or a product.

[0724] "Product information" refers to data provided by manufacturers and providers, including product photos and detailed dimensional information.

[0725] An "emotion analysis engine" is software or hardware that can analyze a user's facial expressions and tone of voice to recognize and determine the user's emotional state.

[0726] A "virtual try-on video" is a video that simulates the experience of trying on clothes, generated by combining a three-dimensional shape model of the user and a three-dimensional shape model of the product.

[0727] "User terminal" refers to an electronic device used by a user to display and operate information, such as a smartphone, tablet, or PC.

[0728] This system provides a virtual try-on experience by 3D scanning a user's entire body with a camera, generating accurate body shape data from the image data, and combining this with an emotion analysis engine. This system is comprised of a user's device, a server that analyzes the data, a server that generates a 3D product model, an emotion analysis engine, and a server that generates a virtual try-on video and sends it to the user's device.

[0729] The process begins when the user launches a dedicated smartphone app and takes a full-body photo. The app displays prompts such as "Please stand facing forward," "Please turn to your right," and "Please stand with your back to the camera," and takes images from each direction based on the user's current posture. After the photos are taken, the image data is temporarily stored on the device.

[0730] The user's device then uploads the stored 3D scan data to a server, which analyzes the data to generate accurate body shape data for the user. The analysis process uses computer vision techniques and machine learning algorithms to extract the user's dimensions from each image and create a 3D mesh model based on that information.

[0731] The server also receives product information from manufacturers, including product photos and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses clothing photos as textures and reproduces the product shape using measurements.

[0732] During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time. An emotion analysis engine recognizes the user's emotions and adjusts the content of the fitting video based on that information. For example, if the user is smiling, the system will determine that they have a positive attitude toward the product and reflect this in the video. If the user's emotions are not positive, the system can also generate a scenario that suggests alternative products.

[0733] The generated virtual try-on video is sent to the user's device, where the user reviews it. During the review, sentiment analysis is performed again and the user's feedback is provided. If the user's satisfaction level is low, the system automatically suggests other products. This allows the user to have a highly accurate virtual try-on experience and also receive more personalized product suggestions based on their emotions.

[0734] Example prompt sentence:

[0735] "Please stand facing forward."

[0736] "Please turn to your right."

[0737] "Please stand with your back to me."

[0738] The above is a specific embodiment based on the present invention for providing a highly accurate virtual try-on experience for the user.

[0739] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0740] Processing flow

[0741] Step 1: 3D scan on your device

[0742] 1.1 The user launches the dedicated app and selects the shooting mode.

[0743] Input: User actions

[0744] Output: Start camera, display prompt

[0745] Specific operation: The user opens the dedicated smartphone app and selects the shooting mode. The app then displays the message, "Please stand facing forward."

[0746] 1.2 The app asks the user to pose and the camera takes the photo.

[0747] Input: current user posture, device camera input

[0748] Output: Image data of the front, right side, and back of the user

[0749] Specific operation: The user follows the instructions of the app to face the front, right side, and back, and the device camera takes a photo of each. The captured image data is temporarily stored on the device.

[0750] Step 2: Upload data to the server and process it

[0751] 2.1 The device uploads data to the server.

[0752] Input: Image data stored on the device

[0753] Output: Uploaded image data

[0754] How it works: Once the user has taken a photo, the app sends the image data to a server, where it is uploaded using a secure protocol.

[0755] 2.2 The server receives and analyzes the data.

[0756] Input: Uploaded image data

[0757] Output: 3D model of the user

[0758] How it works: The server reviews the received image data and analyzes it using computer vision techniques and machine learning algorithms. It extracts the user's dimensions from each image and creates a 3D mesh model based on them.

[0759] Step 3: Import and generate product data

[0760] 3.1 The server receives product information from the manufacturer.

[0761] Input: Product photos and dimensions provided by the manufacturer

[0762] Output: Received product information

[0763] Specific operation: The server receives product information provided by the manufacturer and stores it in a database.

[0764] 3.2 The server generates a 3D model of the product.

[0765] Input: Product information

[0766] Output: 3D shape model of the product

[0767] Specific operation: The server generates a 3D model of the product based on the product information. Specifically, it uses a photo of the clothing as texture and recreates the shape of the product using measurement information.

[0768] Step 4: Incorporating the Emotion Engine

[0769] 4.1 The device collects the user's emotional data.

[0770] Input: User's facial expression data, voice data

[0771] Output: Collected emotion data

[0772] Specific operation: During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time to collect emotional data.

[0773] 4.2 The server analyzes the user's emotions using the emotion engine.

[0774] Input: Collected emotion data

[0775] Output: Parsed emotion information

[0776] How it works: The emotion analysis engine recognizes and analyzes the user's emotions based on the collected data. For example, it detects if the user is smiling and provides feedback based on that situation.

[0777] Step 5: Generate and distribute virtual try-on footage

[0778] 5.1 The server generates the virtual try-on video.

[0779] Input: 3D model of the user, 3D model of the product, analyzed emotion information

[0780] Output: Virtual try-on video

[0781] Specific operation: The server combines the 3D model of the user and the 3D model of the product to generate a virtual try-on video that reflects the user's emotional data.

[0782] 5.2 The server transmits the generated video to the user terminal.

[0783] Input: Virtual try-on video

[0784] Output: Video distribution to user devices

[0785] Specific operation: The generated virtual try-on video is sent to the user's device, allowing the user to view the video on the device.

[0786] Step 6: Video distribution and user feedback

[0787] 6.1 The device re-evaluates the user.

[0788] Input: facial expression data and voice data of the user watching the virtual try-on video

[0789] Output: Re-collected emotion data

[0790] Specific operation: While the user is viewing the virtual fitting video, the device's camera and microphone again capture the user's facial expressions and reactions.

[0791] 6.2 The server makes suggestions based on the feedback.

[0792] Input: Recollected emotion data

[0793] Output: New product proposals for users

[0794] Specific operation: The server analyzes the collected user emotion data again and generates a scenario to suggest a different product if the user's satisfaction is low. If the user's satisfaction is high, the system proceeds to the purchase procedure.

[0795] (Application example 2)

[0796] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0797] Conventional virtual try-on systems can generate try-on videos by combining a user's body data with a 3D product model, but they face challenges in providing highly personalized product recommendations that take into account the user's emotions and feedback, as well as improving the user experience. Furthermore, there is a need for an effective method to reduce product return rates while improving user satisfaction.

[0798] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0799] In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to the user terminal, means for analyzing the user's facial expression and tone of voice and extracting emotion data, means for dynamically adjusting the virtual try-on video based on the extracted emotion data, and means for making product suggestions based on the generated virtual try-on video in accordance with the user's emotions. This enables real-time analysis of user emotions and feedback and personalized product suggestions based on the analysis.

[0800] "User" refers to a person who operates a system, such as a customer or a user.

[0801] "Photography device" refers to a device used to acquire image data, such as a camera or smartphone.

[0802] "Image data" refers to data of photographs or videos captured by a photographing device.

[0803] "3D model" refers to a geometric model of digital data expressed in three dimensions.

[0804] "Product Data" refers to data containing product details, such as product dimensions and photos.

[0805] "Virtual try-on video" refers to a try-on simulation video generated by combining a 3D model of the user and a 3D model of the product.

[0806] "User terminal" refers to an information terminal operated by a user, such as a smartphone or tablet.

[0807] "Facial expression" refers to information that indicates emotions expressed through the movement of facial muscles.

[0808] "Tone of voice" refers to audio information that indicates the pitch, strength, and emotional nuances of a speaking voice.

[0809] "Emotion data" refers to data that indicates the user's emotions analyzed from facial expressions, tone of voice, etc.

[0810] "Dynamic adjustment" refers to the process of changing the video content in response to changing conditions in real time.

[0811] "Personalized product suggestions" refer to product suggestions recommended based on the individual preferences and feelings of each user.

[0812] The system for realizing this invention comprises a user terminal, a server for analyzing data, and a server for managing product data. It also includes hardware and software that are equipped with an emotion engine and can analyze user emotions. A specific embodiment of the system will be described below.

[0813] The process begins when the user launches a dedicated smartphone application and takes a full-body photo. The application displays instructions to the user, such as "Please face forward," "Please face your right side," and "Please turn your back," and prompts the user to take photos in each pose. This process generates 3D scan data of the user.

[0814] The captured 3D scan data is uploaded from the user's device to an analysis server, which then uses the data to generate accurate body shape data for the user. This analysis involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[0815] Next, the product data management server receives the product data, which includes product photos and detailed measurement information. The product data server uses this data to generate a 3D model of the product. In this process, the photos are used as textures, and the shape of the product is reproduced based on the measurement information. The 3D model is structured so that the movement of each part can be flexibly reproduced.

[0816] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice in real time. During the user's fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice, and the emotion engine extracts the user's emotional data. For example, if the user is smiling, the system determines that the user has a positive feeling toward the clothing.

[0817] The analysis server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, the content of the virtual try-on video can be dynamically adjusted based on the user's emotional data analyzed by the emotion engine. For example, if the user's emotion is unfavorable, the system can generate a scenario that suggests a different product that fits better.

[0818] The generated virtual try-on video is sent to the user's device, where the user can review the results. During this review process, the device's camera again analyzes the user's facial expressions and extracts emotional data. If the user is not satisfied, the system suggests other products, but if satisfied, the user can proceed with the purchase.

[0819] Specific examples

[0820] For example, a user can launch a dedicated smartphone app and follow the instructions to take a photo facing forward, then facing right, and finally facing away from the user, generating 3D scan data. The captured data is then uploaded to a server, where it is analyzed and detailed body shape data is generated.

[0821] While the user is trying on a jacket, the device's camera analyzes the user's facial expressions, and the emotion engine dynamically suggests products based on prompts such as, "If the user is smiling, suggest different color and style variations of that jacket." For example, when a user tries on a jacket, the emotion engine can suggest other products or customization options based on prompts such as, "If the user is not satisfied, suggest other products or customization options."

[0822] This provides users with a highly accurate and personalized virtual try-on experience, reducing return rates while increasing consumer satisfaction.

[0823] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0824] Step 1:

[0825] The user launches a dedicated application on their smartphone and takes a full-body photo. The user follows the application's instructions to pose and take photos from the front, right side, and back. Image data from multiple angles of the user is obtained as input. As output, this data is generated as 3D scan data. In this step, each image is captured and temporarily saved on the device.

[0826] Step 2:

[0827] The user terminal uploads the generated 3D scan data to the server. As input, the 3D scan data stored on the terminal is used. As output, the server stores the received image data. In this step, data is transferred over the network.

[0828] Step 3:

[0829] The server analyzes the received 3D scan data and generates accurate body shape data for the user. The 3D scan data is used as input. As output, a 3D mesh model of the user's body shape is generated. In this step, image processing techniques are used to extract dimensional information and create a 3D model based on that data.

[0830] Step 4:

[0831] The product data management server receives the product data. As input, it receives a product photo and detailed measurement information. As output, it generates a 3D model of the product. In this step, it uses the photo as a texture and processes the product shape based on the measurement information.

[0832] Step 5:

[0833] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice. When the user tries on the clothes, the device's camera and microphone capture the user's facial expressions and voice to extract emotion data. Real-time video and audio data are used as input. The output is the user's emotion data. In this step, the emotion analysis algorithm is executed.

[0834] Step 6:

[0835] The analysis server combines the generated user's body shape data with the 3D model of the product to generate a virtual try-on video. The 3D mesh model of the user and the 3D model of the product are used as input. The output is a virtual try-on video. In this step, 3D rendering technology is used to create a visually realistic try-on video.

[0836] Step 7:

[0837] The content of the virtual try-on video is dynamically adjusted based on the user's emotional data analyzed by the emotion engine. The user's emotional data is used as input. The output is an adjusted try-on video based on the user's emotions. In this step, the system changes the video content taking the user's reactions into account.

[0838] Step 8:

[0839] The generated virtual try-on video is sent to the user's device. The adjusted virtual try-on video is used as input. The video is played on the user's device as output. In this step, the data transfer and video playback processes are executed.

[0840] Step 9:

[0841] When the user checks the fitting results, the device's camera again analyzes the user's facial expressions and extracts emotional data. The input is a video of the user watching the fitting video. The output is again emotional data. In this step, emotional analysis is performed to evaluate the user's satisfaction.

[0842] Step 10:

[0843] If the satisfaction level is low, the system will suggest another product. If the satisfaction level is high, the user can proceed to the purchase process. As input, the user's emotional data is used. As output, a product suggestion and purchase process scenario is generated. In this step, the personalized product suggestion and purchase process are executed.

[0844] Through these steps, users can enjoy a highly accurate and personalized virtual try-on experience, resulting in increased satisfaction and reduced return rates.

[0845] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0846] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0847] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0848] [Third embodiment]

[0849] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0850] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0851] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0852] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0853] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0854] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0855] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0856] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0857] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0858] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0859] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0860] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0861] This invention relates to a system that 3D scans a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the image data obtained. The system consists of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0862] Program processing overview

[0863] 1. 3D scanning on your device

[0864] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[0865] 2. Uploading data to the server and processing it

[0866] The generated 3D scan data is uploaded from the user's device to a central server, which analyzes the received data and generates accurate body shape data for the user. This analysis process involves extracting the user's dimensions from each image, generating 3D point cloud data, and then constructing a mesh.

[0867] 3. Product data capture and generation

[0868] The server receives product data provided by the manufacturer, including photos of the garment and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the specific shape of the product from the measurements.

[0869] 4. Creation and distribution of virtual try-on videos

[0870] The server combines the generated user's body data with a 3D product model. This allows the scale of the product to be adjusted to fit the user's body data, and the product is positioned to fit. It also performs stretching and simulation of the clothing to recreate a more realistic fitting experience. The generated virtual fitting video is sent to the user's device.

[0871] Specific examples

[0872] 3D scanning on your device

[0873] The user opens the app and selects the photo mode. The app instructs the user to "stand facing forward," and the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, and then the app displays the instruction "Turn to your back," and the camera takes a rear-side photo. At this point, a 3D scan of the user is generated.

[0874] Data upload to server and processing

[0875] The 3D scan data taken by the user's device is uploaded to the server. The server analyzes the received data and generates data on the user's body shape. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[0876] Product data ingestion and generation

[0877] The server generates a 3D model of the product based on product photos and measurement information provided by the manufacturer. For example, a 3D model of a jacket is generated using a provided photo and detailed measurement information. This 3D model includes details such as sleeve length and collar shape.

[0878] Creation and distribution of virtual try-on videos

[0879] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. For example, the video of the user trying on the jacket they selected accurately shows how the jacket will fit the user's shoulders and waist. The server also simulates light reflection and the stretchiness of the clothing to recreate a more realistic fitting experience. The final virtual try-on video is sent to the user's device, where they can view it.

[0880] This invention allows consumers to choose the clothes that best suit them through highly accurate virtual try-on sessions when shopping online. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0881] The processing flow will be explained below.

[0882] Step 1:

[0883] The user launches the dedicated app on their device, selects 3D scan mode, and is prompted to take photos from three directions in order: the front, right side, and back.

[0884] Step 2:

[0885] The user stands facing forward and the device camera takes a full-body frontal photo of the user. After the frontal photo is taken, the app automatically saves the image temporarily on the device.

[0886] Step 3:

[0887] The user stands facing right, and the device camera takes a photo of the user's whole body from the right side. After the right side image is taken, the app automatically saves this image temporarily on the device.

[0888] Step 4:

[0889] The user stands with their back to the camera and takes a rear-view photo of the user's entire body. After the rear-view photo is taken, the app automatically saves the image temporarily on the device.

[0890] Step 5:

[0891] The device collects the image data of the front, side, and back taken by the device into a single file and sends a request to upload it to the server.

[0892] Step 6:

[0893] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts the user's body dimensions from each image and generates a 3D mesh model based on this.

[0894] Step 7:

[0895] The server receives product data provided by the manufacturer, including photos of the garment and measurements, and uses this data to generate a 3D model of the product.

[0896] Step 8:

[0897] The server combines the user's body data generated by the server with a 3D model of the product to generate a virtual fitting video, which simulates how the product will fit the user's body shape and measurements, and also realistically reproduces the reflection of light and the stretchiness of the clothing.

[0898] Step 9:

[0899] The server sends the generated virtual try-on video to the user's device, which plays the video and displays the try-on results to the user.

[0900] Step 10:

[0901] The user can view the virtual try-on video on the app, select different sizes or designs of clothing as needed, and request the generation of another try-on video. The server will repeat the same process as soon as it receives a new request.

[0902] Example 1

[0903] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0904] With traditional online shopping, customers are unable to try on items, so the size and fit of the purchased item often differ from what they actually are, leading to an increase in returns, which can result in lower consumer satisfaction, increased logistics costs, and an increased environmental impact.

[0905] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0906] In this invention, the server includes means for receiving image data captured by a user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to a user terminal, means for extracting the user's size information from each image and generating 3D point cloud data, means for constructing a mesh from the point cloud data, and means for simulating light reflection and clothing stretch. This allows consumers to experience a highly accurate virtual try-on experience, allowing them to select the most suitable clothing, and also reduces logistics and environmental burdens by reducing returns.

[0907] "Image data taken by a user from multiple directions" refers to multiple sets of image data taken by a user from different directions using an imaging device.

[0908] "Means for generating a three-dimensional model" refers to software and computational processes for digitally reconstructing the three-dimensional shape of a user or product based on received image data.

[0909] "Product data" refers to information including various data such as detailed product descriptions, dimensional information, and photographs.

[0910] "Means for generating virtual try-on footage" refers to software and computational processes that combine a 3D model of the user with a 3D model of the product to visually recreate the state of the user virtually trying on the product.

[0911] "Means for transmitting the virtual try-on video to the user terminal" refers to a communication means for transmitting the generated virtual try-on video to the user's device via a network such as the Internet.

[0912] "Means for extracting dimensional information and generating 3D point cloud data" refers to software and computational processes for measuring the user's body dimensions from the received image data and generating points in 3D space based on that information.

[0913] "Means for constructing a mesh from point cloud data" refers to software and a calculation process for connecting the surface of a three-dimensional shape with triangles or polygons based on the generated point cloud data to generate a continuous mesh.

[0914] "Means for simulating light reflection and clothing stretch" refers to software and computational processes for calculating and reproducing in real time the reflection of light and the dynamic changes that occur as clothing fits the user's body in a virtual environment in which the user is trying on the clothing.

[0915] This invention relates to a system that performs a 3D scan of a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the obtained image data. This system is composed of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[0916] 1. (3D scanning on device)

[0917] The user launches the dedicated smartphone app and selects the shooting mode. The app instructs the user to take three poses: front, right side, and back, and the camera takes a photo for each pose. The captured image data is temporarily stored on the device.

[0918] For example, when a user launches the app and selects the photo mode, the instruction "Please stand facing forward" is displayed. When the user stands facing forward, the camera automatically takes a photo, and then the instruction "Please turn to the right" is displayed. This process is repeated to generate 3D scan data.

[0919] 2. (Uploading data to the server and analyzing it)

[0920] The 3D scan data generated by the device is uploaded to a server. The server extracts the user's dimensional information from each image and generates 3D point cloud data. A mesh is then constructed from the point cloud data. This analysis process uses software such as Python's OpenCV and Point Cloud Library (PCL).

[0921] For example, when a user uploads image data they have taken to a server, the server analyzes the pixel information of each photo and calculates the user's height, shoulder width, waist size, etc. Then, it automatically generates 3D point cloud data and meshes based on this information.

[0922] 3. (Importing product data and generating 3D models)

[0923] The server receives product data (photos and dimensional information) provided by the manufacturer and generates a 3D model of the product based on this data. An image processing library (e.g., OpenCV) is used to use the photo as a texture and to recreate the specific shape from the dimensional information.

[0924] Example: The server receives a photo of a jacket provided by a manufacturer and detailed measurement information, and based on this information, generates a 3D model of the jacket, including details such as sleeve length and collar shape.

[0925] 4. (Generation and distribution of virtual try-on videos)

[0926] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. The scale of the product is adjusted and positioned to fit the user's body. 3D rendering software such as Blender or Unity is also used to simulate light reflection and the stretchiness of the clothing. The final virtual try-on video is sent to the user's device.

[0927] Example: A 3D model of a jacket is adjusted to fit the user's body shape data, and a virtual try-on video is generated. This video shows how the jacket fits perfectly to the user's shoulders and waist. The generated video is sent to the user's device in real time, and the user can view it on their smartphone.

[0928] Examples of prompts:

[0929] "Please generate a video of user A trying on jacket B based on the 3D scan data."

[0930] This invention allows users to perform highly accurate virtual try-on sessions and select clothing that best suits them. It also reduces the burden on logistics and the environmental impact by reducing returns.

[0931] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0932] Step 1:

[0933] The user launches a dedicated app on their smartphone and selects a shooting mode.

[0934] Input: The user operates the app to select a shooting mode.

[0935] Specific operation: The user selects "3D scan mode" and the app launches.

[0936] Output: The app enters shooting mode.

[0937] Step 2:

[0938] The app asks the user to pose from the front, right side, and back.

[0939] Input: The user acts according to the app's instructions.

[0940] Specific behavior: The app displays the instruction "Please stand facing forward," and the user faces forward.

[0941] Output: User facing forward.

[0942] Step 3:

[0943] The device's camera takes a photo for each pose.

[0944] Input: The user strikes a pose and the device camera activates.

[0945] Specific operation: The camera automatically takes photos of the front, right side, and back in sequence.

[0946] Output: Image data from three directions.

[0947] Step 4:

[0948] The captured image data is temporarily stored on the device.

[0949] Input: Image data captured by a camera.

[0950] Specific operation: The captured image data is saved in the smartphone's temporary memory.

[0951] Output: Saved image data.

[0952] Step 5:

[0953] The device uploads the generated 3D scan data to the server.

[0954] Input: Saved image data.

[0955] Specific operation: The device sends image data to the server via an Internet connection.

[0956] Output: Image data uploaded to the server.

[0957] Step 6:

[0958] The server extracts the user's dimensional information from each image and generates 3D point cloud data.

[0959] Input: Image data uploaded to the server.

[0960] How it works: The server uses Python's OpenCV and SciPy libraries to analyze pixel information from each photo and extract measurements such as the user's height, shoulder width, and waist.

[0961] Output: Extracted dimensional information and 3D point cloud data.

[0962] Step 7:

[0963] The server constructs a mesh from the point cloud data.

[0964] Input: Extracted dimensional information and 3D point cloud data.

[0965] Specific operation: The server generates a 3D mesh model using Blender or Point Cloud Library (PCL).

[0966] Output: Reconstructed 3D mesh data.

[0967] Step 8:

[0968] The server receives product data provided by the manufacturer.

[0969] Input: Product data (photos and dimensions) provided by the manufacturer.

[0970] Specific operation: The server receives and stores the product data.

[0971] Output: Received product data.

[0972] Step 9:

[0973] The server generates a 3D model of the product based on the product data.

[0974] Input: Received product data.

[0975] Specific operation: The server uses an image processing library (e.g., OpenCV) to use the photo as a texture and reproduce the specific shape of the product from the dimensional information.

[0976] Output: A 3D model of the generated product.

[0977] Step 10:

[0978] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video.

[0979] Input: User's 3D mesh data and 3D model of the product.

[0980] How it works: The server adjusts the scale of the product and positions it to fit the user's body shape. It also uses Blender or Unity's physics engine to simulate light reflection and clothing stretching.

[0981] Output: Generated virtual try-on video.

[0982] Step 11:

[0983] The server transmits the generated virtual try-on video to the user's terminal.

[0984] Input: Generated virtual try-on footage.

[0985] Specific operation: The server sends the virtual try-on video to the user's device via the Internet.

[0986] Output: A virtual try-on video displayed on the user's device.

[0987] (Application example 1)

[0988] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0989] With traditional online shopping, consumers are unable to actually try on products, resulting in frequent problems such as products not fitting properly or looking different after purchase. This not only increases the return rate, burdens on logistics and the environment, but also reduces consumer satisfaction. Furthermore, existing virtual try-on systems have low accuracy in user body shape data, making it difficult to reproduce an actual fit. To solve these issues, a new system that combines high-precision 3D scanning technology and realistic try-on simulations is needed.

[0990] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0991] In this invention, the server includes: means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data; means for receiving product data and generating a 3D product model from the product data; means for combining the 3D user model and the 3D product model to generate a virtual try-on video; means for transmitting the virtual try-on video to the user device; means for instructing the user to pose when capturing images using the user device's camera or sensor; means for uploading the user's image data to the server and analyzing it to extract detailed dimensional information about the user; means for constructing a 3D model of the user based on the specific dimensional information; means for specifically reproducing the 3D product model based on the product dimensional information; and means for simulating the stretch and shadow of clothing when generating the virtual try-on video. This allows consumers to experience a highly accurate virtual try-on experience and choose the clothing that best suits them. Furthermore, reducing returns can reduce logistics and environmental burdens.

[0992] "Image data taken by a user from multiple directions" refers to a collection of image data taken by a user using a photographing device from different directions such as the front, side, and rear.

[0993] A "3D model" is three-dimensional digital data that reproduces the shape of a user or product from two-dimensional image data.

[0994] "Virtual try-on video" is a video in which a product model is matched to a 3D model of the user, making it appear as if the user is actually trying on the product.

[0995] A "user terminal" is an electronic device that is directly operated by a user, such as a smartphone, tablet, or computer.

[0996] An "imaging device" is a device used to capture images, such as a camera or sensor mounted on a user terminal.

[0997] A "server" is a computer system used for analyzing, storing, and communicating data.

[0998] "Dimensional information" is detailed information about the size of the product and dimensional data of each part of the user's body.

[0999] "Analysis processing" refers to the calculations and processing required to extract necessary information based on received data and generate a model.

[1000] "Detailed dimensional information" refers to the precise dimensions of parts of a user or product, as well as detailed measurement data.

[1001] "Stretching and shadow simulation" is a computer graphics process that recreates the realistic feeling of trying on clothes, including the stretching and shrinking of clothing and the reflection of light.

[1002] The system of the present invention generates a 3D model from image data captured by the user from multiple angles using a camera, and then combines the generated model with a 3D model of the product to provide a virtual try-on video. This system allows users to experience highly accurate virtual try-on from the comfort of their own home, enabling them to choose the clothing that best suits them.

[1003] Hardware and software used

[1004] Device:

[1005] Smartphone (with high-resolution camera and standard IMU (Inertial Measurement Unit))

[1006] server:

[1007] Data analysis server (using TensorFlow and OpenCV)

[1008] 3D rendering server (using Blender)

[1009] Communication (using REST API)

[1010] Data processing and calculation

[1011] Generate a 3D model of the user

[1012] 1. Image capture:

[1013] The user launches the app and takes photos of the front, right side, and back using the device's camera. During this process, the app displays instructions to help the user stand in the correct position and angle.

[1014] 2. Upload image data:

[1015] The captured image data is temporarily stored on the device and then uploaded to the server.

[1016] 3. Analysis process:

[1017] The server analyzes the received image data and extracts detailed dimensional information about the user. Specifically, it uses TensorFlow for image recognition and dimension extraction, OpenCV to generate point cloud data, and Blender to construct the final 3D mesh.

[1018] 3D product model generation

[1019] 1. Product data acquisition:

[1020] The server receives product photos and detailed dimensional information provided by the manufacturer.

[1021] 2. Product model generation:

[1022] Using the required dimensions and photographs, Blender is used to generate a 3D model of the product, including the product's specific shape and texture.

[1023] Virtual try-on video generation and distribution

[1024] 1. Model combination:

[1025] The server combines a 3D model of the user with a 3D model of the product to realistically simulate the fit, simulating the stretch and shadow of the clothing to recreate the feeling of trying it on.

[1026] 2. Image generation:

[1027] The virtual try-on video generated by the above process is sent to the user's device, where the user can view the video and choose the clothing that best suits them.

[1028] Specific examples

[1029] Image capture and upload

[1030] The user opens the app and takes a frontal image following the instruction "Please stand facing forward," then takes a right-side image following the instruction "Please stand facing right," and then takes a backside image following the instruction "Please stand with your back to the camera." This series of images is then uploaded to the server.

[1031] Server-side processing

[1032] The server analyzes the received image data and extracts detailed dimensional information such as the user's height, shoulder width, waist, etc. Based on the extracted dimensional information, point cloud data is generated using TensorFlow and OpenCV, and a 3D mesh is constructed using Blender.

[1033] 3D product model generation

[1034] The server receives jacket photos and measurements provided by the manufacturer and uses Blender to generate a 3D model of the jacket, including sleeve length, collar shape, texture information, and more.

[1035] Virtual try-on video generation

[1036] The generated user model is combined with the jacket model to simulate the fit, light reflection, and stretch of the clothing, generating a realistic video of the user trying on the garment. This video is then sent to the user's device.

[1037] Specific prompt examples

[1038] "Please explain the overview of the system that enables highly accurate virtual try-on based on a full-body 3D scan."

[1039] The details of this system depend on the user's environment, providing a convenient experience that can be experienced individually by each user. This will improve the online shopping experience for consumers, reduce the return rate, and ease the burden on logistics.

[1040] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1041] Step 1:

[1042] Image capture

[1043] Input: Image data from multiple angles according to user instructions

[1044] How it works: The user launches the dedicated app and takes photos of three poses using the device's camera: the front, right side, and back.

[1045] Output: Three images taken by the user (front, right, and back images)

[1046] Step 2:

[1047] Image data upload

[1048] Input: User image data captured in step 1

[1049] Operation: The device uploads the image data it has taken to the server. The app uses the REST API to send the image data to the server.

[1050] Output: User image data transferred to the server

[1051] Step 3:

[1052] Analysis processing

[1053] Input: User image data uploaded to the server

[1054] How it works: The server uses TensorFlow and OpenCV to extract the user's dimensions from each image, generates point cloud data from each image, and builds a 3D mesh using Blender.

[1055] Output: Detailed dimensional information and your 3D model

[1056] Step 4:

[1057] Get product data

[1058] Input: Product photos and dimensions provided by the manufacturer

[1059] How it works: The server receives product data, generates a 3D model based on the necessary dimensions and photos, and uses Blender to incorporate product shape and texture information into the 3D model.

[1060] Output: 3D model of the product

[1061] Step 5:

[1062] Virtual try-on video generation

[1063] Input: 3D model of the user and 3D model of the product

[1064] How it works: The server combines the user model with the product model, simulates fit, stretching, and shadows, and uses Blender to generate a virtual try-on video.

[1065] Output: Virtual try-on video data

[1066] Step 6:

[1067] Distribution of virtual try-on videos

[1068] Input: Generated virtual try-on video data

[1069] Operation: The server generates a virtual fitting video and sends it to the user's device. The user can view the video on their device and check how the clothes feel when they are tried on.

[1070] Output: Virtual try-on video that can be viewed by the user

[1071] In this way, each processing step of the system is combined to provide the user with a highly accurate virtual try-on experience.

[1072] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1073] This invention relates to a system that provides a virtual try-on experience by 3D scanning a user's entire body with a camera such as a smartphone, generating accurate body shape data from the obtained image data, and combining this with an emotion engine. The system consists of a user's device, a server that analyzes the data, a server that generates 3D product models, the emotion engine, and a server that generates virtual try-on videos and sends them to the user's device.

[1074] Program processing overview

[1075] 1. 3D scanning on your device

[1076] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[1077] 2. Uploading data to the server and processing it

[1078] The generated 3D scan data is uploaded from the user's device to a server, which analyzes the received data and generates accurate data on the user's body shape. The analysis process involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[1079] 3. Product data capture and generation

[1080] The server receives product data provided by the manufacturer, including photos of the garments and detailed measurement information, and generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the shape of the product based on the measurement information.

[1081] 4. Incorporating an Emotional Engine

[1082] While the user is trying on the clothes, the device's built-in camera and microphone analyze the user's facial expressions and tone of voice in real time, and the emotion engine recognizes the user's emotions. For example, if the system detects that the user is smiling, it will determine that the user has a positive feeling toward the clothing.

[1083] 5. Creation and distribution of virtual try-on videos

[1084] The server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, it uses the user's emotional data analyzed by the emotion engine to adjust the content of the virtual try-on video. For example, if the user's emotion is unfavorable, it can generate a scenario that suggests a different product that fits better.

[1085] 6. Video distribution and user feedback

[1086] The generated virtual try-on video is sent to the user's device for viewing. The user's try-on experience is then analyzed for emotion, and feedback is provided based on the user's emotions. For example, if the user is dissatisfied, the system automatically suggests other products.

[1087] Specific examples

[1088] 3D scanning on your device

[1089] The user launches the dedicated app and selects the shooting mode. When the app instructs them to "stand facing forward," the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, then the app displays "Turn your back," and the app takes a rear-side photo. At this point, 3D scan data of the user is generated.

[1090] Data upload to server and processing

[1091] The 3D scan data taken by the user's device is uploaded to the server, which then analyzes the data and generates accurate body shape data for the user. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[1092] Product data ingestion and generation

[1093] The server generates a 3D model of the product based on the product photos and measurements provided by the manufacturer. The server uses the provided jacket photos and measurements to generate a 3D model of the jacket, including details such as sleeve length and collar shape.

[1094] Incorporating an emotion engine

[1095] When a user tries on a jacket during the fitting experience, the device's camera analyzes the user's facial expressions and the emotion engine recognizes the user's emotions. For example, if the user is smiling, the emotion engine determines that the user likes the jacket.

[1096] Creation and distribution of virtual try-on videos

[1097] The server combines the user's body shape data with a 3D model of the product to generate a virtual try-on video. The emotion engine adjusts the video content based on the user's emotional data. For example, if the user is smiling, different colors and styles of jackets will be suggested in the video.

[1098] Video distribution and user feedback

[1099] The generated virtual try-on video is sent to the user's device, where the user reviews it. While reviewing, the emotion engine analyzes the user's facial expressions again to determine their level of satisfaction. If their satisfaction is low, the system suggests a different product. If their satisfaction is high, they can proceed with the purchase process.

[1100] This not only allows users to have a highly accurate virtual try-on experience, but also allows them to receive more personalized product recommendations based on their emotions, resulting in increased consumer satisfaction and reduced returns.

[1101] The processing flow will be explained below.

[1102] Step 1:

[1103] The user launches the dedicated app and selects 3D scan mode. The app then prompts the user to take photos from the front, right side, and back in that order.

[1104] Step 2:

[1105] The user stands facing forward, and the device camera takes a full-body frontal photograph of the user. The captured image is immediately saved temporarily on the device.

[1106] Step 3:

[1107] The user stands facing right, and the device camera takes a photo of the user's entire body from the right side. As with the front-facing photo, this image is also temporarily saved on the device.

[1108] Step 4:

[1109] The user stands with their back to the camera and takes a photo of the user's entire body from the back. The image of the user's back is also temporarily stored on the device.

[1110] Step 5:

[1111] The device compiles the 3D scan data of the front, side, and back and sends a request to upload it to the server.

[1112] Step 6:

[1113] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts dimensional information such as the user's height, shoulder width, and waist from each image and generates a 3D mesh model based on this information.

[1114] Step 7:

[1115] The server receives product data provided by the manufacturer, including photos of the garment and dimensions of each part, and generates a 3D model of the product based on this data.

[1116] Step 8:

[1117] The device's camera and microphone capture the user's facial expressions and tone of voice in real time and send the data to the emotion engine, which analyzes this data and recognizes the user's emotions.

[1118] Step 9:

[1119] The emotion engine sends the user's emotion data to the server, which then takes this emotion data into account when generating the virtual try-on video. For example, if the user expresses positive emotion, the video can include content suggesting different variations of the product.

[1120] Step 10:

[1121] The server combines the user's body data with a 3D model of the product to generate a virtual fitting video. The simulation includes the stretching and light reflection of the clothing, providing an experience as close as possible to a real try-on.

[1122] Step 11:

[1123] The generated virtual fitting video is sent to the user's device, where the user can view it. The device continues to send the user's facial expressions to the emotion engine while the video is playing, and analyzes their emotions in real time.

[1124] Step 12:

[1125] The server provides feedback based on the user's emotional data. If the user is not satisfied, the server generates content suggesting alternative products and adjusts it to attract the user's interest.

[1126] Step 13:

[1127] If the user is satisfied, the app will either proceed with the purchase or offer further suggestions, and the entire system will adjust product selection accordingly.

[1128] This not only allows users to have a highly accurate virtual try-on experience, but also allows for more personalized product recommendations based on their emotions, which in turn increases consumer satisfaction and reduces returns.

[1129] Example 2

[1130] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1131] Conventional virtual try-on systems have low accuracy in generating 3D shape models of users and have difficulty in proposing products that appropriately reflect the user's emotions. As a result, user satisfaction declines and ultimately, product returns increase. Furthermore, there is a lack of technology to analyze the user's emotional state in real time and provide feedback to the virtual try-on video.

[1132] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a three-dimensional shape model of the user from this image data, means for receiving product information and generating a three-dimensional shape model of the product from this product information, means for combining the user's three-dimensional shape model and the product's three-dimensional shape model to generate a virtual try-on video, means for analyzing the user's emotional state using an emotion analysis engine and adjusting the content of the virtual try-on video, and means for transmitting the virtual try-on video to the user terminal. This improves the accuracy of the user's three-dimensional shape model, making it possible to recommend products that reflect the user's emotions, which is expected to improve user satisfaction and reduce returned products.

[1133] "Image data taken by a user from multiple directions" refers to image data of the user taken from different directions using a photographing device owned by the user.

[1134] A "three-dimensional shape model" is a model that represents a shape in three-dimensional space and is generated from image data of a user or a product.

[1135] "Product information" refers to data provided by manufacturers and providers, including product photos and detailed dimensional information.

[1136] An "emotion analysis engine" is software or hardware that can analyze a user's facial expressions and tone of voice to recognize and determine the user's emotional state.

[1137] A "virtual try-on video" is a video that simulates the experience of trying on clothes, generated by combining a three-dimensional shape model of the user and a three-dimensional shape model of the product.

[1138] "User terminal" refers to an electronic device used by a user to display and operate information, such as a smartphone, tablet, or PC.

[1139] This system provides a virtual try-on experience by 3D scanning a user's entire body with a camera, generating accurate body shape data from the image data, and combining this with an emotion analysis engine. This system is comprised of a user's device, a server that analyzes the data, a server that generates a 3D product model, an emotion analysis engine, and a server that generates a virtual try-on video and sends it to the user's device.

[1140] The process begins when the user launches a dedicated smartphone app and takes a full-body photo. The app displays prompts such as "Please stand facing forward," "Please turn to your right," and "Please stand with your back to the camera," and takes images from each direction based on the user's current posture. After the photos are taken, the image data is temporarily stored on the device.

[1141] The user's device then uploads the stored 3D scan data to a server, which analyzes the data to generate accurate body shape data for the user. The analysis process uses computer vision techniques and machine learning algorithms to extract the user's dimensions from each image and create a 3D mesh model based on that information.

[1142] The server also receives product information from manufacturers, including product photos and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses clothing photos as textures and reproduces the product shape using measurements.

[1143] During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time. An emotion analysis engine recognizes the user's emotions and adjusts the content of the fitting video based on that information. For example, if the user is smiling, the system will determine that they have a positive attitude toward the product and reflect this in the video. If the user's emotions are not positive, the system can also generate a scenario that suggests alternative products.

[1144] The generated virtual try-on video is sent to the user's device, where the user reviews it. During the review, sentiment analysis is performed again and the user's feedback is provided. If the user's satisfaction level is low, the system automatically suggests other products. This allows the user to have a highly accurate virtual try-on experience and also receive more personalized product suggestions based on their emotions.

[1145] Example prompt sentence:

[1146] "Please stand facing forward."

[1147] "Please turn to your right."

[1148] "Please stand with your back to me."

[1149] The above is a specific embodiment based on the present invention for providing a highly accurate virtual try-on experience for the user.

[1150] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1151] Processing flow

[1152] Step 1: 3D scan on your device

[1153] 1.1 The user launches the dedicated app and selects the shooting mode.

[1154] Input: User actions

[1155] Output: Start camera, display prompt

[1156] Specific operation: The user opens the dedicated smartphone app and selects the shooting mode. The app then displays the message, "Please stand facing forward."

[1157] 1.2 The app asks the user to pose and the camera takes the photo.

[1158] Input: current user posture, device camera input

[1159] Output: Image data of the front, right side, and back of the user

[1160] Specific operation: The user follows the instructions of the app to face the front, right side, and back, and the device camera takes a photo of each. The captured image data is temporarily stored on the device.

[1161] Step 2: Upload data to the server and process it

[1162] 2.1 The device uploads data to the server.

[1163] Input: Image data stored on the device

[1164] Output: Uploaded image data

[1165] How it works: Once the user has taken a photo, the app sends the image data to a server, where it is uploaded using a secure protocol.

[1166] 2.2 The server receives and analyzes the data.

[1167] Input: Uploaded image data

[1168] Output: 3D model of the user

[1169] How it works: The server reviews the received image data and analyzes it using computer vision techniques and machine learning algorithms. It extracts the user's dimensions from each image and creates a 3D mesh model based on them.

[1170] Step 3: Import and generate product data

[1171] 3.1 The server receives product information from the manufacturer.

[1172] Input: Product photos and dimensions provided by the manufacturer

[1173] Output: Received product information

[1174] Specific operation: The server receives product information provided by the manufacturer and stores it in a database.

[1175] 3.2 The server generates a 3D model of the product.

[1176] Input: Product information

[1177] Output: 3D shape model of the product

[1178] Specific operation: The server generates a 3D model of the product based on the product information. Specifically, it uses a photo of the clothing as texture and recreates the shape of the product using measurement information.

[1179] Step 4: Incorporating the Emotion Engine

[1180] 4.1 The device collects the user's emotional data.

[1181] Input: User's facial expression data, voice data

[1182] Output: Collected emotion data

[1183] Specific operation: During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time to collect emotional data.

[1184] 4.2 The server analyzes the user's emotions using the emotion engine.

[1185] Input: Collected emotion data

[1186] Output: Parsed emotion information

[1187] How it works: The emotion analysis engine recognizes and analyzes the user's emotions based on the collected data. For example, it detects if the user is smiling and provides feedback based on that situation.

[1188] Step 5: Generate and distribute virtual try-on footage

[1189] 5.1 The server generates the virtual try-on video.

[1190] Input: 3D model of the user, 3D model of the product, analyzed emotion information

[1191] Output: Virtual try-on video

[1192] Specific operation: The server combines the 3D model of the user and the 3D model of the product to generate a virtual try-on video that reflects the user's emotional data.

[1193] 5.2 The server transmits the generated video to the user terminal.

[1194] Input: Virtual try-on video

[1195] Output: Video distribution to user devices

[1196] Specific operation: The generated virtual try-on video is sent to the user's device, allowing the user to view the video on the device.

[1197] Step 6: Video distribution and user feedback

[1198] 6.1 The device re-evaluates the user.

[1199] Input: facial expression data and voice data of the user watching the virtual try-on video

[1200] Output: Re-collected emotion data

[1201] Specific operation: While the user is viewing the virtual fitting video, the device's camera and microphone again capture the user's facial expressions and reactions.

[1202] 6.2 The server makes suggestions based on the feedback.

[1203] Input: Recollected emotion data

[1204] Output: New product proposals for users

[1205] Specific operation: The server analyzes the collected user emotion data again and generates a scenario to suggest a different product if the user's satisfaction is low. If the user's satisfaction is high, the system proceeds to the purchase procedure.

[1206] (Application example 2)

[1207] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1208] Conventional virtual try-on systems can generate try-on videos by combining a user's body data with a 3D product model, but they face challenges in providing highly personalized product recommendations that take into account the user's emotions and feedback, as well as improving the user experience. Furthermore, there is a need for an effective method to reduce product return rates while improving user satisfaction.

[1209] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1210] In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to the user terminal, means for analyzing the user's facial expression and tone of voice and extracting emotion data, means for dynamically adjusting the virtual try-on video based on the extracted emotion data, and means for making product suggestions based on the generated virtual try-on video in accordance with the user's emotions. This enables real-time analysis of user emotions and feedback and personalized product suggestions based on the analysis.

[1211] "User" refers to a person who operates a system, such as a customer or a user.

[1212] "Photography device" refers to a device used to acquire image data, such as a camera or smartphone.

[1213] "Image data" refers to data of photographs or videos captured by a photographing device.

[1214] "3D model" refers to a geometric model of digital data expressed in three dimensions.

[1215] "Product Data" refers to data containing product details, such as product dimensions and photos.

[1216] "Virtual try-on video" refers to a try-on simulation video generated by combining a 3D model of the user and a 3D model of the product.

[1217] "User terminal" refers to an information terminal operated by a user, such as a smartphone or tablet.

[1218] "Facial expression" refers to information that indicates emotions expressed through the movement of facial muscles.

[1219] "Tone of voice" refers to audio information that indicates the pitch, strength, and emotional nuances of a speaking voice.

[1220] "Emotion data" refers to data that indicates the user's emotions analyzed from facial expressions, tone of voice, etc.

[1221] "Dynamic adjustment" refers to the process of changing the video content in response to changing conditions in real time.

[1222] "Personalized product suggestions" refer to product suggestions recommended based on the individual preferences and feelings of each user.

[1223] The system for realizing this invention comprises a user terminal, a server for analyzing data, and a server for managing product data. It also includes hardware and software that are equipped with an emotion engine and can analyze user emotions. A specific embodiment of the system will be described below.

[1224] The process begins when the user launches a dedicated smartphone application and takes a full-body photo. The application displays instructions to the user, such as "Please face forward," "Please face your right side," and "Please turn your back," and prompts the user to take photos in each pose. This process generates 3D scan data of the user.

[1225] The captured 3D scan data is uploaded from the user's device to an analysis server, which then uses the data to generate accurate body shape data for the user. This analysis involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[1226] Next, the product data management server receives the product data, which includes product photos and detailed measurement information. The product data server uses this data to generate a 3D model of the product. In this process, the photos are used as textures, and the shape of the product is reproduced based on the measurement information. The 3D model is structured so that the movement of each part can be flexibly reproduced.

[1227] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice in real time. During the user's fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice, and the emotion engine extracts the user's emotional data. For example, if the user is smiling, the system determines that the user has a positive feeling toward the clothing.

[1228] The analysis server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, the content of the virtual try-on video can be dynamically adjusted based on the user's emotional data analyzed by the emotion engine. For example, if the user's emotion is unfavorable, the system can generate a scenario that suggests a different product that fits better.

[1229] The generated virtual try-on video is sent to the user's device, where the user can review the results. During this review process, the device's camera again analyzes the user's facial expressions and extracts emotional data. If the user is not satisfied, the system suggests other products, but if satisfied, the user can proceed with the purchase.

[1230] Specific examples

[1231] For example, a user can launch a dedicated smartphone app and follow the instructions to take a photo facing forward, then facing right, and finally facing away from the user, generating 3D scan data. The captured data is then uploaded to a server, where it is analyzed and detailed body shape data is generated.

[1232] While the user is trying on a jacket, the device's camera analyzes the user's facial expressions, and the emotion engine dynamically suggests products based on prompts such as, "If the user is smiling, suggest different color and style variations of that jacket." For example, when a user tries on a jacket, the emotion engine can suggest other products or customization options based on prompts such as, "If the user is not satisfied, suggest other products or customization options."

[1233] This provides users with a highly accurate and personalized virtual try-on experience, reducing return rates while increasing consumer satisfaction.

[1234] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1235] Step 1:

[1236] The user launches a dedicated application on their smartphone and takes a full-body photo. The user follows the application's instructions to pose and take photos from the front, right side, and back. Image data from multiple angles of the user is obtained as input. As output, this data is generated as 3D scan data. In this step, each image is captured and temporarily saved on the device.

[1237] Step 2:

[1238] The user terminal uploads the generated 3D scan data to the server. As input, the 3D scan data stored on the terminal is used. As output, the server stores the received image data. In this step, data is transferred over the network.

[1239] Step 3:

[1240] The server analyzes the received 3D scan data and generates accurate body shape data for the user. The 3D scan data is used as input. As output, a 3D mesh model of the user's body shape is generated. In this step, image processing techniques are used to extract dimensional information and create a 3D model based on that data.

[1241] Step 4:

[1242] The product data management server receives the product data. As input, it receives a product photo and detailed measurement information. As output, it generates a 3D model of the product. In this step, it uses the photo as a texture and processes the product shape based on the measurement information.

[1243] Step 5:

[1244] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice. When the user tries on the clothes, the device's camera and microphone capture the user's facial expressions and voice to extract emotion data. Real-time video and audio data are used as input. The output is the user's emotion data. In this step, the emotion analysis algorithm is executed.

[1245] Step 6:

[1246] The analysis server combines the generated user's body shape data with the 3D model of the product to generate a virtual try-on video. The 3D mesh model of the user and the 3D model of the product are used as input. The output is a virtual try-on video. In this step, 3D rendering technology is used to create a visually realistic try-on video.

[1247] Step 7:

[1248] The content of the virtual try-on video is dynamically adjusted based on the user's emotional data analyzed by the emotion engine. The user's emotional data is used as input. The output is an adjusted try-on video based on the user's emotions. In this step, the system changes the video content taking the user's reactions into account.

[1249] Step 8:

[1250] The generated virtual try-on video is sent to the user's device. The adjusted virtual try-on video is used as input. The video is played on the user's device as output. In this step, the data transfer and video playback processes are executed.

[1251] Step 9:

[1252] When the user checks the fitting results, the device's camera again analyzes the user's facial expressions and extracts emotional data. The input is a video of the user watching the fitting video. The output is again emotional data. In this step, emotional analysis is performed to evaluate the user's satisfaction.

[1253] Step 10:

[1254] If the satisfaction level is low, the system will suggest another product. If the satisfaction level is high, the user can proceed to the purchase process. As input, the user's emotional data is used. As output, a product suggestion and purchase process scenario is generated. In this step, the personalized product suggestion and purchase process are executed.

[1255] Through these steps, users can enjoy a highly accurate and personalized virtual try-on experience, resulting in increased satisfaction and reduced return rates.

[1256] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1257] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1258] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1259] [Fourth embodiment]

[1260] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1261] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1262] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1263] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1264] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1265] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1266] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1267] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1268] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1269] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1270] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1271] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1272] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1273] This invention relates to a system that 3D scans a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the image data obtained. The system consists of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[1274] Program processing overview

[1275] 1. 3D scanning on your device

[1276] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[1277] 2. Uploading data to the server and processing it

[1278] The generated 3D scan data is uploaded from the user's device to a central server, which analyzes the received data and generates accurate body shape data for the user. This analysis process involves extracting the user's dimensions from each image, generating 3D point cloud data, and then constructing a mesh.

[1279] 3. Product data capture and generation

[1280] The server receives product data provided by the manufacturer, including photos of the garment and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the specific shape of the product from the measurements.

[1281] 4. Creation and distribution of virtual try-on videos

[1282] The server combines the generated user's body data with a 3D product model. This allows the scale of the product to be adjusted to fit the user's body data, and the product is positioned to fit. It also performs stretching and simulation of the clothing to recreate a more realistic fitting experience. The generated virtual fitting video is sent to the user's device.

[1283] Specific examples

[1284] 3D scanning on your device

[1285] The user opens the app and selects the photo mode. The app instructs the user to "stand facing forward," and the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, and then the app displays the instruction "Turn to your back," and the camera takes a rear-side photo. At this point, a 3D scan of the user is generated.

[1286] Data upload to server and processing

[1287] The 3D scan data taken by the user's device is uploaded to the server. The server analyzes the received data and generates data on the user's body shape. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[1288] Product data ingestion and generation

[1289] The server generates a 3D model of the product based on product photos and measurement information provided by the manufacturer. For example, a 3D model of a jacket is generated using a provided photo and detailed measurement information. This 3D model includes details such as sleeve length and collar shape.

[1290] Creation and distribution of virtual try-on videos

[1291] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. For example, the video of the user trying on the jacket they selected accurately shows how the jacket will fit the user's shoulders and waist. The server also simulates light reflection and the stretchiness of the clothing to recreate a more realistic fitting experience. The final virtual try-on video is sent to the user's device, where they can view it.

[1292] This invention allows consumers to choose the clothes that best suit them through highly accurate virtual try-on sessions when shopping online. It also reduces the burden on logistics and the environmental impact by reducing returns.

[1293] The processing flow will be explained below.

[1294] Step 1:

[1295] The user launches the dedicated app on their device, selects 3D scan mode, and is prompted to take photos from three directions in order: the front, right side, and back.

[1296] Step 2:

[1297] The user stands facing forward and the device camera takes a full-body frontal photo of the user. After the frontal photo is taken, the app automatically saves the image temporarily on the device.

[1298] Step 3:

[1299] The user stands facing right, and the device camera takes a photo of the user's whole body from the right side. After the right side image is taken, the app automatically saves this image temporarily on the device.

[1300] Step 4:

[1301] The user stands with their back to the camera and takes a rear-view photo of the user's entire body. After the rear-view photo is taken, the app automatically saves the image temporarily on the device.

[1302] Step 5:

[1303] The device collects the image data of the front, side, and back taken by the device into a single file and sends a request to upload it to the server.

[1304] Step 6:

[1305] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts the user's body dimensions from each image and generates a 3D mesh model based on this.

[1306] Step 7:

[1307] The server receives product data provided by the manufacturer, including photos of the garment and measurements, and uses this data to generate a 3D model of the product.

[1308] Step 8:

[1309] The server combines the user's body data generated by the server with a 3D model of the product to generate a virtual fitting video, which simulates how the product will fit the user's body shape and measurements, and also realistically reproduces the reflection of light and the stretchiness of the clothing.

[1310] Step 9:

[1311] The server sends the generated virtual try-on video to the user's device, which plays the video and displays the try-on results to the user.

[1312] Step 10:

[1313] The user can view the virtual try-on video on the app, select different sizes or designs of clothing as needed, and request the generation of another try-on video. The server will repeat the same process as soon as it receives a new request.

[1314] Example 1

[1315] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1316] With traditional online shopping, customers are unable to try on items, so the size and fit of the purchased item often differ from what they actually are, leading to an increase in returns, which can result in lower consumer satisfaction, increased logistics costs, and an increased environmental impact.

[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1318] In this invention, the server includes means for receiving image data captured by a user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to a user terminal, means for extracting the user's size information from each image and generating 3D point cloud data, means for constructing a mesh from the point cloud data, and means for simulating light reflection and clothing stretch. This allows consumers to experience a highly accurate virtual try-on experience, allowing them to select the most suitable clothing, and also reduces logistics and environmental burdens by reducing returns.

[1319] "Image data taken by a user from multiple directions" refers to multiple sets of image data taken by a user from different directions using an imaging device.

[1320] "Means for generating a three-dimensional model" refers to software and computational processes for digitally reconstructing the three-dimensional shape of a user or product based on received image data.

[1321] "Product data" refers to information including various data such as detailed product descriptions, dimensional information, and photographs.

[1322] "Means for generating virtual try-on footage" refers to software and computational processes that combine a 3D model of the user with a 3D model of the product to visually recreate the state of the user virtually trying on the product.

[1323] "Means for transmitting the virtual try-on video to the user terminal" refers to a communication means for transmitting the generated virtual try-on video to the user's device via a network such as the Internet.

[1324] "Means for extracting dimensional information and generating 3D point cloud data" refers to software and computational processes for measuring the user's body dimensions from the received image data and generating points in 3D space based on that information.

[1325] "Means for constructing a mesh from point cloud data" refers to software and a calculation process for connecting the surface of a three-dimensional shape with triangles or polygons based on the generated point cloud data to generate a continuous mesh.

[1326] "Means for simulating light reflection and clothing stretch" refers to software and computational processes for calculating and reproducing in real time the reflection of light and the dynamic changes that occur as clothing fits the user's body in a virtual environment in which the user is trying on the clothing.

[1327] This invention relates to a system that performs a 3D scan of a user's entire body using a camera such as a smartphone, and generates accurate body shape data from the obtained image data. This system is composed of the user's camera (terminal), a server that analyzes the data, a server that generates 3D models of products, and a server that generates virtual fitting videos and sends them to the terminal.

[1328] 1. (3D scanning on device)

[1329] The user launches the dedicated smartphone app and selects the shooting mode. The app instructs the user to take three poses: front, right side, and back, and the camera takes a photo for each pose. The captured image data is temporarily stored on the device.

[1330] For example, when a user launches the app and selects the photo mode, the instruction "Please stand facing forward" is displayed. When the user stands facing forward, the camera automatically takes a photo, and then the instruction "Please turn to the right" is displayed. This process is repeated to generate 3D scan data.

[1331] 2. (Uploading data to the server and analyzing it)

[1332] The 3D scan data generated by the device is uploaded to a server. The server extracts the user's dimensional information from each image and generates 3D point cloud data. A mesh is then constructed from the point cloud data. This analysis process uses software such as Python's OpenCV and Point Cloud Library (PCL).

[1333] For example, when a user uploads image data they have taken to a server, the server analyzes the pixel information of each photo and calculates the user's height, shoulder width, waist size, etc. Then, it automatically generates 3D point cloud data and meshes based on this information.

[1334] 3. (Importing product data and generating 3D models)

[1335] The server receives product data (photos and dimensional information) provided by the manufacturer and generates a 3D model of the product based on this data. An image processing library (e.g., OpenCV) is used to use the photo as a texture and to recreate the specific shape from the dimensional information.

[1336] Example: The server receives a photo of a jacket provided by a manufacturer and detailed measurement information, and based on this information, generates a 3D model of the jacket, including details such as sleeve length and collar shape.

[1337] 4. (Generation and distribution of virtual try-on videos)

[1338] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video. The scale of the product is adjusted and positioned to fit the user's body. 3D rendering software such as Blender or Unity is also used to simulate light reflection and the stretchiness of the clothing. The final virtual try-on video is sent to the user's device.

[1339] Example: A 3D model of a jacket is adjusted to fit the user's body shape data, and a virtual try-on video is generated. This video shows how the jacket fits perfectly to the user's shoulders and waist. The generated video is sent to the user's device in real time, and the user can view it on their smartphone.

[1340] Examples of prompts:

[1341] "Please generate a video of user A trying on jacket B based on the 3D scan data."

[1342] This invention allows users to perform highly accurate virtual try-on sessions and select clothing that best suits them. It also reduces the burden on logistics and the environmental impact by reducing returns.

[1343] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1344] Step 1:

[1345] The user launches a dedicated app on their smartphone and selects a shooting mode.

[1346] Input: The user operates the app to select a shooting mode.

[1347] Specific operation: The user selects "3D scan mode" and the app launches.

[1348] Output: The app enters shooting mode.

[1349] Step 2:

[1350] The app asks the user to pose from the front, right side, and back.

[1351] Input: The user acts according to the app's instructions.

[1352] Specific behavior: The app displays the instruction "Please stand facing forward," and the user faces forward.

[1353] Output: User facing forward.

[1354] Step 3:

[1355] The device's camera takes a photo for each pose.

[1356] Input: The user strikes a pose and the device camera activates.

[1357] Specific operation: The camera automatically takes photos of the front, right side, and back in sequence.

[1358] Output: Image data from three directions.

[1359] Step 4:

[1360] The captured image data is temporarily stored on the device.

[1361] Input: Image data captured by a camera.

[1362] Specific operation: The captured image data is saved in the smartphone's temporary memory.

[1363] Output: Saved image data.

[1364] Step 5:

[1365] The device uploads the generated 3D scan data to the server.

[1366] Input: Saved image data.

[1367] Specific operation: The device sends image data to the server via an Internet connection.

[1368] Output: Image data uploaded to the server.

[1369] Step 6:

[1370] The server extracts the user's dimensional information from each image and generates 3D point cloud data.

[1371] Input: Image data uploaded to the server.

[1372] How it works: The server uses Python's OpenCV and SciPy libraries to analyze pixel information from each photo and extract measurements such as the user's height, shoulder width, and waist.

[1373] Output: Extracted dimensional information and 3D point cloud data.

[1374] Step 7:

[1375] The server constructs a mesh from the point cloud data.

[1376] Input: Extracted dimensional information and 3D point cloud data.

[1377] Specific operation: The server generates a 3D mesh model using Blender or Point Cloud Library (PCL).

[1378] Output: Reconstructed 3D mesh data.

[1379] Step 8:

[1380] The server receives product data provided by the manufacturer.

[1381] Input: Product data (photos and dimensions) provided by the manufacturer.

[1382] Specific operation: The server receives and stores the product data.

[1383] Output: Received product data.

[1384] Step 9:

[1385] The server generates a 3D model of the product based on the product data.

[1386] Input: Received product data.

[1387] Specific operation: The server uses an image processing library (e.g., OpenCV) to use the photo as a texture and reproduce the specific shape of the product from the dimensional information.

[1388] Output: A 3D model of the generated product.

[1389] Step 10:

[1390] The server combines the user's body data with a 3D model of the product to generate a virtual try-on video.

[1391] Input: User's 3D mesh data and 3D model of the product.

[1392] How it works: The server adjusts the scale of the product and positions it to fit the user's body shape. It also uses Blender or Unity's physics engine to simulate light reflection and clothing stretching.

[1393] Output: Generated virtual try-on video.

[1394] Step 11:

[1395] The server transmits the generated virtual try-on video to the user's terminal.

[1396] Input: Generated virtual try-on footage.

[1397] Specific operation: The server sends the virtual try-on video to the user's device via the Internet.

[1398] Output: A virtual try-on video displayed on the user's device.

[1399] (Application example 1)

[1400] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1401] With traditional online shopping, consumers are unable to actually try on products, resulting in frequent problems such as products not fitting properly or looking different after purchase. This not only increases the return rate, burdens on logistics and the environment, but also reduces consumer satisfaction. Furthermore, existing virtual try-on systems have low accuracy in user body shape data, making it difficult to reproduce an actual fit. To solve these issues, a new system that combines high-precision 3D scanning technology and realistic try-on simulations is needed.

[1402] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1403] In this invention, the server includes: means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data; means for receiving product data and generating a 3D product model from the product data; means for combining the 3D user model and the 3D product model to generate a virtual try-on video; means for transmitting the virtual try-on video to the user device; means for instructing the user to pose when capturing images using the user device's camera or sensor; means for uploading the user's image data to the server and analyzing it to extract detailed dimensional information about the user; means for constructing a 3D model of the user based on the specific dimensional information; means for specifically reproducing the 3D product model based on the product dimensional information; and means for simulating the stretch and shadow of clothing when generating the virtual try-on video. This allows consumers to experience a highly accurate virtual try-on experience and choose the clothing that best suits them. Furthermore, reducing returns can reduce logistics and environmental burdens.

[1404] "Image data taken by a user from multiple directions" refers to a collection of image data taken by a user using a photographing device from different directions such as the front, side, and rear.

[1405] A "3D model" is three-dimensional digital data that reproduces the shape of a user or product from two-dimensional image data.

[1406] "Virtual try-on video" is a video in which a product model is matched to a 3D model of the user, making it appear as if the user is actually trying on the product.

[1407] A "user terminal" is an electronic device that is directly operated by a user, such as a smartphone, tablet, or computer.

[1408] An "imaging device" is a device used to capture images, such as a camera or sensor mounted on a user terminal.

[1409] A "server" is a computer system used for analyzing, storing, and communicating data.

[1410] "Dimensional information" is detailed information about the size of the product and dimensional data of each part of the user's body.

[1411] "Analysis processing" refers to the calculations and processing required to extract necessary information based on received data and generate a model.

[1412] "Detailed dimensional information" refers to the precise dimensions of parts of a user or product, as well as detailed measurement data.

[1413] "Stretching and shadow simulation" is a computer graphics process that recreates the realistic feeling of trying on clothes, including the stretching and shrinking of clothing and the reflection of light.

[1414] The system of the present invention generates a 3D model from image data captured by the user from multiple angles using a camera, and then combines the generated model with a 3D model of the product to provide a virtual try-on video. This system allows users to experience highly accurate virtual try-on from the comfort of their own home, enabling them to choose the clothing that best suits them.

[1415] Hardware and software used

[1416] Device:

[1417] Smartphone (with high-resolution camera and standard IMU (Inertial Measurement Unit))

[1418] server:

[1419] Data analysis server (using TensorFlow and OpenCV)

[1420] 3D rendering server (using Blender)

[1421] Communication (using REST API)

[1422] Data processing and calculation

[1423] Generate a 3D model of the user

[1424] 1. Image capture:

[1425] The user launches the app and takes photos of the front, right side, and back using the device's camera. During this process, the app displays instructions to help the user stand in the correct position and angle.

[1426] 2. Upload image data:

[1427] The captured image data is temporarily stored on the device and then uploaded to the server.

[1428] 3. Analysis process:

[1429] The server analyzes the received image data and extracts detailed dimensional information about the user. Specifically, it uses TensorFlow for image recognition and dimension extraction, OpenCV to generate point cloud data, and Blender to construct the final 3D mesh.

[1430] 3D product model generation

[1431] 1. Product data acquisition:

[1432] The server receives product photos and detailed dimensional information provided by the manufacturer.

[1433] 2. Product model generation:

[1434] Using the required dimensions and photographs, Blender is used to generate a 3D model of the product, including the product's specific shape and texture.

[1435] Virtual try-on video generation and distribution

[1436] 1. Model combination:

[1437] The server combines a 3D model of the user with a 3D model of the product to realistically simulate the fit, simulating the stretch and shadow of the clothing to recreate the feeling of trying it on.

[1438] 2. Image generation:

[1439] The virtual try-on video generated by the above process is sent to the user's device, where the user can view the video and choose the clothing that best suits them.

[1440] Specific examples

[1441] Image capture and upload

[1442] The user opens the app and takes a frontal image following the instruction "Please stand facing forward," then takes a right-side image following the instruction "Please stand facing right," and then takes a backside image following the instruction "Please stand with your back to the camera." This series of images is then uploaded to the server.

[1443] Server-side processing

[1444] The server analyzes the received image data and extracts detailed dimensional information such as the user's height, shoulder width, waist, etc. Based on the extracted dimensional information, point cloud data is generated using TensorFlow and OpenCV, and a 3D mesh is constructed using Blender.

[1445] 3D product model generation

[1446] The server receives jacket photos and measurements provided by the manufacturer and uses Blender to generate a 3D model of the jacket, including sleeve length, collar shape, texture information, and more.

[1447] Virtual try-on video generation

[1448] The generated user model is combined with the jacket model to simulate the fit, light reflection, and stretch of the clothing, generating a realistic video of the user trying on the garment. This video is then sent to the user's device.

[1449] Specific prompt examples

[1450] "Please explain the overview of the system that enables highly accurate virtual try-on based on a full-body 3D scan."

[1451] The details of this system depend on the user's environment, providing a convenient experience that can be experienced individually by each user. This will improve the online shopping experience for consumers, reduce the return rate, and ease the burden on logistics.

[1452] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1453] Step 1:

[1454] Image capture

[1455] Input: Image data from multiple angles according to user instructions

[1456] How it works: The user launches the dedicated app and takes photos of three poses using the device's camera: the front, right side, and back.

[1457] Output: Three images taken by the user (front, right, and back images)

[1458] Step 2:

[1459] Image data upload

[1460] Input: User image data captured in step 1

[1461] Operation: The device uploads the image data it has taken to the server. The app uses the REST API to send the image data to the server.

[1462] Output: User image data transferred to the server

[1463] Step 3:

[1464] Analysis processing

[1465] Input: User image data uploaded to the server

[1466] How it works: The server uses TensorFlow and OpenCV to extract the user's dimensions from each image, generates point cloud data from each image, and builds a 3D mesh using Blender.

[1467] Output: Detailed dimensional information and your 3D model

[1468] Step 4:

[1469] Get product data

[1470] Input: Product photos and dimensions provided by the manufacturer

[1471] How it works: The server receives product data, generates a 3D model based on the necessary dimensions and photos, and uses Blender to incorporate product shape and texture information into the 3D model.

[1472] Output: 3D model of the product

[1473] Step 5:

[1474] Virtual try-on video generation

[1475] Input: 3D model of the user and 3D model of the product

[1476] How it works: The server combines the user model with the product model, simulates fit, stretching, and shadows, and uses Blender to generate a virtual try-on video.

[1477] Output: Virtual try-on video data

[1478] Step 6:

[1479] Distribution of virtual try-on videos

[1480] Input: Generated virtual try-on video data

[1481] Operation: The server generates a virtual fitting video and sends it to the user's device. The user can view the video on their device and check how the clothes feel when they are tried on.

[1482] Output: Virtual try-on video that can be viewed by the user

[1483] In this way, each processing step of the system is combined to provide the user with a highly accurate virtual try-on experience.

[1484] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1485] This invention relates to a system that provides a virtual try-on experience by 3D scanning a user's entire body with a camera such as a smartphone, generating accurate body shape data from the obtained image data, and combining this with an emotion engine. The system consists of a user's device, a server that analyzes the data, a server that generates 3D product models, the emotion engine, and a server that generates virtual try-on videos and sends them to the user's device.

[1486] Program processing overview

[1487] 1. 3D scanning on your device

[1488] The user launches a dedicated smartphone app and takes a full-body photo. The app then asks the user to pose in three ways: from the front, right side, and back, and takes an image of each pose. The captured images are temporarily saved on the device. This process generates 3D scan data of the user.

[1489] 2. Uploading data to the server and processing it

[1490] The generated 3D scan data is uploaded from the user's device to a server, which analyzes the received data and generates accurate data on the user's body shape. The analysis process involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[1491] 3. Product data capture and generation

[1492] The server receives product data provided by the manufacturer, including photos of the garments and detailed measurement information, and generates a 3D model of the product based on this data. Specifically, the server uses the photos as textures and recreates the shape of the product based on the measurement information.

[1493] 4. Incorporating an Emotional Engine

[1494] While the user is trying on the clothes, the device's built-in camera and microphone analyze the user's facial expressions and tone of voice in real time, and the emotion engine recognizes the user's emotions. For example, if the system detects that the user is smiling, it will determine that the user has a positive feeling toward the clothing.

[1495] 5. Creation and distribution of virtual try-on videos

[1496] The server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, it uses the user's emotional data analyzed by the emotion engine to adjust the content of the virtual try-on video. For example, if the user's emotion is unfavorable, it can generate a scenario that suggests a different product that fits better.

[1497] 6. Video distribution and user feedback

[1498] The generated virtual try-on video is sent to the user's device for viewing. The user's try-on experience is then analyzed for emotion, and feedback is provided based on the user's emotions. For example, if the user is dissatisfied, the system automatically suggests other products.

[1499] Specific examples

[1500] 3D scanning on your device

[1501] The user launches the dedicated app and selects the shooting mode. When the app instructs them to "stand facing forward," the user stands facing forward. The device's camera takes a frontal photo, then the app displays "Turn to your right," and the user stands facing right. The camera then takes a right-side photo, then the app displays "Turn your back," and the app takes a rear-side photo. At this point, 3D scan data of the user is generated.

[1502] Data upload to server and processing

[1503] The 3D scan data taken by the user's device is uploaded to the server, which then analyzes the data and generates accurate body shape data for the user. For example, detailed dimensional information such as the user's height, shoulder width, and waist size is extracted and reproduced as a 3D model.

[1504] Product data ingestion and generation

[1505] The server generates a 3D model of the product based on the product photos and measurements provided by the manufacturer. The server uses the provided jacket photos and measurements to generate a 3D model of the jacket, including details such as sleeve length and collar shape.

[1506] Incorporating an emotion engine

[1507] When a user tries on a jacket during the fitting experience, the device's camera analyzes the user's facial expressions and the emotion engine recognizes the user's emotions. For example, if the user is smiling, the emotion engine determines that the user likes the jacket.

[1508] Creation and distribution of virtual try-on videos

[1509] The server combines the user's body shape data with a 3D model of the product to generate a virtual try-on video. The emotion engine adjusts the video content based on the user's emotional data. For example, if the user is smiling, different colors and styles of jackets will be suggested in the video.

[1510] Video distribution and user feedback

[1511] The generated virtual try-on video is sent to the user's device, where the user reviews it. While reviewing, the emotion engine analyzes the user's facial expressions again to determine their level of satisfaction. If their satisfaction is low, the system suggests a different product. If their satisfaction is high, they can proceed with the purchase process.

[1512] This not only allows users to have a highly accurate virtual try-on experience, but also allows them to receive more personalized product recommendations based on their emotions, resulting in increased consumer satisfaction and reduced returns.

[1513] The processing flow will be explained below.

[1514] Step 1:

[1515] The user launches the dedicated app and selects 3D scan mode. The app then prompts the user to take photos from the front, right side, and back in that order.

[1516] Step 2:

[1517] The user stands facing forward, and the device camera takes a full-body frontal photograph of the user. The captured image is immediately saved temporarily on the device.

[1518] Step 3:

[1519] The user stands facing right, and the device camera takes a photo of the user's entire body from the right side. As with the front-facing photo, this image is also temporarily saved on the device.

[1520] Step 4:

[1521] The user stands with their back to the camera and takes a photo of the user's entire body from the back. The image of the user's back is also temporarily stored on the device.

[1522] Step 5:

[1523] The device compiles the 3D scan data of the front, side, and back and sends a request to upload it to the server.

[1524] Step 6:

[1525] The server analyzes the image data received from the device and generates the user's body shape data. Specifically, it extracts dimensional information such as the user's height, shoulder width, and waist from each image and generates a 3D mesh model based on this information.

[1526] Step 7:

[1527] The server receives product data provided by the manufacturer, including photos of the garment and dimensions of each part, and generates a 3D model of the product based on this data.

[1528] Step 8:

[1529] The device's camera and microphone capture the user's facial expressions and tone of voice in real time and send the data to the emotion engine, which analyzes this data and recognizes the user's emotions.

[1530] Step 9:

[1531] The emotion engine sends the user's emotion data to the server, which then takes this emotion data into account when generating the virtual try-on video. For example, if the user expresses positive emotion, the video can include content suggesting different variations of the product.

[1532] Step 10:

[1533] The server combines the user's body data with a 3D model of the product to generate a virtual fitting video. The simulation includes the stretching and light reflection of the clothing, providing an experience as close as possible to a real try-on.

[1534] Step 11:

[1535] The generated virtual fitting video is sent to the user's device, where the user can view it. The device continues to send the user's facial expressions to the emotion engine while the video is playing, and analyzes their emotions in real time.

[1536] Step 12:

[1537] The server provides feedback based on the user's emotional data. If the user is not satisfied, the server generates content suggesting alternative products and adjusts it to attract the user's interest.

[1538] Step 13:

[1539] If the user is satisfied, the app will either proceed with the purchase or offer further suggestions, and the entire system will adjust product selection accordingly.

[1540] This not only allows users to have a highly accurate virtual try-on experience, but also allows for more personalized product recommendations based on their emotions, which in turn increases consumer satisfaction and reduces returns.

[1541] Example 2

[1542] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1543] Conventional virtual try-on systems have low accuracy in generating 3D shape models of users and have difficulty in proposing products that appropriately reflect the user's emotions. As a result, user satisfaction declines and ultimately, product returns increase. Furthermore, there is a lack of technology to analyze the user's emotional state in real time and provide feedback to the virtual try-on video.

[1544] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a three-dimensional shape model of the user from this image data, means for receiving product information and generating a three-dimensional shape model of the product from this product information, means for combining the user's three-dimensional shape model and the product's three-dimensional shape model to generate a virtual try-on video, means for analyzing the user's emotional state using an emotion analysis engine and adjusting the content of the virtual try-on video, and means for transmitting the virtual try-on video to the user terminal. This improves the accuracy of the user's three-dimensional shape model, making it possible to recommend products that reflect the user's emotions, which is expected to improve user satisfaction and reduce returned products.

[1545] "Image data taken by a user from multiple directions" refers to image data of the user taken from different directions using a photographing device owned by the user.

[1546] A "three-dimensional shape model" is a model that represents a shape in three-dimensional space and is generated from image data of a user or a product.

[1547] "Product information" refers to data provided by manufacturers and providers, including product photos and detailed dimensional information.

[1548] An "emotion analysis engine" is software or hardware that can analyze a user's facial expressions and tone of voice to recognize and determine the user's emotional state.

[1549] A "virtual try-on video" is a video that simulates the experience of trying on clothes, generated by combining a three-dimensional shape model of the user and a three-dimensional shape model of the product.

[1550] "User terminal" refers to an electronic device used by a user to display and operate information, such as a smartphone, tablet, or PC.

[1551] This system provides a virtual try-on experience by 3D scanning a user's entire body with a camera, generating accurate body shape data from the image data, and combining this with an emotion analysis engine. This system is comprised of a user's device, a server that analyzes the data, a server that generates a 3D product model, an emotion analysis engine, and a server that generates a virtual try-on video and sends it to the user's device.

[1552] The process begins when the user launches a dedicated smartphone app and takes a full-body photo. The app displays prompts such as "Please stand facing forward," "Please turn to your right," and "Please stand with your back to the camera," and takes images from each direction based on the user's current posture. After the photos are taken, the image data is temporarily stored on the device.

[1553] The user's device then uploads the stored 3D scan data to a server, which analyzes the data to generate accurate body shape data for the user. The analysis process uses computer vision techniques and machine learning algorithms to extract the user's dimensions from each image and create a 3D mesh model based on that information.

[1554] The server also receives product information from manufacturers, including product photos and detailed measurements. The server then generates a 3D model of the product based on this data. Specifically, the server uses clothing photos as textures and reproduces the product shape using measurements.

[1555] During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time. An emotion analysis engine recognizes the user's emotions and adjusts the content of the fitting video based on that information. For example, if the user is smiling, the system will determine that they have a positive attitude toward the product and reflect this in the video. If the user's emotions are not positive, the system can also generate a scenario that suggests alternative products.

[1556] The generated virtual try-on video is sent to the user's device, where the user reviews it. During the review, sentiment analysis is performed again and the user's feedback is provided. If the user's satisfaction level is low, the system automatically suggests other products. This allows the user to have a highly accurate virtual try-on experience and also receive more personalized product suggestions based on their emotions.

[1557] Example prompt sentence:

[1558] "Please stand facing forward."

[1559] "Please turn to your right."

[1560] "Please stand with your back to me."

[1561] The above is a specific embodiment based on the present invention for providing a highly accurate virtual try-on experience for the user.

[1562] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1563] Processing flow

[1564] Step 1: 3D scan on your device

[1565] 1.1 The user launches the dedicated app and selects the shooting mode.

[1566] Input: User actions

[1567] Output: Start camera, display prompt

[1568] Specific operation: The user opens the dedicated smartphone app and selects the shooting mode. The app then displays the message, "Please stand facing forward."

[1569] 1.2 The app asks the user to pose and the camera takes the photo.

[1570] Input: current user posture, device camera input

[1571] Output: Image data of the front, right side, and back of the user

[1572] Specific operation: The user follows the instructions of the app to face the front, right side, and back, and the device camera takes a photo of each. The captured image data is temporarily stored on the device.

[1573] Step 2: Upload data to the server and process it

[1574] 2.1 The device uploads data to the server.

[1575] Input: Image data stored on the device

[1576] Output: Uploaded image data

[1577] How it works: Once the user has taken a photo, the app sends the image data to a server, where it is uploaded using a secure protocol.

[1578] 2.2 The server receives and analyzes the data.

[1579] Input: Uploaded image data

[1580] Output: 3D model of the user

[1581] How it works: The server reviews the received image data and analyzes it using computer vision techniques and machine learning algorithms. It extracts the user's dimensions from each image and creates a 3D mesh model based on them.

[1582] Step 3: Import and generate product data

[1583] 3.1 The server receives product information from the manufacturer.

[1584] Input: Product photos and dimensions provided by the manufacturer

[1585] Output: Received product information

[1586] Specific operation: The server receives product information provided by the manufacturer and stores it in a database.

[1587] 3.2 The server generates a 3D model of the product.

[1588] Input: Product information

[1589] Output: 3D shape model of the product

[1590] Specific operation: The server generates a 3D model of the product based on the product information. Specifically, it uses a photo of the clothing as texture and recreates the shape of the product using measurement information.

[1591] Step 4: Incorporating the Emotion Engine

[1592] 4.1 The device collects the user's emotional data.

[1593] Input: User's facial expression data, voice data

[1594] Output: Collected emotion data

[1595] Specific operation: During the fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice in real time to collect emotional data.

[1596] 4.2 The server analyzes the user's emotions using the emotion engine.

[1597] Input: Collected emotion data

[1598] Output: Parsed emotion information

[1599] How it works: The emotion analysis engine recognizes and analyzes the user's emotions based on the collected data. For example, it detects if the user is smiling and provides feedback based on that situation.

[1600] Step 5: Generate and distribute virtual try-on footage

[1601] 5.1 The server generates the virtual try-on video.

[1602] Input: 3D model of the user, 3D model of the product, analyzed emotion information

[1603] Output: Virtual try-on video

[1604] Specific operation: The server combines the 3D model of the user and the 3D model of the product to generate a virtual try-on video that reflects the user's emotional data.

[1605] 5.2 The server transmits the generated video to the user terminal.

[1606] Input: Virtual try-on video

[1607] Output: Video distribution to user devices

[1608] Specific operation: The generated virtual try-on video is sent to the user's device, allowing the user to view the video on the device.

[1609] Step 6: Video distribution and user feedback

[1610] 6.1 The device re-evaluates the user.

[1611] Input: facial expression data and voice data of the user watching the virtual try-on video

[1612] Output: Re-collected emotion data

[1613] Specific operation: While the user is viewing the virtual fitting video, the device's camera and microphone again capture the user's facial expressions and reactions.

[1614] 6.2 The server makes suggestions based on the feedback.

[1615] Input: Recollected emotion data

[1616] Output: New product proposals for users

[1617] Specific operation: The server analyzes the collected user emotion data again and generates a scenario to suggest a different product if the user's satisfaction is low. If the user's satisfaction is high, the system proceeds to the purchase procedure.

[1618] (Application example 2)

[1619] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1620] Conventional virtual try-on systems can generate try-on videos by combining a user's body data with a 3D product model, but they face challenges in providing highly personalized product recommendations that take into account the user's emotions and feedback, as well as improving the user experience. Furthermore, there is a need for an effective method to reduce product return rates while improving user satisfaction.

[1621] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1622] In this invention, the server includes means for receiving image data captured by the user from multiple angles and generating a 3D model of the user from the image data, means for receiving product data and generating a 3D product model from the product data, means for combining the 3D user model and the 3D product model to generate a virtual try-on video, means for transmitting the virtual try-on video to the user terminal, means for analyzing the user's facial expression and tone of voice and extracting emotion data, means for dynamically adjusting the virtual try-on video based on the extracted emotion data, and means for making product suggestions based on the generated virtual try-on video in accordance with the user's emotions. This enables real-time analysis of user emotions and feedback and personalized product suggestions based on the analysis.

[1623] "User" refers to a person who operates a system, such as a customer or a user.

[1624] "Photography device" refers to a device used to acquire image data, such as a camera or smartphone.

[1625] "Image data" refers to data of photographs or videos captured by a photographing device.

[1626] "3D model" refers to a geometric model of digital data expressed in three dimensions.

[1627] "Product Data" refers to data containing product details, such as product dimensions and photos.

[1628] "Virtual try-on video" refers to a try-on simulation video generated by combining a 3D model of the user and a 3D model of the product.

[1629] "User terminal" refers to an information terminal operated by a user, such as a smartphone or tablet.

[1630] "Facial expression" refers to information that indicates emotions expressed through the movement of facial muscles.

[1631] "Tone of voice" refers to audio information that indicates the pitch, strength, and emotional nuances of a speaking voice.

[1632] "Emotion data" refers to data that indicates the user's emotions analyzed from facial expressions, tone of voice, etc.

[1633] "Dynamic adjustment" refers to the process of changing the video content in response to changing conditions in real time.

[1634] "Personalized product suggestions" refer to product suggestions recommended based on the individual preferences and feelings of each user.

[1635] The system for realizing this invention comprises a user terminal, a server for analyzing data, and a server for managing product data. It also includes hardware and software that are equipped with an emotion engine and can analyze user emotions. A specific embodiment of the system will be described below.

[1636] The process begins when the user launches a dedicated smartphone application and takes a full-body photo. The application displays instructions to the user, such as "Please face forward," "Please face your right side," and "Please turn your back," and prompts the user to take photos in each pose. This process generates 3D scan data of the user.

[1637] The captured 3D scan data is uploaded from the user's device to an analysis server, which then uses the data to generate accurate body shape data for the user. This analysis involves extracting the user's dimensional information from each image and creating a 3D mesh model based on that information.

[1638] Next, the product data management server receives the product data, which includes product photos and detailed measurement information. The product data server uses this data to generate a 3D model of the product. In this process, the photos are used as textures, and the shape of the product is reproduced based on the measurement information. The 3D model is structured so that the movement of each part can be flexibly reproduced.

[1639] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice in real time. During the user's fitting experience, the device's camera and microphone analyze the user's facial expressions and tone of voice, and the emotion engine extracts the user's emotional data. For example, if the user is smiling, the system determines that the user has a positive feeling toward the clothing.

[1640] The analysis server combines the generated user's body shape data with a 3D product model to generate a virtual try-on video. Furthermore, the content of the virtual try-on video can be dynamically adjusted based on the user's emotional data analyzed by the emotion engine. For example, if the user's emotion is unfavorable, the system can generate a scenario that suggests a different product that fits better.

[1641] The generated virtual try-on video is sent to the user's device, where the user can review the results. During this review process, the device's camera again analyzes the user's facial expressions and extracts emotional data. If the user is not satisfied, the system suggests other products, but if satisfied, the user can proceed with the purchase.

[1642] Specific examples

[1643] For example, a user can launch a dedicated smartphone app and follow the instructions to take a photo facing forward, then facing right, and finally facing away from the user, generating 3D scan data. The captured data is then uploaded to a server, where it is analyzed and detailed body shape data is generated.

[1644] While the user is trying on a jacket, the device's camera analyzes the user's facial expressions, and the emotion engine dynamically suggests products based on prompts such as, "If the user is smiling, suggest different color and style variations of that jacket." For example, when a user tries on a jacket, the emotion engine can suggest other products or customization options based on prompts such as, "If the user is not satisfied, suggest other products or customization options."

[1645] This provides users with a highly accurate and personalized virtual try-on experience, reducing return rates while increasing consumer satisfaction.

[1646] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1647] Step 1:

[1648] The user launches a dedicated application on their smartphone and takes a full-body photo. The user follows the application's instructions to pose and take photos from the front, right side, and back. Image data from multiple angles of the user is obtained as input. As output, this data is generated as 3D scan data. In this step, each image is captured and temporarily saved on the device.

[1649] Step 2:

[1650] The user terminal uploads the generated 3D scan data to the server. As input, the 3D scan data stored on the terminal is used. As output, the server stores the received image data. In this step, data is transferred over the network.

[1651] Step 3:

[1652] The server analyzes the received 3D scan data and generates accurate body shape data for the user. The 3D scan data is used as input. As output, a 3D mesh model of the user's body shape is generated. In this step, image processing techniques are used to extract dimensional information and create a 3D model based on that data.

[1653] Step 4:

[1654] The product data management server receives the product data. As input, it receives a product photo and detailed measurement information. As output, it generates a 3D model of the product. In this step, it uses the photo as a texture and processes the product shape based on the measurement information.

[1655] Step 5:

[1656] The user device is equipped with an emotion engine that analyzes facial expressions and tone of voice. When the user tries on the clothes, the device's camera and microphone capture the user's facial expressions and voice to extract emotion data. Real-time video and audio data are used as input. The output is the user's emotion data. In this step, the emotion analysis algorithm is executed.

[1657] Step 6:

[1658] The analysis server combines the generated user's body shape data with the 3D model of the product to generate a virtual try-on video. The 3D mesh model of the user and the 3D model of the product are used as input. The output is a virtual try-on video. In this step, 3D rendering technology is used to create a visually realistic try-on video.

[1659] Step 7:

[1660] The content of the virtual try-on video is dynamically adjusted based on the user's emotional data analyzed by the emotion engine. The user's emotional data is used as input. The output is an adjusted try-on video based on the user's emotions. In this step, the system changes the video content taking the user's reactions into account.

[1661] Step 8:

[1662] The generated virtual try-on video is sent to the user's device. The adjusted virtual try-on video is used as input. The video is played on the user's device as output. In this step, the data transfer and video playback processes are executed.

[1663] Step 9:

[1664] When the user checks the fitting results, the device's camera again analyzes the user's facial expressions and extracts emotional data. The input is a video of the user watching the fitting video. The output is again emotional data. In this step, emotional analysis is performed to evaluate the user's satisfaction.

[1665] Step 10:

[1666] If the satisfaction level is low, the system will suggest another product. If the satisfaction level is high, the user can proceed to the purchase process. As input, the user's emotional data is used. As output, a product suggestion and purchase process scenario is generated. In this step, the personalized product suggestion and purchase process are executed.

[1667] Through these steps, users can enjoy a highly accurate and personalized virtual try-on experience, resulting in increased satisfaction and reduced return rates.

[1668] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1669] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1670] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1671] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1672] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1673] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1674] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1675] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1676] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1677] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1678] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1679] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1680] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1681] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1682] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1683] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1684] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1685] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1686] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1687] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1688] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1689] The following is further disclosed regarding the above embodiment.

[1690] (Claim 1)

[1691] means for receiving image data captured by a user from multiple directions and generating a three-dimensional model of the user from the image data;

[1692] means for receiving product data and generating a three-dimensional model of the product from the product data;

[1693] A means for generating a virtual try-on video by combining a 3D model of the user and a 3D model of the product;

[1694] The system includes a means for transmitting a virtual try-on video to a user terminal.

[1695] (Claim 2)

[1696] 10. The system according to claim 1, wherein image data of a user from multiple directions is received from an image capture device carried by the user.

[1697] (Claim 3)

[1698] 2. The system according to claim 1, further comprising means for generating a three-dimensional model of the product based on data including dimensional information of the product.

[1699] (Claim 4)

[1700] 2. The system according to claim 1, further comprising means for generating, in the generated virtual try-on video, a video of multiple products being tried on at the same time.

[1701] (Claim 5)

[1702] 2. The system according to claim 1, further comprising means for displaying in real time on the screen of the user's terminal a display content in which a three-dimensional model of the product is combined with a three-dimensional model of the user.

[1703] (Claim 6)

[1704] 2. The system according to claim 1, further comprising means for storing the generated three-dimensional model of the user on a server, and generating and distributing a virtual try-on video of the product to a plurality of users.

[1705] (Claim 7)

[1706] 2. The system according to claim 1, further comprising means for simulating light reflection and stretching of clothing in real time in the generated virtual fitting video.

[1707] "Example 1"

[1708] (Claim 1)

[1709] means for receiving image data captured by a user from multiple directions and generating a three-dimensional model of the user from the image data;

[1710] means for receiving product data and generating a three-dimensional model of the product from the product data;

[1711] A means for generating a virtual try-on video by combining a 3D model of the user and a 3D model of the product;

[1712] means for transmitting the virtual try-on video to a user terminal;

[1713] A means for extracting user dimensional information from each image to generate three-dimensional point cloud data;

[1714] a means for constructing a mesh from the point cloud data;

[1715] A means to simulate light reflection and the stretching of clothing,

[1716] A system including:

[1717] (Claim 2)

[1718] 10. The system according to claim 1, wherein image data of a user from multiple directions is received from an image capture device carried by the user.

[1719] (Claim 3)

[1720] 2. The system according to claim 1, further comprising means for generating a three-dimensional model of the product based on data including dimensional information of the product.

[1721] "Application Example 1"

[1722] (Claim 1)

[1723] means for receiving image data captured by a user from multiple directions and generating a three-dimensional model of the user from the image data;

[1724] means for receiving product data and generating a three-dimensional model of the product from the product data;

[1725] A means for generating a virtual try-on video by combining a 3D model of the user and a 3D model of the product;

[1726] means for transmitting the virtual try-on video to a user terminal;

[1727] means for instructing a user to pose when taking an image using a camera or sensor of the user terminal;

[1728] A means for uploading user image data to a server and extracting detailed dimensional information of the user through analysis processing;

[1729] A means for constructing a 3D model of the user based on specific dimensional information;

[1730] A means for specifically reproducing a three-dimensional model of the product based on the product's dimensional information;

[1731] A method for simulating the stretching and shadowing of clothing when generating virtual fitting footage

[1732] A system including:

[1733] (Claim 2)

[1734] 2. The system according to claim 1, wherein image data of a user from multiple directions is received from a photographing device carried by the user, and the system instructs the user to pose when photographing.

[1735] (Claim 3)

[1736] The system of claim 1 generates a three-dimensional model of the product based on the product's dimensional information and detailed specifications, and reproduces a realistic fit.

[1737] "Example 2: Combining Emotion Engines"

[1738] (Claim 1)

[1739] means for receiving image data captured by a user from multiple directions and generating a three-dimensional shape model of the user from the image data;

[1740] means for receiving product information and generating a three-dimensional shape model of the product from the product information;

[1741] a means for generating a virtual try-on video by combining a three-dimensional shape model of the user and a three-dimensional shape model of the product;

[1742] a means for analyzing the emotional state of a user using an emotion analysis engine and adjusting the content of the virtual try-on video;

[1743] The system includes a means for transmitting a virtual try-on video to a user terminal.

[1744] (Claim 2)

[1745] 10. The system according to claim 1, wherein image data of a user from multiple directions is received from an image capture device carried by the user.

[1746] (Claim 3)

[1747] 2. The system according to claim 1, further comprising means for generating a three-dimensional shape model of the product based on data including dimensional information of the product.

[1748] "Application example 2 when combining emotion engines"

[1749] (Claim 1)

[1750] means for receiving image data captured by a user from multiple directions and generating a three-dimensional model of the user from the image data;

[1751] means for receiving product data and generating a three-dimensional model of the product from the product data;

[1752] A means for generating a virtual try-on video by combining a 3D model of the user and a 3D model of the product;

[1753] means for transmitting the virtual try-on video to a user terminal;

[1754] A means for analyzing a user's facial expression and tone of voice to extract emotional data;

[1755] means for dynamically adjusting the virtual fitting video based on the extracted emotion data;

[1756] A means of suggesting products based on the user's emotions based on the generated virtual try-on video

[1757] A system including:

[1758] (Claim 2)

[1759] 10. The system according to claim 1, wherein image data of a user from multiple directions is received from an image capture device carried by the user.

[1760] (Claim 3)

[1761] 2. The system according to claim 1, further comprising means for generating a three-dimensional model of the product based on data including dimensional information of the product. [Explanation of symbols]

[1762] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving image data captured by a user from multiple directions and generating a three-dimensional model of the user from the image data; means for receiving product data and generating a three-dimensional model of the product from the product data; A means for generating a virtual try-on video by combining a 3D model of the user and a 3D model of the product; The system includes a means for transmitting a virtual try-on video to a user terminal.

2. 2. The system according to claim 1, wherein image data of a user from a plurality of directions is received from a photographing device carried by the user.

3. 2. The system according to claim 1, further comprising means for generating a three-dimensional model of the product based on data including dimensional information of the product.

4. 2. The system according to claim 1, further comprising means for generating an image of a virtual try-on video in which a plurality of products are tried on at the same time.

5. 2. The system according to claim 1, further comprising means for displaying in real time on the screen of the user's terminal a display content in which a three-dimensional model of the product is combined with the user's three-dimensional model.

6. 2. The system according to claim 1, further comprising means for storing the generated three-dimensional model of the user on a server, and generating and distributing a virtual try-on video of the product to a plurality of users.

7. 2. The system according to claim 1, further comprising means for simulating light reflection and stretching of the clothing in real time in the generated virtual fitting video.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A