System

The system addresses the challenge of conveying product size and fit through three-dimensional silhouette analysis and generative AI, allowing retailers to create effective promotional videos without specialized skills, enhancing consumer confidence and sales.

JP2026017299APending Publication Date: 2026-02-04SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024118081
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-02-04

AI Technical Summary

Technical Problem

Small and medium-sized retailers face challenges in accurately conveying the size and fit of fashion products, as flat images are inadequate, and they lack the resources to produce high-quality product videos, hindering consumer purchases.

Method used

A system that allows users to upload product photos, analyze them to generate a three-dimensional silhouette, and use generative AI models to create videos of virtual models wearing the products, enabling users to select the most suitable video without requiring technical skills.

Benefits of technology

Enables retailers to create high-quality product promotional videos, reducing consumer anxiety about size and fit, and increasing sales by providing detailed product views.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026017299000001_ABST
    Figure 2026017299000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving and storing a commodity picture; means for analyzing the received commodity picture to generate a three dimensional silhouette of the commodity; means for generating a moving image in which a AI model wears the commodity using a generative virtual model based on the generated three dimensional silhouette; means for presenting a plurality of generated moving images to a user; and means for storing and providing a moving image selected by the user as a final high-resolution moving image.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When selling fashion products, flat images often make it difficult to accurately convey the size and fit of the product, which often makes consumers hesitant to purchase. Small and medium-sized retailers, in particular, are at a disadvantage when it comes to promoting purchases because they have limited skills and budgets to produce high-quality product videos. There is also a need for a simple and efficient way to actually show how a product looks on multiple body types. [Means for solving the problem]

[0005] To solve this problem, the present invention provides the following system. First, a user uploads a product photo to the system, which is received and stored by a server. Next, the server analyzes the received product photo and generates a three-dimensional silhouette of the product. Based on the generated three-dimensional silhouette, the server uses a generative AI model to generate a video of a virtual model wearing the product. This takes into account the model's physique and background information specified by the user. The server presents multiple generated videos to the user, and the user selects the most appropriate one. Finally, the server saves the selected video as a final high-resolution video and provides it to the user. This allows users to easily create and publish videos of models wearing products without requiring technical knowledge or skills, contributing to promoting purchases on e-commerce sites.

[0006] A "product photo" is a photograph of a flat image showing the appearance of a product.

[0007] A "three-dimensional silhouette" is a model that represents the shape and size of a product in three-dimensional space.

[0008] A "generative AI model" is an algorithm that uses machine learning technology to generate new data, images, and videos based on given data.

[0009] A "virtual model" is a digital representation of a person or object created using computer graphics techniques.

[0010] "Video" is a media format that expresses movement by playing a series of image frames.

[0011] "User" means a person or entity that uses the System to upload product photos and generate or select videos.

[0012] A "server" is a central computer system that processes and manages data for a system.

[0013] "High-definition video" means video with detailed image quality and providing high visual clarity.

[0014] A "system" is a collection of hardware and software combined to accomplish a particular purpose. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] The system of the present invention is a technology that allows users to upload product photos and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. The program of the system of the present invention is explained below in natural language.

[0037] System program processing explanation

[0038] 1. Upload product photos

[0039] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[0040] The terminal transmits the selected photo to the server.

[0041] The server stores the received photos and organizes them into folders for analysis.

[0042] 2. Creating a three-dimensional silhouette of a product

[0043] The server analyzes the received product photos and runs image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[0044] The server uses 3D modeling technology, such as Structure-from-Motion technology, to generate a three-dimensional silhouette of the product using the identified product shape information.

[0045] 3. Specifying the model and background

[0046] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[0047] The terminal transmits this specification information to the server.

[0048] 4. Video creation using generative AI models

[0049] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[0050] The server generates an interface for suggesting the generated multiple videos to the user.

[0051] 5. Select a video

[0052] The user reviews the suggested videos and selects the most suitable one.

[0053] The terminal transmits the user's selection information to the server.

[0054] 6. Generate and save the final video

[0055] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[0056] The server stores the generated high-resolution video therein so that the user can download it, and provides it to the user.

[0057] Specific examples

[0058] Example 1: Creating a shirt video

[0059] 1. A user (retailer) uploads five photos of a new shirt to the system, including photos from the front, back, left and right, and one-quarter angles.

[0060] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[0061] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[0062] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[0063] 5. The user selects the most suitable video from the suggested videos.

[0064] 6. The server stores the selected video in high resolution and makes it available for download by the user.

[0065] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills, thereby effectively supporting the promotion of fashion merchandise purchases.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[0069] Step 2:

[0070] The terminal transmits the selected photo to the server.

[0071] Step 3:

[0072] The server stores the received photo files and organizes them for further processing.

[0073] Step 4:

[0074] The server analyzes the stored product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[0075] Step 5:

[0076] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[0077] Step 6:

[0078] The user uses the system interface to specify the model's physical information (e.g., height, weight) and background information.

[0079] Step 7:

[0080] The terminal transmits the information specified by the user (physique information and background information of the model) to the server.

[0081] Step 8:

[0082] The server integrates the 3D silhouette with the user-specified information, uses a generative AI model to dress the virtual model in the product, and generates a video against the specified background, using generative technologies such as GANs and NeRF.

[0083] Step 9:

[0084] The server generates an interface for suggesting the generated multiple videos to the user.

[0085] Step 10:

[0086] The terminal displays the suggested videos to the user.

[0087] Step 11:

[0088] The user selects the most suitable video from the displayed videos.

[0089] Step 12:

[0090] The terminal transmits the user's selection information to the server.

[0091] Step 13:

[0092] The server processes the user's selected video in high resolution and generates the final video file.

[0093] Step 14:

[0094] The server stores and provides the generated high-resolution video for users to download.

[0095] Step 15:

[0096] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] Traditional online shopping has the problem that customers cannot actually try on products, which can cause anxiety about the size and fit of the product. Furthermore, product photos and simple descriptions alone make it difficult to fully check the product's details and actual appearance, which can lead to hesitation in purchasing. Furthermore, the need for specialized knowledge and skills to take product photos and create videos places a significant burden on product sellers.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes a means for receiving and saving product photos, a means for analyzing the photos to generate a three-dimensional silhouette of the product, a means for generating a video of a virtual model wearing the product using a generative AI model based on the generated silhouette, a means for presenting multiple videos to the user, and a means for saving and providing the video selected by the user as a final high-resolution video. This allows users to check the appearance and fit of products in detail without actually trying them on, eliminating any anxiety about purchasing. Furthermore, product sellers can create high-quality product promotional videos without specialized knowledge of photography or editing, leading to increased sales.

[0102] "Product photos" are image data of products uploaded by users.

[0103] "Means for storing" refers to the system's function of storing received product photos in a database or storage.

[0104] The "means for analyzing" refers to a function of the system that executes an image processing algorithm to identify the contours and shape of the product using a product photograph and generate a three-dimensional silhouette.

[0105] A "three-dimensional silhouette" is three-dimensional shape information of a product generated from a product photo.

[0106] A "generative AI model" is an algorithm that uses deep learning technology to generate new data based on input data, and in this invention is used to generate videos of virtual models.

[0107] A "virtual model" is a model of a person or object that is digitally generated using a generative AI model.

[0108] The "means for generating video" is a system function that creates video data based on the generated three-dimensional silhouette and virtual model.

[0109] The "presentation means" is a system function that displays the generated multiple videos so that the user can check them.

[0110] "High-resolution video" refers to high-quality video data that can display every detail clearly.

[0111] The system of the present invention is a technology that allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Specific embodiments of the system of the present invention are described below.

[0112] 1. Upload product photos:

[0113] Users take photos of the items they want to sell using a digital camera or smartphone.

[0114] The user logs into the system's web interface and selects a product photo in the upload window.

[0115] The terminal transmits the selected photo data to the server using an HTTP request.

[0116] The server stores the received photos in a database and organizes them into folders for analysis, using a cloud storage service such as Amazon S3.

[0117] 2. Generate a 3D silhouette of the product:

[0118] The server analyzes the stored photos using an image processing algorithm (e.g., OpenCV).

[0119] The server uses the analysis results to identify the contours and shape of the product, using edge detection and segmentation techniques in the process.

[0120] The server uses Structure-from-Motion (SfM) technology to generate a three-dimensional silhouette of the product based on the identified shape information.

[0121] The server saves the generated three-dimensional silhouette data as a 3D model file in preparation for the next process.

[0122] 3. Model and background specification:

[0123] The user inputs the desired model's physical information (e.g., height, weight, etc.) and background setting (e.g., studio background or outdoor background) through the system interface.

[0124] The terminal sends the setting information entered by the user to the server in JSON format.

[0125] The server saves the received setting information in a file and prepares for the next video generation process based on this information.

[0126] 4. Video creation using generative AI models:

[0127] The server integrates the three-dimensional silhouette data with the model's physique information and background settings specified by the user.

[0128] The server uses deep learning libraries (e.g., TensorFlow and PyTorch) to run generative AI models (e.g., Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF)).

[0129] The server inputs a prompt statement (e.g., "Generate a video of the person wearing a shirt and place it against a specified background") into the generative AI model, causing it to generate multiple videos.

[0130] The server displays the generated videos on a web interface and asks the user for confirmation.

[0131] 5. Video Selection:

[0132] The user can view the suggested videos on a web interface and select the one they like best.

[0133] The terminal captures the user's selection actions and transmits the information to the server.

[0134] The server processes the selected video data again at high resolution and saves it as the final video.

[0135] 6. Generate and save the final video:

[0136] The server encodes the selected video in high resolution and generates the final video file, using libraries such as FFmpeg.

[0137] The server stores the generated high-resolution video in cloud storage such as Amazon S3 and generates a link that allows users to download it.

[0138] The server notifies the user of the download link and provides it through a web interface.

[0139] Example: Creating a shirt video

[0140] 1. A user (retailer) uploads five photos of a new shirt to the system: front, back, left, right, and right-angle views.

[0141] 2. The server analyzes the five received photos, identifies the outline of the shirt using OpenCV, and generates a three-dimensional silhouette of the shirt using SfM technology.

[0142] 3. The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background) into the system's interface.

[0143] 4. The server uses the generative AI model to generate multiple videos based on the specified criteria and suggests those videos to the user on a web interface.

[0144] 5. The user selects the most suitable video from the suggested videos and sends that information to the server.

[0145] 6. The server regenerates the selected video in high resolution, stores it on Amazon S3, and then provides a link for the user to download it.

[0146] An example of a prompt sentence for this invention is, "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background." Using this prompt sentence, the generative AI model can generate the optimal video based on the required information.

[0147] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0148] Step 1: Upload your product photos

[0149] The user takes photos of the product they want to sell (for example, photos taken from the front, back, left and right, and diagonal angles), which are used as input data.

[0150] The user logs into the system's web interface, selects a photo of the product, and clicks the upload button.

[0151] The terminal transmits the selected photo data to the server using an HTTP request.

[0152] The server stores the received photo data using a cloud storage service such as Amazon S3.

[0153] Output: Image data stored in a folder for analysis, as the photos are saved on the server.

[0154] Step 2: Generate a 3D silhouette of the product

[0155] The server runs image processing algorithms such as OpenCV on the stored photos to analyze the contours and shapes of the products, using edge detection and segmentation techniques. This analysis serves as input data.

[0156] Based on the analyzed data, the server generates a 3D silhouette of the product using Structure-from-Motion (SfM) technology, a process that reconstructs a three-dimensional shape from multiple photographs.

[0157] Output: Three-dimensional silhouette data (3D model file) of the generated product.

[0158] Step 3: Specifying the model and background

[0159] The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background or outdoors) on the system interface. This becomes the input data.

[0160] The terminal converts the information entered by the user into JSON format and sends it to the server using an HTTP request.

[0161] The server saves the received JSON data to a file and uses this information when running the generative AI model.

[0162] Output: A JSON file with the saved model's physique and background information.

[0163] Step 4: Creating videos with generative AI models

[0164] The server takes in and integrates the three-dimensional silhouette data with the model and background information specified by the user. This integration becomes the input data.

[0165] The server starts running generative AI models (Generative Adversarial Networks (GANs) or Neural Radiance Fields (NeRF)) using deep learning libraries (e.g., TensorFlow or PyTorch).

[0166] The server inputs a prompt statement (e.g., "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background.") into the generative AI model and generates multiple videos.

[0167] Output: Multiple generated video files.

[0168] Step 5: Select a video

[0169] The user reviews the proposed videos on the system's web interface and selects the most suitable one, which becomes the input data.

[0170] The device captures information about the user's selected video, converts it into JSON format, and sends it to the server in an HTTP request.

[0171] The server reprocesses the selected video data in its final high resolution and prepares it for storage.

[0172] Output: The video data selected by the user.

[0173] Step 6: Generate and save the final video

[0174] The server uses a library such as FFmpeg to encode the selected video in high resolution and generate the final video file, which becomes the input data.

[0175] The server stores the high-resolution video in cloud storage such as Amazon S3 and generates a link that users can use to download it.

[0176] The server notifies the user of the download link through a web interface.

[0177] Output: A link to a high-resolution video file that users can download.

[0178] (Application example 1)

[0179] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0180] Traditional try-on experiences in brick-and-mortar stores required customers to physically try on the products to check them, which was time-consuming and labor-intensive. Also, trying on multiple products required customers to use a fitting room and take the products in and out, which was inefficient for both customers and stores. Furthermore, if a particular size or model was not in stock at the store, customers were unable to try on that product. To solve these problems, new technology was needed to enable virtual try-on in real time.

[0181] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0182] In this invention, the server includes means for receiving and storing product images, means for analyzing the received product images to generate a three-dimensional skeleton of the product, and means for generating a motion video in which the product is applied to a virtual model using a generative AI model based on the generated three-dimensional skeleton. This allows users to obtain product images in a physical store using a mobile device or visual aid device and have a virtual try-on experience in real time. This eliminates the hassle of trying on products and allows users to efficiently try on many products, thereby increasing customer purchasing motivation beyond the constraints of stock and size.

[0183] "Product images" are photos or image data of products taken or scanned by users.

[0184] A "three-dimensional skeleton" is three-dimensional shape information of a product generated from two-dimensional image data.

[0185] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate virtual fitting videos from image data.

[0186] "Motion video" is a video that shows a virtual model moving while wearing the product.

[0187] A "mobile terminal" is a portable information processing terminal such as a smartphone or tablet.

[0188] A "visual aid" is a device that assists with visual information, such as smart glasses or a head-mounted display.

[0189] The "user interface" refers to a screen or input device that allows the user to interact with the system, and includes functions for specifying body type information and background.

[0190] "High-definition video" refers to video data that has high resolution and quality and contains detailed information.

[0191] "Real-time" is a term that refers to processing occurring immediately, with little time delay.

[0192] "Virtual try-on experience" refers to the use of digital technology to provide the experience of trying on products without actually trying them on.

[0193] "Inventory" refers to the quantity and variety of products stored in a store or warehouse.

[0194] This technology allows users to upload product images, and generates a moving video of a virtual model wearing the product based on a three-dimensional skeleton generated from the image. To realize this technology, the following steps and system configuration are required.

[0195] Uploading product images

[0196] Users take or scan product images in physical stores using a mobile device or visual aid. The device sends the captured image to a server, which then stores it. Image processing libraries such as OpenCV are used to identify the position and shape of the product image.

[0197] Three-dimensional skeleton generation

[0198] The server analyzes the stored product images and generates a three-dimensional skeleton of the product, using Structure-from-Motion technology to extract the product's three-dimensional shape information from multiple image data.

[0199] Generation of motion images

[0200] Next, the server uses a generative AI model (e.g., GANs or NeRF) based on the generated three-dimensional skeleton to generate a moving image of the product applied to the virtual model. The user can specify the model's physical information (e.g., height, weight) and background information through the system's user interface.

[0201] Presentation and selection of motion images

[0202] The server generates an interface for presenting the generated motion images to the user, who can then review the motion images and select the most suitable one.

[0203] Final high-definition video generation and storage

[0204] The server stores the video selected by the user as high-definition video and allows the user to download it, providing a real-time virtual try-on experience and a comfortable shopping experience.

[0205] Hardware and software used

[0206] Hardware: mobile devices, smart glasses, head-mounted displays, servers

[0207] Software: OpenCV, Structure-from-Motion library, GANs, NeRF

[0208] Specific example explanation

[0209] As an example, consider the case where a user wants to virtually try on a shirt in a physical store.

[0210] 1. The user (customer) takes five photos of a shirt with their smartphone in a physical store (from the front, back, left and right, and diagonal).

[0211] 2. The device sends the captured photo to the server.

[0212] 3. The server analyzes the received photo and generates a three-dimensional skeleton of the shirt.

[0213] 4. The user specifies the model's physical information (e.g., height 170 cm, weight 60 kg) and the background (e.g., studio background).

[0214] 5. The server uses the generative AI model to generate multiple motion videos of the shirt being worn by the specified virtual model and suggests them to the user.

[0215] 6. The user selects the most suitable video from the suggested videos.

[0216] 7. The server stores the selected footage as high definition footage and makes it available for user download.

[0217] Prompt Sentence Examples

[0218] Photo path = upload_photo('Shirt photo taken in store.jpg')

[0219] Silhouette = generate_3d_silhouette(photopath)

[0220] User information = {'height': 170, 'weight': 60} The user enters their body type information.

[0221] background = 'store background' Use the store background as an example

[0222] Video = create_virtual_model(silhouette, user information, background)

[0223] save_video(video, 'Virtual Try-On Shirt 2023.mp4')

[0224] By following the above-described procedure, the present invention can generate a three-dimensional skeleton from a product image and provide a virtual try-on experience in real time.

[0225] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0226] Step 1:

[0227] Uploading product images

[0228] A user takes product images in a physical store using a mobile device or visual aid. The device sends the captured images to a server, which then stores the received images. Specifically, the user takes photos from multiple angles and the device uploads the image data to the server. The input is the captured product image, and the output is the image data stored on the server.

[0229] Step 2:

[0230] Three-dimensional skeleton generation

[0231] The server analyzes the stored product images and generates a three-dimensional skeleton of the product. Image processing libraries such as OpenCV are used for the analysis, and Structure-from-Motion technology is applied to extract the product's three-dimensional shape information from multiple image data. The input is the stored product image, and the output is three-dimensional skeleton data. Specifically, an image processing algorithm is executed on the server to measure the shape and size of the product.

[0232] Step 3:

[0233] Specifying the model and background

[0234] The user uses a user interface within the system to specify the model's physical information (e.g., height, weight) and background information. The input is the physical and background information entered by the user, and the output is a dataset containing that information. Specifically, the user enters the required information using drop-down menus and input forms in the user interface, and the information is sent to the server.

[0235] Step 4:

[0236] Virtual Model Generation

[0237] The server integrates the generated 3D skeleton with the model information specified by the user and uses a generative AI model (e.g., GANs or NeRF) to generate a moving image of the product applied to the virtual model. The input is the 3D skeleton data, model information, and background information, and the output is multiple moving images. Specifically, the generative AI model simulates the product's movement and executes the process of applying it to the virtual model.

[0238] Step 5:

[0239] Presentation and selection of motion images

[0240] The server generates an interface to present the generated multiple motion videos to the user. The user reviews the presented motion videos and selects the most suitable one. The input at this time is the generated motion video, and the output is the motion video selected by the user. Specifically, the user watches multiple videos on the interface and performs an operation to select the best one.

[0241] Step 6:

[0242] Final high-definition video generation and storage

[0243] The server saves the video selected by the user as high-definition video and makes it available for download. The input is the motion video selected by the user, and the output is the final high-definition video data. Specifically, the selected video is regenerated in high resolution on the server and provided to the user. During this process, the final video data is saved in a dedicated folder on the server.

[0244] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0245] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, the present invention enables more effective video generation and presentation.

[0246] System program processing explanation

[0247] 1. Upload product photos

[0248] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[0249] The terminal transmits the selected photo to the server.

[0250] The server stores the received photos and organizes them into folders for analysis.

[0251] 2. Creating a three-dimensional silhouette of a product

[0252] The server analyzes the received product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[0253] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[0254] 3. Specifying the model and background

[0255] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[0256] The terminal transmits this specification information to the server.

[0257] 4. Video creation using generative AI models

[0258] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[0259] The server generates an interface for suggesting the generated multiple videos to the user.

[0260] 5. Recognition of user emotions by emotion engine

[0261] The server runs an emotion engine that recognizes the user's emotions in real time during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc.

[0262] The device uses the user's camera and microphone to send emotion data to the emotion engine.

[0263] Based on the analyzed emotional information, the emotion engine identifies the videos that the user is most likely to be interested in and changes the order in which they are presented.

[0264] 6. Video Presentation and Selection

[0265] The server presents the video generated in the order adjusted by the emotion engine to the user.

[0266] The terminal displays the suggested videos to the user, and the user selects the most suitable video from the displayed videos.

[0267] The terminal transmits the user's selection information to the server.

[0268] 7. Generate and save the final video

[0269] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[0270] The server stores and provides the generated high-resolution video for users to download.

[0271] Specific examples

[0272] Example 1: Creating a shirt video

[0273] 1. A user (retailer) uploads five photos of a new shirt to the system, including front, back, left and right views, and three-quarter views.

[0274] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[0275] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[0276] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[0277] 5. The server runs an emotion engine during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc., and adjusts the presentation order.

[0278] 6. The user selects the most suitable video from the suggested videos.

[0279] 7. The server stores the selected video in high resolution and makes it available for download by the user.

[0280] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present optimal videos according to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[0281] The processing flow will be explained below.

[0282] Step 1:

[0283] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[0284] Step 2:

[0285] The terminal transmits the selected photo file to the server.

[0286] Step 3:

[0287] The server stores the received photo files and organizes them into folders for further processing.

[0288] Step 4:

[0289] The server opens the stored product photos and uses image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[0290] Step 5:

[0291] The server then integrates the shape information from the analyzed photos to generate a three-dimensional silhouette of the product, sometimes using Structure-from-Motion technology.

[0292] Step 6:

[0293] The user specifies the model's physical information (e.g., height, weight) and background information through an interface within the system.

[0294] Step 7:

[0295] The terminal transmits the physique information and background information input by the user to the server.

[0296] Step 8:

[0297] The server uses a generative AI model (e.g., GANs or NeRF) to generate a video of a virtual model wearing the product based on the three-dimensional silhouette and the physique and background information specified by the user.

[0298] Step 9:

[0299] The server generates an interface for suggesting the generated multiple videos to the user.

[0300] Step 10:

[0301] The server runs an emotion engine to acquire emotion data in real time from the user's camera and microphone while displaying the suggested video.

[0302] Step 11:

[0303] The terminal inputs the user's facial expressions, voice, click patterns, etc. into an emotion engine and analyzes the emotion data.

[0304] Step 12:

[0305] The emotion engine identifies the user's emotional state based on the analyzed emotion information, identifies the videos that the user is most interested in, and changes the presentation order of those videos.

[0306] Step 13:

[0307] The server presents the multiple videos to the user in an order adjusted by the emotion engine.

[0308] Step 14:

[0309] The terminal displays the videos to the user in the adjusted order, and the user selects the most suitable video from the suggested videos.

[0310] Step 15:

[0311] The terminal transmits the user's selection information to the server.

[0312] Step 16:

[0313] The server processes the video selected by the user as a high-definition video and generates the final video file.

[0314] Step 17:

[0315] The server stores and provides the generated high-resolution video for users to download.

[0316] Step 18:

[0317] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[0318] Example 2

[0319] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0320] Conventional systems have had difficulty generating a three-dimensional silhouette using product photos and creating videos in which a virtual model wears the product. Furthermore, the order in which the generated videos are presented does not take into account the user's interests or emotions, which makes it difficult to efficiently select the most suitable video. Furthermore, the interface for users to specify the model's physique and background information is inadequate, making it difficult to customize the system to meet individual needs.

[0321] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0322] In this invention, the server includes means for receiving and saving product photos via a user terminal, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user terminal, means for saving and providing a video selected by the user based on selection information transmitted from the user terminal as a final high-resolution video, and means for adjusting the video presentation order, which includes an emotion engine that collects and analyzes user emotion data using the camera and microphone of the user terminal. This significantly improves the user experience, enabling optimal video presentation and customized video generation according to individual needs.

[0323] "User terminal" refers to any device that a user uses to connect to the Internet and send and receive data.

[0324] A "server" is a computer system that provides data processing and storage functions over a network.

[0325] "Product Photos" refers to product image data uploaded by users.

[0326] "Three-dimensional silhouette" refers to a three-dimensional shape model of a product generated based on a product photo.

[0327] A "generative AI model" is a model that uses artificial intelligence algorithms to generate data to perform a specific task (in this case, video generation).

[0328] "Virtual Model" refers to a digital sculpture or animation of a digitally generated figure or form that simulates a tangible product or scene.

[0329] "High-definition video" refers to video files produced at high-quality image resolution.

[0330] An "emotion engine" refers to technology and algorithms that analyze a user's facial expressions, voice, behavioral patterns, etc. to recognize emotions in real time.

[0331] "Means for adjusting the presentation order" refers to a function that rearranges the order in which videos are displayed based on the analysis results of the emotion engine.

[0332] MODE FOR CARRYING OUT THE INVENTION

[0333] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. By combining this with an emotion engine that recognizes the user's emotions, the system presents the most appropriate video to the user.

[0334] The system of the present invention includes the following hardware and software.

[0335] User terminal: This can be a smartphone, tablet, PC, or other device. The user terminal connects to the Internet and sends and receives data via the system interface.

[0336] Server: A computer system that provides data processing and storage functions over a network. The server receives and analyzes photo data, generates 3D silhouettes, generates videos using generative AI models, runs the emotion engine, and displays and stores videos.

[0337] Image processing libraries: Image processing libraries such as OpenCV are used to analyze product photos and generate 3D silhouettes.

[0338] Generative AI models: Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF) are used to generate videos of virtual models wearing products.

[0339] Emotion engine: Analyzes the user's facial expressions, voice, and click patterns to adjust the order in which videos are presented.

[0340] Specific Examples

[0341] 1. Product photo upload example

[0342] A user (e.g., a representative from an online retailer) uses the system's interface to upload five photos of a new shirt: photos taken from the front, back, left and right, and an angle.

[0343] The terminal divides the uploaded photo into packets and sends them to the server.

[0344] 2. Product photo analysis and three-dimensional silhouette generation example

[0345] The server analyzes the received product photos and extracts feature points from each photo using OpenCV.

[0346] The server uses Structure-from-Motion technology to reconstruct the three-dimensional shape of the product based on the extracted feature points and generate a three-dimensional silhouette.

[0347] 3. Example of specifying a model and background

[0348] The user inputs the virtual model's physical information (for example, height 170 cm, weight 60 kg) and background information (for example, studio background) through an interface provided within the system.

[0349] The terminal transmits this information to the server.

[0350] 4. Example of video creation using generative AI model

[0351] The server combines the generated three-dimensional silhouette with specified model information and uses a generative AI model such as GANs or NeRF to generate a video of the virtual model wearing the product.

[0352] The server provides an interface to present the generated multiple animation versions to the user.

[0353] 5. Example of user emotion recognition using emotion engine

[0354] The device uses a camera and microphone to capture the user's facial expressions and voice while watching videos, and also records click patterns.

[0355] The device transmits the captured data in real time to a server, where an emotion engine built into the server analyzes the data.

[0356] The server adjusts the presentation order of the videos based on the analysis results.

[0357] 6. Video Presentation and Selection Examples

[0358] The server presents the generated video in the adjusted order to the user terminal.

[0359] The user selects the most suitable video from the displayed videos.

[0360] 7. Example of generating and saving the final video

[0361] The server renders the user's selected video in high resolution and generates the final video file.

[0362] The server stores the generated high-resolution video in cloud storage or on a dedicated server and provides a link for users to download it.

[0363] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present videos that are optimally tailored to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[0364] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0365] The flow of this system's program processing

[0366] Step 1:

[0367] Users upload product photos using the upload interface on the system. The input is the product photo selected by the user (front, back, left, right, oblique, etc.), and the output is a photo file sent to the server via the terminal.

[0368] Step 2:

[0369] The terminal divides the received product photos into packets and securely sends them to the server. The input is the photo data uploaded by the user, and the output is the divided photo data sent to the server. Specific operations include data packetization and encrypted communication.

[0370] Step 3:

[0371] The server saves the received photo data and organizes it in a folder for analysis. The input is the photo data sent from the terminal, and the output is the saved photo file and product information recorded in the database.

[0372] Step 4:

[0373] The server analyzes product photos using an image processing library such as OpenCV. The input is the saved product photo, and the output is the extracted feature point data. Specific operations include image reading and feature point detection.

[0374] Step 5:

[0375] The server matches feature points between multiple photos and reconstructs the three-dimensional shape of the product using Structure-from-Motion technology. The input is feature point data extracted from multiple photos, and the output is a three-dimensional silhouette (3D model) of the product. Specific operations include matching corresponding feature points, estimating camera position, and generating a point cloud.

[0376] Step 6:

[0377] The user specifies model physique information and background information through the interface. The input is the model and background information entered by the user, and the output is the model information and background information sent from the terminal.

[0378] Step 7:

[0379] The terminal transmits the specified model information and background information to the server. The input is the model information and background information input by the user, and the output is the model information and background information transmitted to the server.

[0380] Step 8:

[0381] The server integrates the model information and the 3D silhouette and uses a generative AI model to generate a video of the virtual model wearing the product. The input is the model information and the 3D silhouette, and the output is multiple generated videos. Specific operations include inputting data into the generative AI model and simulating the movement of the product.

[0382] Step 9:

[0383] The server generates an interface for proposing the generated videos to the user, where the input is the generated videos and the output is the video suggestion interface displayed on the user terminal.

[0384] Step 10:

[0385] The device collects emotional data using the user's camera and microphone and sends it to a server. The input is the user's facial expressions, voice, and click patterns, and the output is the emotional data sent to the server. Specific operations include facial expression recognition and voice analysis.

[0386] Step 11:

[0387] The server analyzes the data collected by the emotion engine and adjusts the video presentation order. The input is the transmitted emotion data, and the output is the adjusted video presentation order. Specific operations include data analysis and rearrangement of the presentation order.

[0388] Step 12:

[0389] The user selects the most suitable video from the displayed videos. The input is the video presented by the server, and the output is the selected video information.

[0390] Step 13:

[0391] The terminal transmits the user's selection information to the server, where the input is the video information selected by the user and the output is the selection information transmitted to the server.

[0392] Step 14:

[0393] The server renders the selected video in high resolution and generates and saves the final video file. The input is the selected video information, and the output is a high-resolution video file. Specific operations include high-resolution rendering and video file saving.

[0394] Step 15:

[0395] The server generates a link that allows the user to download the generated high-resolution video and provides it to the user. The input is the generated high-resolution video, and the output is the download link.

[0396] (Application example 2)

[0397] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0398] In today's online shopping environment, users cannot actually touch the product, making it difficult to accurately grasp its appearance and feel. There are also limited ways to maximize the product's appeal and effectively communicate it. Furthermore, there is a lack of ways to recognize users' purchasing intentions and interests and make optimal suggestions based on them, creating a need to improve the user experience.

[0399] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0400] In this invention, the server includes means for receiving and saving product photos, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user, means for saving and providing a video selected by the user as a final high-resolution video, means for recognizing and analyzing the user's emotions, means for adjusting the presentation order of videos that are likely to be of most interest based on the analyzed emotional information, and means for providing an interface for the user to specify the model's physique information and background when displaying the generated videos. This allows users to accurately grasp the appearance and feel of the product in real time, and makes it possible to suggest optimal videos according to their individual interests and emotions.

[0401] "Product Photos" are images taken or provided by users that capture the exterior of a product from multiple angles.

[0402] A "three-dimensional silhouette" is a digital model that reproduces the shape and size of a product in three dimensions based on photographs of multiple products.

[0403] A "generative AI model" is an artificial intelligence model that uses deep learning techniques to generate new data from data provided by users, and includes, for example, Generative Adversarial Networks (GANs).

[0404] A "virtual model" is a digital human model created by a generative AI model based on physique information specified by the user.

[0405] "Emotion recognition" is a technology that analyzes a user's facial expressions, voice, click patterns, etc. to identify their emotional state at that time.

[0406] "High-definition video" is a video file created with a custom resolution setting that is higher than the standard resolution.

[0407] The "presentation order" is the order in which multiple generated videos and information are displayed to the user.

[0408] "Physical information" is data that represents a person's physical characteristics, such as the user's height, weight, and gender.

[0409] An "interface" is a software part that provides the screen and operating means for the user to interact with the system.

[0410] "Background" means the digital background scene or environment in which the virtual model is displayed, as selected by the user.

[0411] The system for realizing the present invention is configured using the following specific hardware and software.

[0412] Hardware:

[0413] Devices: Smartphone, camera, microphone

[0414] Server: A high-performance computer, either a cloud service or on-premise

[0415] software:

[0416] OpenCV: Image processing library

[0417] Emotion Recognition Engine: Software that recognizes user emotions in real time

[0418] Generative AI models: Deep learning techniques, such as Generative Adversarial Networks (GANs), to generate new data based on user-uploaded photos.

[0419] User Interface: The part of the software that allows the user to interact with the system.

[0420] System Details:

[0421] 1. Upload product photos

[0422] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles and upload these photos to the app, which then sends the uploaded photos to the server.

[0423] 2. Creating a three-dimensional silhouette of a product

[0424] The server analyzes the received product photos and uses image processing libraries such as OpenCV to identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product.

[0425] 3. Specifying the model and background

[0426] The user specifies his / her physical information (e.g., height, weight) and background image through an interface provided within the system. The terminal transmits this information to the server.

[0427] 4. Video Generation

[0428] Based on the three-dimensional silhouette and the specified model information, the server uses a generative AI model to generate a video in which the virtual model wears the product, taking into account the background information specified by the user.

[0429] 5. Emotional Engine Optimization

[0430] The server uses an emotion recognition engine to analyze the user's facial expressions, voice, click patterns, etc. in real time, and adjusts the order of videos presented to them based on the user's likely interest.

[0431] 6. Video Presentation and Selection

[0432] The generated videos are presented to the user in the order adjusted by the emotion engine. The user selects the most suitable video from these videos. The selected video is then sent from the device to the server.

[0433] 7. Generate and save the final video

[0434] The server finally stores the user's selected video in high resolution and makes it available for the user to download.

[0435] Program operation description:

[0436] This system operates by combining the above hardware and software. First, the device sends user input to the server, which then analyzes and processes the received data. OpenCV is used for image analysis, and a generative AI model (such as GANs) generates the video. The emotion recognition engine collects and analyzes the user's emotional data. The user interface is the means by which the user inputs and makes selections, and also presents the final video.

[0437] Examples:

[0438] For example, if a user wants to try on a new dress, the process would be as follows:

[0439] 1. The user takes photos of the dress from multiple angles and uploads them to the app.

[0440] 2. The server generates a three-dimensional silhouette of the dress based on these photos.

[0441] 3. The user specifies their physical information (e.g., height 170 cm, weight 60 kg) and a background image (e.g., a beach scene).

[0442] 4. The server uses the generative AI model to generate multiple videos of the virtual model wearing the dress.

[0443] 5. The emotion engine analyzes the user's facial expressions and click patterns to prioritize presenting the most interesting videos to the user.

[0444] 6. Users can select their favorite videos and save and share them as final high-resolution videos.

[0445] Example prompt sentence:

[0446] 1. Take photos of your product from multiple angles and upload them to the app.

[0447] 2. Select your physical information and background.

[0448] 3. Review the videos generated by the app and choose the one you like the most.

[0449] This allows users to accurately understand the appearance and feel of a product in real time, making it possible to suggest the most suitable video based on their individual interests and emotions.

[0450] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0451] Step 1:

[0452] Users take and upload product photos

[0453] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles. These photos are then uploaded to the app. The input is the photos of multiple products taken, and the output is the photo data sent to the server. Specifically, users use the app's photo upload interface, and their device sends these photos to the server.

[0454] Step 2:

[0455] The server receives and saves the product photos.

[0456] The server receives product photos sent from the device and saves them in a folder for analysis. The input is the photo data sent from the device, and the output is a photo file stored in a storage folder on the server. The server first saves the received files in a specific directory and organizes them within the folder.

[0457] Step 3:

[0458] The server analyzes the product photo and generates a three-dimensional silhouette

[0459] The server uses image processing libraries such as OpenCV to analyze the received product photos and identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product. The input is multiple stored photos, and the output is three-dimensional silhouette data. The server then runs image analysis algorithms to detect edges and extract specific features from the photos.

[0460] Step 4:

[0461] User specifies model and background information

[0462] The user specifies the model's physical information (e.g., height, weight) and background information through an interface provided within the system. The input is the user's physical information and background image selection data, and the output is that this specified information is sent to the server. The user enters information using pull-down menus and text boxes, and the terminal sends this data to the server.

[0463] Step 5:

[0464] The server generates the video

[0465] The server integrates the 3D silhouette with the specified model information and uses a generative AI model (e.g., GANs) to generate a video in which the virtual model wears the product. At this time, it also takes into account background information specified by the user. The input is the 3D silhouette, physique information, and background information, and the output is the generated video data. Specifically, the server runs the generative AI model and generates a video in which the virtual model wears the product based on the input data.

[0466] Step 6:

[0467] The server uses an emotion engine to recognize and analyze the user's emotions in real time.

[0468] The server runs an emotion recognition engine and analyzes the user's facial expressions, voice, click patterns, etc. in real time. The input is the user's emotional data (facial expressions, voice, click patterns), and the output is analyzed emotional information. The device uses a camera and microphone to acquire user data and sends it to the server. The server sends this data to the emotion engine and returns the analysis results.

[0469] Step 7:

[0470] The server adjusts the video presentation order based on emotional information

[0471] The server adjusts the presentation order of videos that the user is most likely to be interested in based on the analyzed emotional information. The input is the analyzed emotional information, and the output is the adjusted video presentation order. The server rearranges the video presentation order based on the results of the emotion engine.

[0472] Step 8:

[0473] Presentation and selection of generated videos

[0474] The server presents the videos generated in the order adjusted by the emotion engine to the user, and the user selects the most appropriate video from the displayed videos. The input is the video presented by the server, and the output is the user's selection information. The terminal displays the video to the user and sends the user's selection to the server.

[0475] Step 9:

[0476] Generate and save the final video

[0477] The server saves the user-selected video in high resolution and generates the final video file. The input is the user-selected video, and the output is the final high-resolution video file. The server saves this in cloud storage and makes it available for user download.

[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0481] [Second embodiment]

[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0494] The system of the present invention is a technology that allows users to upload product photos and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. The program of the system of the present invention is explained below in natural language.

[0495] System program processing explanation

[0496] 1. Upload product photos

[0497] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[0498] The terminal transmits the selected photo to the server.

[0499] The server stores the received photos and organizes them into folders for analysis.

[0500] 2. Creating a three-dimensional silhouette of a product

[0501] The server analyzes the received product photos and runs image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[0502] The server uses 3D modeling technology, such as Structure-from-Motion technology, to generate a three-dimensional silhouette of the product using the identified product shape information.

[0503] 3. Specifying the model and background

[0504] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[0505] The terminal transmits this specification information to the server.

[0506] 4. Video creation using generative AI models

[0507] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[0508] The server generates an interface for suggesting the generated multiple videos to the user.

[0509] 5. Select a video

[0510] The user reviews the suggested videos and selects the most suitable one.

[0511] The terminal transmits the user's selection information to the server.

[0512] 6. Generate and save the final video

[0513] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[0514] The server stores the generated high-resolution video therein so that the user can download it, and provides it to the user.

[0515] Specific examples

[0516] Example 1: Creating a shirt video

[0517] 1. A user (retailer) uploads five photos of a new shirt to the system, including photos from the front, back, left and right, and one-quarter angles.

[0518] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[0519] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[0520] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[0521] 5. The user selects the most suitable video from the suggested videos.

[0522] 6. The server stores the selected video in high resolution and makes it available for download by the user.

[0523] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills, thereby effectively supporting the promotion of fashion merchandise purchases.

[0524] The processing flow will be explained below.

[0525] Step 1:

[0526] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[0527] Step 2:

[0528] The terminal transmits the selected photo to the server.

[0529] Step 3:

[0530] The server stores the received photo files and organizes them for further processing.

[0531] Step 4:

[0532] The server analyzes the stored product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[0533] Step 5:

[0534] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[0535] Step 6:

[0536] The user uses the system interface to specify the model's physical information (e.g., height, weight) and background information.

[0537] Step 7:

[0538] The terminal transmits the information specified by the user (physique information and background information of the model) to the server.

[0539] Step 8:

[0540] The server integrates the 3D silhouette with the user-specified information, uses a generative AI model to dress the virtual model in the product, and generates a video against the specified background, using generative technologies such as GANs and NeRF.

[0541] Step 9:

[0542] The server generates an interface for suggesting the generated multiple videos to the user.

[0543] Step 10:

[0544] The terminal displays the suggested videos to the user.

[0545] Step 11:

[0546] The user selects the most suitable video from the displayed videos.

[0547] Step 12:

[0548] The terminal transmits the user's selection information to the server.

[0549] Step 13:

[0550] The server processes the user's selected video in high resolution and generates the final video file.

[0551] Step 14:

[0552] The server stores and provides the generated high-resolution video for users to download.

[0553] Step 15:

[0554] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[0555] Example 1

[0556] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0557] Traditional online shopping has the problem that customers cannot actually try on products, which can cause anxiety about the size and fit of the product. Furthermore, product photos and simple descriptions alone make it difficult to fully check the product's details and actual appearance, which can lead to hesitation in purchasing. Furthermore, the need for specialized knowledge and skills to take product photos and create videos places a significant burden on product sellers.

[0558] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0559] In this invention, the server includes a means for receiving and saving product photos, a means for analyzing the photos to generate a three-dimensional silhouette of the product, a means for generating a video of a virtual model wearing the product using a generative AI model based on the generated silhouette, a means for presenting multiple videos to the user, and a means for saving and providing the video selected by the user as a final high-resolution video. This allows users to check the appearance and fit of products in detail without actually trying them on, eliminating any anxiety about purchasing. Furthermore, product sellers can create high-quality product promotional videos without specialized knowledge of photography or editing, leading to increased sales.

[0560] "Product photos" are image data of products uploaded by users.

[0561] "Means for storing" refers to the system's function of storing received product photos in a database or storage.

[0562] The "means for analyzing" refers to a function of the system that executes an image processing algorithm to identify the contours and shape of the product using a product photograph and generate a three-dimensional silhouette.

[0563] A "three-dimensional silhouette" is three-dimensional shape information of a product generated from a product photo.

[0564] A "generative AI model" is an algorithm that uses deep learning technology to generate new data based on input data, and in this invention is used to generate videos of virtual models.

[0565] A "virtual model" is a model of a person or object that is digitally generated using a generative AI model.

[0566] The "means for generating video" is a system function that creates video data based on the generated three-dimensional silhouette and virtual model.

[0567] The "presentation means" is a system function that displays the generated multiple videos so that the user can check them.

[0568] "High-resolution video" refers to high-quality video data that can display every detail clearly.

[0569] The system of the present invention is a technology that allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Specific embodiments of the system of the present invention are described below.

[0570] 1. Upload product photos:

[0571] Users take photos of the items they want to sell using a digital camera or smartphone.

[0572] The user logs into the system's web interface and selects a product photo in the upload window.

[0573] The terminal transmits the selected photo data to the server using an HTTP request.

[0574] The server stores the received photos in a database and organizes them into folders for analysis, using a cloud storage service such as Amazon S3.

[0575] 2. Generate a 3D silhouette of the product:

[0576] The server analyzes the stored photos using an image processing algorithm (e.g., OpenCV).

[0577] The server uses the analysis results to identify the contours and shape of the product, using edge detection and segmentation techniques in the process.

[0578] The server uses Structure-from-Motion (SfM) technology to generate a three-dimensional silhouette of the product based on the identified shape information.

[0579] The server saves the generated three-dimensional silhouette data as a 3D model file in preparation for the next process.

[0580] 3. Specify the model and background:

[0581] The user inputs the desired model's physical information (e.g., height, weight, etc.) and background setting (e.g., studio background or outdoor background) through the system interface.

[0582] The terminal sends the setting information entered by the user to the server in JSON format.

[0583] The server saves the received setting information in a file and prepares for the next video generation process based on this information.

[0584] 4. Video creation using generative AI models:

[0585] The server integrates the three-dimensional silhouette data with the model's physique information and background settings specified by the user.

[0586] The server uses deep learning libraries (e.g., TensorFlow and PyTorch) to run generative AI models (e.g., Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF)).

[0587] The server inputs a prompt statement (e.g., "Generate a video of the person wearing a shirt and place it against a specified background") into the generative AI model, causing it to generate multiple videos.

[0588] The server displays the generated videos on a web interface and asks the user for confirmation.

[0589] 5. Video Selection:

[0590] The user can view the suggested videos on a web interface and select the one they like best.

[0591] The terminal captures the user's selection actions and transmits the information to the server.

[0592] The server processes the selected video data again at high resolution and saves it as the final video.

[0593] 6. Generate and save the final video:

[0594] The server encodes the selected video in high resolution and generates the final video file, using libraries such as FFmpeg.

[0595] The server stores the generated high-resolution video in cloud storage such as Amazon S3 and generates a link that allows users to download it.

[0596] The server notifies the user of the download link and provides it through a web interface.

[0597] Example: Creating a shirt video

[0598] 1. A user (retailer) uploads five photos of a new shirt to the system: front, back, left, right, and right-angle views.

[0599] 2. The server analyzes the five received photos, identifies the outline of the shirt using OpenCV, and generates a three-dimensional silhouette of the shirt using SfM technology.

[0600] 3. The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background) into the system's interface.

[0601] 4. The server uses the generative AI model to generate multiple videos based on the specified criteria and suggests those videos to the user on a web interface.

[0602] 5. The user selects the most suitable video from the suggested videos and sends that information to the server.

[0603] 6. The server regenerates the selected video in high resolution, stores it on Amazon S3, and then provides a link for the user to download it.

[0604] An example of a prompt sentence for this invention is, "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background." Using this prompt sentence, the generative AI model can generate the optimal video based on the required information.

[0605] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0606] Step 1: Upload your product photos

[0607] The user takes photos of the product they want to sell (for example, photos taken from the front, back, left and right, and diagonal angles), which are used as input data.

[0608] The user logs into the system's web interface, selects a photo of the product, and clicks the upload button.

[0609] The terminal transmits the selected photo data to the server using an HTTP request.

[0610] The server stores the received photo data using a cloud storage service such as Amazon S3.

[0611] Output: Image data stored in a folder for analysis, as the photos are saved on the server.

[0612] Step 2: Generate a 3D silhouette of the product

[0613] The server runs image processing algorithms such as OpenCV on the stored photos to analyze the contours and shapes of the products, using edge detection and segmentation techniques. This analysis serves as input data.

[0614] Based on the analyzed data, the server generates a 3D silhouette of the product using Structure-from-Motion (SfM) technology, a process that reconstructs a three-dimensional shape from multiple photographs.

[0615] Output: Three-dimensional silhouette data (3D model file) of the generated product.

[0616] Step 3: Specifying the model and background

[0617] The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background or outdoors) on the system interface. This becomes the input data.

[0618] The terminal converts the information entered by the user into JSON format and sends it to the server using an HTTP request.

[0619] The server saves the received JSON data to a file and uses this information when running the generative AI model.

[0620] Output: A JSON file with the saved model's physique and background information.

[0621] Step 4: Creating videos with generative AI models

[0622] The server takes in and integrates the three-dimensional silhouette data with the model and background information specified by the user. This integration becomes the input data.

[0623] The server starts running generative AI models (Generative Adversarial Networks (GANs) or Neural Radiance Fields (NeRF)) using deep learning libraries (e.g., TensorFlow or PyTorch).

[0624] The server inputs a prompt statement (e.g., "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background.") into the generative AI model and generates multiple videos.

[0625] Output: Multiple generated video files.

[0626] Step 5: Select a video

[0627] The user reviews the proposed videos on the system's web interface and selects the most suitable one, which becomes the input data.

[0628] The device captures information about the user's selected video, converts it into JSON format, and sends it to the server in an HTTP request.

[0629] The server reprocesses the selected video data in its final high resolution and prepares it for storage.

[0630] Output: The video data selected by the user.

[0631] Step 6: Generate and save the final video

[0632] The server uses a library such as FFmpeg to encode the selected video in high resolution and generate the final video file, which becomes the input data.

[0633] The server stores the high-resolution video in cloud storage such as Amazon S3 and generates a link that users can use to download it.

[0634] The server notifies the user of the download link through a web interface.

[0635] Output: A link to a high-resolution video file that users can download.

[0636] (Application example 1)

[0637] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0638] Traditional try-on experiences in brick-and-mortar stores required customers to physically try on the products to check them, which was time-consuming and labor-intensive. Also, trying on multiple products required customers to use a fitting room and take the products in and out, which was inefficient for both customers and stores. Furthermore, if a particular size or model was not in stock at the store, customers were unable to try on that product. To solve these problems, new technology was needed to enable virtual try-on in real time.

[0639] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0640] In this invention, the server includes means for receiving and storing product images, means for analyzing the received product images to generate a three-dimensional skeleton of the product, and means for generating a motion video in which the product is applied to a virtual model using a generative AI model based on the generated three-dimensional skeleton. This allows users to obtain product images in a physical store using a mobile device or visual aid device and have a virtual try-on experience in real time. This eliminates the hassle of trying on products and allows users to efficiently try on many products, thereby increasing customer purchasing motivation beyond the constraints of stock and size.

[0641] "Product images" are photos or image data of products taken or scanned by users.

[0642] A "three-dimensional skeleton" is three-dimensional shape information of a product generated from two-dimensional image data.

[0643] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate virtual fitting videos from image data.

[0644] "Motion video" is a video that shows a virtual model moving while wearing the product.

[0645] A "mobile terminal" is a portable information processing terminal such as a smartphone or tablet.

[0646] A "visual aid" is a device that assists with visual information, such as smart glasses or a head-mounted display.

[0647] The "user interface" refers to a screen or input device that allows the user to interact with the system, and includes functions for specifying body type information and background.

[0648] "High-definition video" refers to video data that has high resolution and quality and contains detailed information.

[0649] "Real-time" is a term that refers to processing occurring immediately, with little time delay.

[0650] "Virtual try-on experience" means using digital technology to provide the experience of trying on products without actually trying them on.

[0651] "Inventory" refers to the quantity and variety of products stored in a store or warehouse.

[0652] This technology allows users to upload product images, and generates a moving video of a virtual model wearing the product based on a three-dimensional skeleton generated from the image. To realize this technology, the following steps and system configuration are required.

[0653] Uploading product images

[0654] Users take or scan product images in physical stores using a mobile device or visual aid. The device sends the captured image to a server, which then stores it. Image processing libraries such as OpenCV are used to identify the position and shape of the product image.

[0655] Three-dimensional skeleton generation

[0656] The server analyzes the stored product images and generates a three-dimensional skeleton of the product, using Structure-from-Motion technology to extract the product's three-dimensional shape information from multiple image data.

[0657] Generation of motion images

[0658] Next, the server uses a generative AI model (e.g., GANs or NeRF) based on the generated three-dimensional skeleton to generate a moving image of the product applied to the virtual model. The user can specify the model's physical information (e.g., height, weight) and background information through the system's user interface.

[0659] Presentation and selection of motion images

[0660] The server generates an interface for presenting the generated motion images to the user, who can then review the motion images and select the most suitable one.

[0661] Final high-definition video generation and storage

[0662] The server stores the video selected by the user as high-definition video and allows the user to download it, providing a real-time virtual try-on experience and a comfortable shopping experience.

[0663] Hardware and software used

[0664] Hardware: mobile devices, smart glasses, head-mounted displays, servers

[0665] Software: OpenCV, Structure-from-Motion library, GANs, NeRF

[0666] Specific example explanation

[0667] As an example, consider the case where a user wants to virtually try on a shirt in a physical store.

[0668] 1. The user (customer) takes five photos of a shirt with their smartphone in a physical store (from the front, back, left and right, and diagonal).

[0669] 2. The device sends the captured photo to the server.

[0670] 3. The server analyzes the received photo and generates a three-dimensional skeleton of the shirt.

[0671] 4. The user specifies the model's physical information (e.g., height 170 cm, weight 60 kg) and the background (e.g., studio background).

[0672] 5. The server uses the generative AI model to generate multiple motion videos of the shirt being worn by the specified virtual model and suggests them to the user.

[0673] 6. The user selects the most suitable video from the suggested videos.

[0674] 7. The server stores the selected footage as high definition footage and makes it available for user download.

[0675] Prompt Sentence Examples

[0676] Photo path = upload_photo('Shirt photo taken in store.jpg')

[0677] Silhouette = generate_3d_silhouette(photopath)

[0678] User information = {'height': 170, 'weight': 60} The user enters their body type information.

[0679] background = 'store background' Use the store background as an example

[0680] Video = create_virtual_model(silhouette, user information, background)

[0681] save_video(video, 'Virtual Try-On Shirt 2023.mp4')

[0682] By following the above-described procedure, the present invention can generate a three-dimensional skeleton from a product image and provide a virtual try-on experience in real time.

[0683] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0684] Step 1:

[0685] Uploading product images

[0686] A user takes product images in a physical store using a mobile device or visual aid. The device sends the captured images to a server, which then stores the received images. Specifically, the user takes photos from multiple angles and the device uploads the image data to the server. The input is the captured product image, and the output is the image data stored on the server.

[0687] Step 2:

[0688] Three-dimensional skeleton generation

[0689] The server analyzes the stored product images and generates a three-dimensional skeleton of the product. Image processing libraries such as OpenCV are used for the analysis, and Structure-from-Motion technology is applied to extract the product's three-dimensional shape information from multiple image data. The input is the stored product image, and the output is three-dimensional skeleton data. Specifically, an image processing algorithm is executed on the server to measure the shape and size of the product.

[0690] Step 3:

[0691] Specifying the model and background

[0692] The user uses a user interface within the system to specify the model's physical information (e.g., height, weight) and background information. The input is the physical and background information entered by the user, and the output is a dataset containing that information. Specifically, the user enters the required information using drop-down menus and input forms in the user interface, and the information is sent to the server.

[0693] Step 4:

[0694] Virtual Model Generation

[0695] The server integrates the generated 3D skeleton with the model information specified by the user and uses a generative AI model (e.g., GANs or NeRF) to generate a moving image of the product applied to the virtual model. The input is the 3D skeleton data, model information, and background information, and the output is multiple moving images. Specifically, the generative AI model simulates the product's movement and executes the process of applying it to the virtual model.

[0696] Step 5:

[0697] Presentation and selection of motion images

[0698] The server generates an interface to present the generated multiple motion videos to the user. The user reviews the presented motion videos and selects the most suitable one. The input at this time is the generated motion video, and the output is the motion video selected by the user. Specifically, the user watches multiple videos on the interface and performs an operation to select the best one.

[0699] Step 6:

[0700] Final high-definition video generation and storage

[0701] The server saves the video selected by the user as high-definition video and makes it available for download. The input is the motion video selected by the user, and the output is the final high-definition video data. Specifically, the selected video is regenerated in high resolution on the server and provided to the user. During this process, the final video data is saved in a dedicated folder on the server.

[0702] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0703] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, the present invention enables more effective video generation and presentation.

[0704] System program processing explanation

[0705] 1. Upload product photos

[0706] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[0707] The terminal transmits the selected photo to the server.

[0708] The server stores the received photos and organizes them into folders for analysis.

[0709] 2. Creating a three-dimensional silhouette of a product

[0710] The server analyzes the received product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[0711] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[0712] 3. Specifying the model and background

[0713] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[0714] The terminal transmits this specification information to the server.

[0715] 4. Video creation using generative AI models

[0716] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[0717] The server generates an interface for suggesting the generated multiple videos to the user.

[0718] 5. Recognition of user emotions by emotion engine

[0719] The server runs an emotion engine that recognizes the user's emotions in real time during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc.

[0720] The device uses the user's camera and microphone to send emotion data to the emotion engine.

[0721] Based on the analyzed emotional information, the emotion engine identifies the videos that the user is most likely to be interested in and changes the order in which they are presented.

[0722] 6. Video Presentation and Selection

[0723] The server presents the video generated in the order adjusted by the emotion engine to the user.

[0724] The terminal displays the suggested videos to the user, and the user selects the most suitable video from the displayed videos.

[0725] The terminal transmits the user's selection information to the server.

[0726] 7. Generate and save the final video

[0727] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[0728] The server stores and provides the generated high-resolution video for users to download.

[0729] Specific examples

[0730] Example 1: Creating a shirt video

[0731] 1. A user (retailer) uploads five photos of a new shirt to the system, including front, back, left and right views, and three-quarter views.

[0732] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[0733] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[0734] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[0735] 5. The server runs an emotion engine during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc., and adjusts the presentation order.

[0736] 6. The user selects the most suitable video from the suggested videos.

[0737] 7. The server stores the selected video in high resolution and makes it available for download by the user.

[0738] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present optimal videos according to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[0739] The processing flow will be explained below.

[0740] Step 1:

[0741] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[0742] Step 2:

[0743] The terminal transmits the selected photo file to the server.

[0744] Step 3:

[0745] The server stores the received photo files and organizes them into folders for further processing.

[0746] Step 4:

[0747] The server opens the stored product photos and uses image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[0748] Step 5:

[0749] The server then integrates the shape information from the analyzed photos to generate a three-dimensional silhouette of the product, sometimes using Structure-from-Motion technology.

[0750] Step 6:

[0751] The user specifies the model's physical information (e.g., height, weight) and background information through an interface within the system.

[0752] Step 7:

[0753] The terminal transmits the physique information and background information input by the user to the server.

[0754] Step 8:

[0755] The server uses a generative AI model (e.g., GANs or NeRF) to generate a video of a virtual model wearing the product based on the three-dimensional silhouette and the physique and background information specified by the user.

[0756] Step 9:

[0757] The server generates an interface for suggesting the generated multiple videos to the user.

[0758] Step 10:

[0759] The server runs an emotion engine to acquire emotion data in real time from the user's camera and microphone while displaying the suggested video.

[0760] Step 11:

[0761] The terminal inputs the user's facial expressions, voice, click patterns, etc. into an emotion engine and analyzes the emotion data.

[0762] Step 12:

[0763] The emotion engine identifies the user's emotional state based on the analyzed emotion information, identifies the videos that the user is most interested in, and changes the presentation order of those videos.

[0764] Step 13:

[0765] The server presents the multiple videos to the user in an order adjusted by the emotion engine.

[0766] Step 14:

[0767] The terminal displays the videos to the user in the adjusted order, and the user selects the most suitable video from the suggested videos.

[0768] Step 15:

[0769] The terminal transmits the user's selection information to the server.

[0770] Step 16:

[0771] The server processes the video selected by the user as a high-definition video and generates the final video file.

[0772] Step 17:

[0773] The server stores and provides the generated high-resolution video for users to download.

[0774] Step 18:

[0775] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[0776] Example 2

[0777] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0778] Conventional systems have had difficulty generating a three-dimensional silhouette using product photos and creating videos in which a virtual model wears the product. Furthermore, the order in which the generated videos are presented does not take into account the user's interests or emotions, which makes it difficult to efficiently select the most suitable video. Furthermore, the interface for users to specify the model's physique and background information is inadequate, making it difficult to customize the system to meet individual needs.

[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0780] In this invention, the server includes means for receiving and saving product photos via a user terminal, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user terminal, means for saving and providing a video selected by the user based on selection information transmitted from the user terminal as a final high-resolution video, and means for adjusting the video presentation order, which includes an emotion engine that collects and analyzes user emotion data using the camera and microphone of the user terminal. This significantly improves the user experience, enabling optimal video presentation and customized video generation according to individual needs.

[0781] "User terminal" refers to any device that a user uses to connect to the Internet and send and receive data.

[0782] A "server" is a computer system that provides data processing and storage functions over a network.

[0783] "Product Photos" refers to product image data uploaded by users.

[0784] "Three-dimensional silhouette" refers to a three-dimensional shape model of a product generated based on a product photo.

[0785] A "generative AI model" is a model that uses artificial intelligence algorithms to generate data to perform a specific task (in this case, video generation).

[0786] "Virtual Model" refers to a digital sculpture or animation of a digitally generated figure or form that simulates a tangible product or scene.

[0787] "High-definition video" refers to video files produced at high-quality image resolution.

[0788] An "emotion engine" refers to technology and algorithms that analyze a user's facial expressions, voice, behavioral patterns, etc. to recognize emotions in real time.

[0789] "Means for adjusting the presentation order" refers to a function that rearranges the order in which videos are displayed based on the analysis results of the emotion engine.

[0790] MODE FOR CARRYING OUT THE INVENTION

[0791] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. By combining this with an emotion engine that recognizes the user's emotions, the system presents the most appropriate video to the user.

[0792] The system of the present invention includes the following hardware and software.

[0793] User terminal: This can be a smartphone, tablet, PC, or other device. The user terminal connects to the Internet and sends and receives data via the system interface.

[0794] Server: A computer system that provides data processing and storage functions over a network. The server receives and analyzes photo data, generates 3D silhouettes, generates videos using generative AI models, runs the emotion engine, and displays and stores videos.

[0795] Image processing libraries: Image processing libraries such as OpenCV are used to analyze product photos and generate 3D silhouettes.

[0796] Generative AI models: Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF) are used to generate videos of virtual models wearing products.

[0797] Emotion engine: Analyzes the user's facial expressions, voice, and click patterns to adjust the order in which videos are presented.

[0798] Specific Examples

[0799] 1. Product photo upload example

[0800] A user (e.g., a representative from an online retailer) uses the system's interface to upload five photos of a new shirt: photos taken from the front, back, left and right, and an angle.

[0801] The terminal divides the uploaded photo into packets and sends them to the server.

[0802] 2. Product photo analysis and three-dimensional silhouette generation example

[0803] The server analyzes the received product photos and extracts feature points from each photo using OpenCV.

[0804] The server uses Structure-from-Motion technology to reconstruct the three-dimensional shape of the product based on the extracted feature points and generate a three-dimensional silhouette.

[0805] 3. Example of specifying a model and background

[0806] The user inputs the virtual model's physical information (for example, height 170 cm, weight 60 kg) and background information (for example, studio background) through an interface provided within the system.

[0807] The terminal transmits this information to the server.

[0808] 4. Example of video creation using generative AI model

[0809] The server combines the generated three-dimensional silhouette with specified model information and uses a generative AI model such as GANs or NeRF to generate a video of the virtual model wearing the product.

[0810] The server provides an interface to present the generated multiple animation versions to the user.

[0811] 5. Example of user emotion recognition using emotion engine

[0812] The device uses a camera and microphone to capture the user's facial expressions and voice while watching videos, and also records click patterns.

[0813] The device transmits the captured data in real time to a server, where an emotion engine built into the server analyzes the data.

[0814] The server adjusts the presentation order of the videos based on the analysis results.

[0815] 6. Video Presentation and Selection Examples

[0816] The server presents the generated video in the adjusted order to the user terminal.

[0817] The user selects the most suitable video from the displayed videos.

[0818] 7. Example of generating and saving the final video

[0819] The server renders the user's selected video in high resolution and generates the final video file.

[0820] The server stores the generated high-resolution video in cloud storage or on a dedicated server and provides a link for users to download it.

[0821] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present videos that are optimally tailored to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[0822] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0823] The flow of this system's program processing

[0824] Step 1:

[0825] Users upload product photos using the upload interface on the system. The input is the product photo selected by the user (front, back, left, right, oblique, etc.), and the output is a photo file sent to the server via the terminal.

[0826] Step 2:

[0827] The terminal divides the received product photos into packets and securely sends them to the server. The input is the photo data uploaded by the user, and the output is the divided photo data sent to the server. Specific operations include data packetization and encrypted communication.

[0828] Step 3:

[0829] The server saves the received photo data and organizes it in a folder for analysis. The input is the photo data sent from the terminal, and the output is the saved photo file and product information recorded in the database.

[0830] Step 4:

[0831] The server analyzes product photos using an image processing library such as OpenCV. The input is the saved product photo, and the output is the extracted feature point data. Specific operations include image reading and feature point detection.

[0832] Step 5:

[0833] The server matches feature points between multiple photos and reconstructs the three-dimensional shape of the product using Structure-from-Motion technology. The input is feature point data extracted from multiple photos, and the output is a three-dimensional silhouette (3D model) of the product. Specific operations include matching corresponding feature points, estimating camera position, and generating a point cloud.

[0834] Step 6:

[0835] The user specifies model physique information and background information through the interface. The input is the model and background information entered by the user, and the output is the model information and background information sent from the terminal.

[0836] Step 7:

[0837] The terminal transmits the specified model information and background information to the server. The input is the model information and background information input by the user, and the output is the model information and background information transmitted to the server.

[0838] Step 8:

[0839] The server integrates the model information and the 3D silhouette and uses a generative AI model to generate a video of the virtual model wearing the product. The input is the model information and the 3D silhouette, and the output is multiple generated videos. Specific operations include inputting data into the generative AI model and simulating the movement of the product.

[0840] Step 9:

[0841] The server generates an interface for proposing the generated videos to the user, where the input is the generated videos and the output is the video suggestion interface displayed on the user terminal.

[0842] Step 10:

[0843] The device collects emotional data using the user's camera and microphone and sends it to a server. The input is the user's facial expressions, voice, and click patterns, and the output is the emotional data sent to the server. Specific operations include facial expression recognition and voice analysis.

[0844] Step 11:

[0845] The server analyzes the data collected by the emotion engine and adjusts the video presentation order. The input is the transmitted emotion data, and the output is the adjusted video presentation order. Specific operations include data analysis and rearrangement of the presentation order.

[0846] Step 12:

[0847] The user selects the most suitable video from the displayed videos. The input is the video presented by the server, and the output is the selected video information.

[0848] Step 13:

[0849] The terminal transmits the user's selection information to the server, where the input is the video information selected by the user and the output is the selection information transmitted to the server.

[0850] Step 14:

[0851] The server renders the selected video in high resolution and generates and saves the final video file. The input is the selected video information, and the output is a high-resolution video file. Specific operations include high-resolution rendering and video file saving.

[0852] Step 15:

[0853] The server generates a link that allows the user to download the generated high-resolution video and provides it to the user. The input is the generated high-resolution video, and the output is the download link.

[0854] (Application example 2)

[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0856] In today's online shopping environment, users cannot actually touch the product, making it difficult to accurately grasp its appearance and feel. There are also limited ways to maximize the product's appeal and effectively communicate it. Furthermore, there is a lack of ways to recognize users' purchasing intentions and interests and make optimal suggestions based on them, creating a need to improve the user experience.

[0857] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0858] In this invention, the server includes means for receiving and saving product photos, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user, means for saving and providing a video selected by the user as a final high-resolution video, means for recognizing and analyzing the user's emotions, means for adjusting the presentation order of videos that are likely to be of most interest based on the analyzed emotional information, and means for providing an interface for the user to specify the model's physique information and background when displaying the generated videos. This allows users to accurately grasp the appearance and feel of the product in real time, and makes it possible to suggest optimal videos according to their individual interests and emotions.

[0859] "Product Photos" are images taken or provided by users that capture the exterior of a product from multiple angles.

[0860] A "three-dimensional silhouette" is a digital model that reproduces the shape and size of a product in three dimensions based on photographs of multiple products.

[0861] A "generative AI model" is an artificial intelligence model that uses deep learning techniques to generate new data from data provided by users, and includes, for example, Generative Adversarial Networks (GANs).

[0862] A "virtual model" is a digital human model created by a generative AI model based on physique information specified by the user.

[0863] "Emotion recognition" is a technology that analyzes a user's facial expressions, voice, click patterns, etc. to identify their emotional state at that time.

[0864] "High-definition video" is a video file created with a custom resolution setting that is higher than the standard resolution.

[0865] The "presentation order" is the order in which multiple generated videos and information are displayed to the user.

[0866] "Physical information" is data that represents a person's physical characteristics, such as the user's height, weight, and gender.

[0867] An "interface" is a software part that provides the screen and operating means for the user to interact with the system.

[0868] "Background" means the digital background scene or environment in which the virtual model is displayed, as selected by the user.

[0869] The system for realizing the present invention is configured using the following specific hardware and software.

[0870] Hardware:

[0871] Devices: Smartphone, camera, microphone

[0872] Server: A high-performance computer, either a cloud service or on-premise

[0873] software:

[0874] OpenCV: Image processing library

[0875] Emotion Recognition Engine: Software that recognizes user emotions in real time

[0876] Generative AI models: Deep learning techniques, such as Generative Adversarial Networks (GANs), to generate new data based on user-uploaded photos.

[0877] User Interface: The part of the software that allows the user to interact with the system.

[0878] System Details:

[0879] 1. Upload product photos

[0880] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles and upload these photos to the app, which then sends the uploaded photos to the server.

[0881] 2. Creating a three-dimensional silhouette of a product

[0882] The server analyzes the received product photos and uses image processing libraries such as OpenCV to identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product.

[0883] 3. Specifying the model and background

[0884] The user specifies his / her physical information (e.g., height, weight) and background image through an interface provided within the system. The terminal transmits this information to the server.

[0885] 4. Video Generation

[0886] Based on the three-dimensional silhouette and the specified model information, the server uses a generative AI model to generate a video in which the virtual model wears the product, taking into account the background information specified by the user.

[0887] 5. Emotional Engine Optimization

[0888] The server uses an emotion recognition engine to analyze the user's facial expressions, voice, click patterns, etc. in real time, and adjusts the order of videos presented to them based on the user's likely interest.

[0889] 6. Video Presentation and Selection

[0890] The generated videos are presented to the user in the order adjusted by the emotion engine. The user selects the most suitable video from these videos. The selected video is then sent from the device to the server.

[0891] 7. Generate and save the final video

[0892] The server finally stores the user's selected video in high resolution and makes it available for the user to download.

[0893] Program operation description:

[0894] This system operates by combining the above hardware and software. First, the device sends user input to the server, which then analyzes and processes the received data. OpenCV is used for image analysis, and a generative AI model (such as GANs) generates the video. The emotion recognition engine collects and analyzes the user's emotional data. The user interface is the means by which the user inputs and makes selections, and also presents the final video.

[0895] Examples:

[0896] For example, if a user wants to try on a new dress, the process would be as follows:

[0897] 1. The user takes photos of the dress from multiple angles and uploads them to the app.

[0898] 2. The server generates a three-dimensional silhouette of the dress based on these photos.

[0899] 3. The user specifies their physical information (e.g., height 170 cm, weight 60 kg) and a background image (e.g., a beach scene).

[0900] 4. The server uses the generative AI model to generate multiple videos of the virtual model wearing the dress.

[0901] 5. The emotion engine analyzes the user's facial expressions and click patterns to prioritize presenting the most interesting videos to the user.

[0902] 6. Users can select their favorite videos and save and share them as final high-resolution videos.

[0903] Example prompt sentence:

[0904] 1. Take photos of your product from multiple angles and upload them to the app.

[0905] 2. Select your physical information and background.

[0906] 3. Review the videos generated by the app and choose the one you like the most.

[0907] This allows users to accurately understand the appearance and feel of a product in real time, making it possible to suggest the most suitable video based on their individual interests and emotions.

[0908] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0909] Step 1:

[0910] Users take and upload product photos

[0911] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles. These photos are then uploaded to the app. The input is the photos of multiple products taken, and the output is the photo data sent to the server. Specifically, users use the app's photo upload interface, and their device sends these photos to the server.

[0912] Step 2:

[0913] The server receives and saves the product photos.

[0914] The server receives product photos sent from the device and saves them in a folder for analysis. The input is the photo data sent from the device, and the output is a photo file stored in a storage folder on the server. The server first saves the received files in a specific directory and organizes them within the folder.

[0915] Step 3:

[0916] The server analyzes the product photo and generates a three-dimensional silhouette

[0917] The server uses image processing libraries such as OpenCV to analyze the received product photos and identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product. The input is multiple stored photos, and the output is three-dimensional silhouette data. The server then runs image analysis algorithms to detect edges and extract specific features from the photos.

[0918] Step 4:

[0919] User specifies model and background information

[0920] The user specifies the model's physical information (e.g., height, weight) and background information through an interface provided within the system. The input is the user's physical information and background image selection data, and the output is that this specified information is sent to the server. The user enters information using pull-down menus and text boxes, and the terminal sends this data to the server.

[0921] Step 5:

[0922] The server generates the video

[0923] The server integrates the 3D silhouette with the specified model information and uses a generative AI model (e.g., GANs) to generate a video in which the virtual model wears the product. At this time, it also takes into account background information specified by the user. The input is the 3D silhouette, physique information, and background information, and the output is the generated video data. Specifically, the server runs the generative AI model and generates a video in which the virtual model wears the product based on the input data.

[0924] Step 6:

[0925] The server uses an emotion engine to recognize and analyze the user's emotions in real time.

[0926] The server runs an emotion recognition engine and analyzes the user's facial expressions, voice, click patterns, etc. in real time. The input is the user's emotional data (facial expressions, voice, click patterns), and the output is analyzed emotional information. The device uses a camera and microphone to acquire user data and sends it to the server. The server sends this data to the emotion engine and returns the analysis results.

[0927] Step 7:

[0928] The server adjusts the video presentation order based on emotional information

[0929] The server adjusts the presentation order of videos that the user is most likely to be interested in based on the analyzed emotional information. The input is the analyzed emotional information, and the output is the adjusted video presentation order. The server rearranges the video presentation order based on the results of the emotion engine.

[0930] Step 8:

[0931] Presentation and selection of generated videos

[0932] The server presents the videos generated in the order adjusted by the emotion engine to the user, and the user selects the most appropriate video from the displayed videos. The input is the video presented by the server, and the output is the user's selection information. The terminal displays the video to the user and sends the user's selection to the server.

[0933] Step 9:

[0934] Generate and save the final video

[0935] The server saves the user-selected video in high resolution and generates the final video file. The input is the user-selected video, and the output is the final high-resolution video file. The server saves this in cloud storage and makes it available for user download.

[0936] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0937] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0938] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0939] [Third embodiment]

[0940] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0941] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0942] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0943] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0944] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0945] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0946] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0947] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0948] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0949] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0950] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0951] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0952] The system of the present invention is a technology that allows users to upload product photos and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. The program of the system of the present invention is explained below in natural language.

[0953] System program processing explanation

[0954] 1. Upload product photos

[0955] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[0956] The terminal transmits the selected photo to the server.

[0957] The server stores the received photos and organizes them into folders for analysis.

[0958] 2. Creating a three-dimensional silhouette of a product

[0959] The server analyzes the received product photos and runs image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[0960] The server uses 3D modeling technology, such as Structure-from-Motion technology, to generate a three-dimensional silhouette of the product using the identified product shape information.

[0961] 3. Specifying the model and background

[0962] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[0963] The terminal transmits this specification information to the server.

[0964] 4. Video creation using generative AI models

[0965] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[0966] The server generates an interface for suggesting the generated multiple videos to the user.

[0967] 5. Select a video

[0968] The user reviews the suggested videos and selects the most suitable one.

[0969] The terminal transmits the user's selection information to the server.

[0970] 6. Generate and save the final video

[0971] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[0972] The server stores the generated high-resolution video therein so that the user can download it, and provides it to the user.

[0973] Specific examples

[0974] Example 1: Creating a shirt video

[0975] 1. A user (retailer) uploads five photos of a new shirt to the system, including photos from the front, back, left and right, and one-quarter angles.

[0976] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[0977] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[0978] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[0979] 5. The user selects the most suitable video from the suggested videos.

[0980] 6. The server stores the selected video in high resolution and makes it available for download by the user.

[0981] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills, thereby effectively supporting the promotion of fashion merchandise purchases.

[0982] The processing flow will be explained below.

[0983] Step 1:

[0984] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[0985] Step 2:

[0986] The terminal transmits the selected photo to the server.

[0987] Step 3:

[0988] The server stores the received photo files and organizes them for further processing.

[0989] Step 4:

[0990] The server analyzes the stored product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[0991] Step 5:

[0992] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[0993] Step 6:

[0994] The user uses the system interface to specify the model's physical information (e.g., height, weight) and background information.

[0995] Step 7:

[0996] The terminal transmits the information specified by the user (physique information and background information of the model) to the server.

[0997] Step 8:

[0998] The server integrates the 3D silhouette with the user-specified information, uses a generative AI model to dress the virtual model in the product, and generates a video against the specified background, using generative technologies such as GANs and NeRF.

[0999] Step 9:

[1000] The server generates an interface for suggesting the generated multiple videos to the user.

[1001] Step 10:

[1002] The terminal displays the suggested videos to the user.

[1003] Step 11:

[1004] The user selects the most suitable video from the displayed videos.

[1005] Step 12:

[1006] The terminal transmits the user's selection information to the server.

[1007] Step 13:

[1008] The server processes the user's selected video in high resolution and generates the final video file.

[1009] Step 14:

[1010] The server stores and provides the generated high-resolution video for users to download.

[1011] Step 15:

[1012] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[1013] Example 1

[1014] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1015] Traditional online shopping has the problem that customers cannot actually try on products, which can cause anxiety about the size and fit of the product. Furthermore, product photos and simple descriptions alone make it difficult to fully check the product's details and actual appearance, which can lead to hesitation in purchasing. Furthermore, the need for specialized knowledge and skills to take product photos and create videos places a significant burden on product sellers.

[1016] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1017] In this invention, the server includes a means for receiving and saving product photos, a means for analyzing the photos to generate a three-dimensional silhouette of the product, a means for generating a video of a virtual model wearing the product using a generative AI model based on the generated silhouette, a means for presenting multiple videos to the user, and a means for saving and providing the video selected by the user as a final high-resolution video. This allows users to check the appearance and fit of products in detail without actually trying them on, eliminating any anxiety about purchasing. Furthermore, product sellers can create high-quality product promotional videos without specialized knowledge of photography or editing, leading to increased sales.

[1018] "Product photos" are image data of products uploaded by users.

[1019] "Means for storing" refers to the system's function of storing received product photos in a database or storage.

[1020] The "means for analyzing" refers to a function of the system that executes an image processing algorithm to identify the contours and shape of the product using a product photograph and generate a three-dimensional silhouette.

[1021] A "three-dimensional silhouette" is three-dimensional shape information of a product generated from a product photo.

[1022] A "generative AI model" is an algorithm that uses deep learning technology to generate new data based on input data, and in this invention is used to generate videos of virtual models.

[1023] A "virtual model" is a model of a person or object that is digitally generated using a generative AI model.

[1024] The "means for generating video" is a system function that creates video data based on the generated three-dimensional silhouette and virtual model.

[1025] The "presentation means" is a system function that displays the generated multiple videos so that the user can check them.

[1026] "High-resolution video" refers to high-quality video data that can display every detail clearly.

[1027] The system of the present invention is a technology that allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Specific embodiments of the system of the present invention are described below.

[1028] 1. Upload product photos:

[1029] Users take photos of the items they want to sell using a digital camera or smartphone.

[1030] The user logs into the system's web interface and selects a product photo in the upload window.

[1031] The terminal transmits the selected photo data to the server using an HTTP request.

[1032] The server stores the received photos in a database and organizes them into folders for analysis, using a cloud storage service such as Amazon S3.

[1033] 2. Generate a 3D silhouette of the product:

[1034] The server analyzes the stored photos using an image processing algorithm (e.g., OpenCV).

[1035] The server uses the analysis results to identify the contours and shape of the product, using edge detection and segmentation techniques in the process.

[1036] The server uses Structure-from-Motion (SfM) technology to generate a three-dimensional silhouette of the product based on the identified shape information.

[1037] The server saves the generated three-dimensional silhouette data as a 3D model file in preparation for the next process.

[1038] 3. Specify the model and background:

[1039] The user inputs the desired model's physical information (e.g., height, weight, etc.) and background setting (e.g., studio background or outdoor background) through the system interface.

[1040] The terminal sends the setting information entered by the user to the server in JSON format.

[1041] The server saves the received setting information in a file and prepares for the next video generation process based on this information.

[1042] 4. Video creation using generative AI models:

[1043] The server integrates the three-dimensional silhouette data with the model's physique information and background settings specified by the user.

[1044] The server uses deep learning libraries (e.g., TensorFlow and PyTorch) to run generative AI models (e.g., Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF)).

[1045] The server inputs a prompt statement (e.g., "Generate a video of the person wearing a shirt and place it against a specified background") into the generative AI model, causing it to generate multiple videos.

[1046] The server displays the generated videos on a web interface and asks the user for confirmation.

[1047] 5. Video Selection:

[1048] The user can view the suggested videos on a web interface and select the one they like best.

[1049] The terminal captures the user's selection actions and transmits the information to the server.

[1050] The server processes the selected video data again at high resolution and saves it as the final video.

[1051] 6. Generate and save the final video:

[1052] The server encodes the selected video in high resolution and generates the final video file, using libraries such as FFmpeg.

[1053] The server stores the generated high-resolution video in cloud storage such as Amazon S3 and generates a link that allows users to download it.

[1054] The server notifies the user of the download link and provides it through a web interface.

[1055] Example: Creating a shirt video

[1056] 1. A user (retailer) uploads five photos of a new shirt to the system: front, back, left, right, and right-angle views.

[1057] 2. The server analyzes the five received photos, identifies the outline of the shirt using OpenCV, and generates a three-dimensional silhouette of the shirt using SfM technology.

[1058] 3. The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background) into the system's interface.

[1059] 4. The server uses the generative AI model to generate multiple videos based on the specified criteria and suggests those videos to the user on a web interface.

[1060] 5. The user selects the most suitable video from the suggested videos and sends that information to the server.

[1061] 6. The server regenerates the selected video in high resolution, stores it on Amazon S3, and then provides a link for the user to download it.

[1062] An example of a prompt sentence for this invention is, "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background." Using this prompt sentence, the generative AI model can generate the optimal video based on the required information.

[1063] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1064] Step 1: Upload your product photos

[1065] The user takes photos of the product they want to sell (for example, photos taken from the front, back, left and right, and diagonal angles), which are used as input data.

[1066] The user logs into the system's web interface, selects a photo of the product, and clicks the upload button.

[1067] The terminal transmits the selected photo data to the server using an HTTP request.

[1068] The server stores the received photo data using a cloud storage service such as Amazon S3.

[1069] Output: Image data stored in a folder for analysis, as the photos are saved on the server.

[1070] Step 2: Generate a 3D silhouette of the product

[1071] The server runs image processing algorithms such as OpenCV on the stored photos to analyze the contours and shapes of the products, using edge detection and segmentation techniques. This analysis serves as input data.

[1072] Based on the analyzed data, the server generates a 3D silhouette of the product using Structure-from-Motion (SfM) technology, a process that reconstructs a three-dimensional shape from multiple photographs.

[1073] Output: Three-dimensional silhouette data (3D model file) of the generated product.

[1074] Step 3: Specifying the model and background

[1075] The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background or outdoors) on the system interface. This becomes the input data.

[1076] The terminal converts the information entered by the user into JSON format and sends it to the server using an HTTP request.

[1077] The server saves the received JSON data to a file and uses this information when running the generative AI model.

[1078] Output: A JSON file with the saved model's physique and background information.

[1079] Step 4: Creating videos with generative AI models

[1080] The server takes in and integrates the three-dimensional silhouette data with the model and background information specified by the user. This integration becomes the input data.

[1081] The server starts running generative AI models (Generative Adversarial Networks (GANs) or Neural Radiance Fields (NeRF)) using deep learning libraries (e.g., TensorFlow or PyTorch).

[1082] The server inputs a prompt statement (e.g., "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background.") into the generative AI model and generates multiple videos.

[1083] Output: Multiple generated video files.

[1084] Step 5: Select a video

[1085] The user reviews the proposed videos on the system's web interface and selects the most suitable one, which becomes the input data.

[1086] The device captures information about the user's selected video, converts it into JSON format, and sends it to the server in an HTTP request.

[1087] The server reprocesses the selected video data in its final high resolution and prepares it for storage.

[1088] Output: The video data selected by the user.

[1089] Step 6: Generate and save the final video

[1090] The server uses a library such as FFmpeg to encode the selected video in high resolution and generate the final video file, which becomes the input data.

[1091] The server stores the high-resolution video in cloud storage such as Amazon S3 and generates a link that users can use to download it.

[1092] The server notifies the user of the download link through a web interface.

[1093] Output: A link to a high-resolution video file that users can download.

[1094] (Application example 1)

[1095] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1096] Traditional try-on experiences in brick-and-mortar stores required customers to physically try on the products to check them, which was time-consuming and labor-intensive. Also, trying on multiple products required customers to use a fitting room and take the products in and out, which was inefficient for both customers and stores. Furthermore, if a particular size or model was not in stock at the store, customers were unable to try on that product. To solve these problems, new technology was needed to enable virtual try-on in real time.

[1097] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1098] In this invention, the server includes means for receiving and storing product images, means for analyzing the received product images to generate a three-dimensional skeleton of the product, and means for generating a motion video in which the product is applied to a virtual model using a generative AI model based on the generated three-dimensional skeleton. This allows users to obtain product images in a physical store using a mobile device or visual aid device and have a virtual try-on experience in real time. This eliminates the hassle of trying on products and allows users to efficiently try on many products, thereby increasing customer purchasing motivation beyond the constraints of stock and size.

[1099] "Product images" are photos or image data of products taken or scanned by users.

[1100] A "three-dimensional skeleton" is three-dimensional shape information of a product generated from two-dimensional image data.

[1101] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate virtual fitting videos from image data.

[1102] "Motion video" is a video that shows a virtual model moving while wearing the product.

[1103] A "mobile terminal" is a portable information processing terminal such as a smartphone or tablet.

[1104] A "visual aid" is a device that assists with visual information, such as smart glasses or a head-mounted display.

[1105] The "user interface" refers to a screen or input device that allows the user to interact with the system, and includes functions for specifying body type information and background.

[1106] "High-definition video" refers to video data that has high resolution and quality and contains detailed information.

[1107] "Real-time" is a term that refers to processing occurring immediately, with little time delay.

[1108] "Virtual try-on experience" refers to the use of digital technology to provide the experience of trying on products without actually trying them on.

[1109] "Inventory" refers to the quantity and variety of products stored in a store or warehouse.

[1110] This technology allows users to upload product images, and generates a moving video of a virtual model wearing the product based on a three-dimensional skeleton generated from the image. To realize this technology, the following steps and system configuration are required.

[1111] Uploading product images

[1112] Users take or scan product images in physical stores using a mobile device or visual aid. The device sends the captured image to a server, which then stores it. Image processing libraries such as OpenCV are used to identify the position and shape of the product image.

[1113] Three-dimensional skeleton generation

[1114] The server analyzes the stored product images and generates a three-dimensional skeleton of the product, using Structure-from-Motion technology to extract the product's three-dimensional shape information from multiple image data.

[1115] Generation of motion images

[1116] Next, the server uses a generative AI model (e.g., GANs or NeRF) based on the generated three-dimensional skeleton to generate a moving image of the product applied to the virtual model. The user can specify the model's physical information (e.g., height, weight) and background information through the system's user interface.

[1117] Presentation and selection of motion images

[1118] The server generates an interface for presenting the generated motion images to the user, who can then review the motion images and select the most suitable one.

[1119] Final high-definition video generation and storage

[1120] The server stores the video selected by the user as high-definition video and allows the user to download it, providing a real-time virtual try-on experience and a comfortable shopping experience.

[1121] Hardware and software used

[1122] Hardware: mobile devices, smart glasses, head-mounted displays, servers

[1123] Software: OpenCV, Structure-from-Motion library, GANs, NeRF

[1124] Specific example explanation

[1125] As an example, consider the case where a user wants to virtually try on a shirt in a physical store.

[1126] 1. The user (customer) takes five photos of a shirt with their smartphone in a physical store (from the front, back, left and right, and diagonal).

[1127] 2. The device sends the captured photo to the server.

[1128] 3. The server analyzes the received photo and generates a three-dimensional skeleton of the shirt.

[1129] 4. The user specifies the model's physical information (e.g., height 170 cm, weight 60 kg) and the background (e.g., studio background).

[1130] 5. The server uses the generative AI model to generate multiple motion videos of the shirt being worn by the specified virtual model and suggests them to the user.

[1131] 6. The user selects the most suitable video from the suggested videos.

[1132] 7. The server stores the selected footage as high definition footage and makes it available for user download.

[1133] Prompt Sentence Examples

[1134] Photo path = upload_photo('Shirt photo taken in store.jpg')

[1135] Silhouette = generate_3d_silhouette(photopath)

[1136] User information = {'height': 170, 'weight': 60} The user enters their body type information.

[1137] background = 'store background' Use the store background as an example

[1138] Video = create_virtual_model(silhouette, user information, background)

[1139] save_video(video, 'Virtual Try-On Shirt 2023.mp4')

[1140] By following the above-described procedure, the present invention can generate a three-dimensional skeleton from a product image and provide a virtual try-on experience in real time.

[1141] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1142] Step 1:

[1143] Uploading product images

[1144] A user takes product images in a physical store using a mobile device or visual aid. The device sends the captured images to a server, which then stores the received images. Specifically, the user takes photos from multiple angles and the device uploads the image data to the server. The input is the captured product image, and the output is the image data stored on the server.

[1145] Step 2:

[1146] Three-dimensional skeleton generation

[1147] The server analyzes the stored product images and generates a three-dimensional skeleton of the product. Image processing libraries such as OpenCV are used for the analysis, and Structure-from-Motion technology is applied to extract the product's three-dimensional shape information from multiple image data. The input is the stored product image, and the output is three-dimensional skeleton data. Specifically, an image processing algorithm is executed on the server to measure the shape and size of the product.

[1148] Step 3:

[1149] Specifying the model and background

[1150] The user uses a user interface within the system to specify the model's physical information (e.g., height, weight) and background information. The input is the physical and background information entered by the user, and the output is a dataset containing that information. Specifically, the user enters the required information using drop-down menus and input forms in the user interface, and the information is sent to the server.

[1151] Step 4:

[1152] Virtual model generation

[1153] The server integrates the generated 3D skeleton with the model information specified by the user and uses a generative AI model (e.g., GANs or NeRF) to generate a moving image of the product applied to the virtual model. The input is the 3D skeleton data, model information, and background information, and the output is multiple moving images. Specifically, the generative AI model simulates the product's movement and executes the process of applying it to the virtual model.

[1154] Step 5:

[1155] Presentation and selection of motion images

[1156] The server generates an interface to present the generated multiple motion videos to the user. The user reviews the presented motion videos and selects the most suitable one. The input at this time is the generated motion video, and the output is the motion video selected by the user. Specifically, the user watches multiple videos on the interface and performs an operation to select the best one.

[1157] Step 6:

[1158] Final high-definition video generation and storage

[1159] The server saves the video selected by the user as high-definition video and makes it available for download. The input is the motion video selected by the user, and the output is the final high-definition video data. Specifically, the selected video is regenerated in high resolution on the server and provided to the user. During this process, the final video data is saved in a dedicated folder on the server.

[1160] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1161] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, the present invention enables more effective video generation and presentation.

[1162] System program processing explanation

[1163] 1. Upload product photos

[1164] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[1165] The terminal transmits the selected photo to the server.

[1166] The server stores the received photos and organizes them into folders for analysis.

[1167] 2. Creating a three-dimensional silhouette of a product

[1168] The server analyzes the received product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[1169] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[1170] 3. Specifying the model and background

[1171] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[1172] The terminal transmits this specification information to the server.

[1173] 4. Video creation using generative AI models

[1174] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[1175] The server generates an interface for suggesting the generated multiple videos to the user.

[1176] 5. Recognition of user emotions by emotion engine

[1177] The server runs an emotion engine that recognizes the user's emotions in real time during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc.

[1178] The device uses the user's camera and microphone to send emotion data to the emotion engine.

[1179] Based on the analyzed emotional information, the emotion engine identifies the videos that the user is most likely to be interested in and changes the order in which they are presented.

[1180] 6. Video Presentation and Selection

[1181] The server presents the video generated in the order adjusted by the emotion engine to the user.

[1182] The terminal displays the suggested videos to the user, and the user selects the most suitable video from the displayed videos.

[1183] The terminal transmits the user's selection information to the server.

[1184] 7. Generate and save the final video

[1185] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[1186] The server stores and provides the generated high-resolution video for users to download.

[1187] Specific examples

[1188] Example 1: Creating a shirt video

[1189] 1. A user (retailer) uploads five photos of a new shirt to the system, including front, back, left and right views, and three-quarter views.

[1190] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[1191] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[1192] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[1193] 5. The server runs an emotion engine during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc., and adjusts the presentation order.

[1194] 6. The user selects the most suitable video from the suggested videos.

[1195] 7. The server stores the selected video in high resolution and makes it available for download by the user.

[1196] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present optimal videos according to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[1197] The processing flow will be explained below.

[1198] Step 1:

[1199] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[1200] Step 2:

[1201] The terminal transmits the selected photo file to the server.

[1202] Step 3:

[1203] The server stores the received photo files and organizes them into folders for further processing.

[1204] Step 4:

[1205] The server opens the stored product photos and uses image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[1206] Step 5:

[1207] The server then integrates the shape information from the analyzed photos to generate a three-dimensional silhouette of the product, sometimes using Structure-from-Motion technology.

[1208] Step 6:

[1209] The user specifies the model's physical information (e.g., height, weight) and background information through an interface within the system.

[1210] Step 7:

[1211] The terminal transmits the physique information and background information input by the user to the server.

[1212] Step 8:

[1213] The server uses a generative AI model (e.g., GANs or NeRF) to generate a video of a virtual model wearing the product based on the three-dimensional silhouette and the physique and background information specified by the user.

[1214] Step 9:

[1215] The server generates an interface for suggesting the generated multiple videos to the user.

[1216] Step 10:

[1217] The server runs an emotion engine to acquire emotion data in real time from the user's camera and microphone while displaying the suggested video.

[1218] Step 11:

[1219] The terminal inputs the user's facial expressions, voice, click patterns, etc. into an emotion engine and analyzes the emotion data.

[1220] Step 12:

[1221] The emotion engine identifies the user's emotional state based on the analyzed emotion information, identifies the videos that the user is most interested in, and changes the presentation order of those videos.

[1222] Step 13:

[1223] The server presents the multiple videos to the user in an order adjusted by the emotion engine.

[1224] Step 14:

[1225] The terminal displays the videos to the user in the adjusted order, and the user selects the most suitable video from the suggested videos.

[1226] Step 15:

[1227] The terminal transmits the user's selection information to the server.

[1228] Step 16:

[1229] The server processes the video selected by the user as a high-definition video and generates the final video file.

[1230] Step 17:

[1231] The server stores and provides the generated high-resolution video for users to download.

[1232] Step 18:

[1233] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[1234] Example 2

[1235] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1236] Conventional systems have had difficulty generating a three-dimensional silhouette using product photos and creating videos in which a virtual model wears the product. Furthermore, the order in which the generated videos are presented does not take into account the user's interests or emotions, which makes it difficult to efficiently select the most suitable video. Furthermore, the interface for users to specify the model's physique and background information is inadequate, making it difficult to customize the system to meet individual needs.

[1237] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1238] In this invention, the server includes means for receiving and saving product photos via a user terminal, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user terminal, means for saving and providing a video selected by the user based on selection information transmitted from the user terminal as a final high-resolution video, and means for adjusting the video presentation order, which includes an emotion engine that collects and analyzes user emotion data using the camera and microphone of the user terminal. This significantly improves the user experience, enabling optimal video presentation and customized video generation according to individual needs.

[1239] "User terminal" refers to any device that a user uses to connect to the Internet and send and receive data.

[1240] A "server" is a computer system that provides data processing and storage functions over a network.

[1241] "Product Photos" refers to product image data uploaded by users.

[1242] "Three-dimensional silhouette" refers to a three-dimensional shape model of a product generated based on a product photo.

[1243] A "generative AI model" is a model that uses artificial intelligence algorithms to generate data to perform a specific task (in this case, video generation).

[1244] "Virtual Model" refers to a digital sculpture or animation of a digitally generated figure or form that simulates a tangible product or scene.

[1245] "High-definition video" refers to video files produced at high-quality image resolution.

[1246] An "emotion engine" refers to technology and algorithms that analyze a user's facial expressions, voice, behavioral patterns, etc. to recognize emotions in real time.

[1247] "Means for adjusting the presentation order" refers to a function that rearranges the order in which videos are displayed based on the analysis results of the emotion engine.

[1248] MODE FOR CARRYING OUT THE INVENTION

[1249] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. By combining this with an emotion engine that recognizes the user's emotions, the system presents the most appropriate video to the user.

[1250] The system of the present invention includes the following hardware and software.

[1251] User terminal: This can be a smartphone, tablet, PC, or other device. The user terminal connects to the Internet and sends and receives data via the system interface.

[1252] Server: A computer system that provides data processing and storage functions over a network. The server receives and analyzes photo data, generates 3D silhouettes, generates videos using generative AI models, runs the emotion engine, and displays and stores videos.

[1253] Image processing libraries: Image processing libraries such as OpenCV are used to analyze product photos and generate 3D silhouettes.

[1254] Generative AI models: Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF) are used to generate videos of virtual models wearing products.

[1255] Emotion engine: Analyzes the user's facial expressions, voice, and click patterns to adjust the order in which videos are presented.

[1256] Specific Examples

[1257] 1. Product photo upload example

[1258] A user (e.g., a representative from an online retailer) uses the system's interface to upload five photos of a new shirt: photos taken from the front, back, left and right, and an angle.

[1259] The terminal divides the uploaded photo into packets and sends them to the server.

[1260] 2. Product photo analysis and three-dimensional silhouette generation example

[1261] The server analyzes the received product photos and extracts feature points from each photo using OpenCV.

[1262] The server uses Structure-from-Motion technology to reconstruct the three-dimensional shape of the product based on the extracted feature points and generate a three-dimensional silhouette.

[1263] 3. Example of specifying a model and background

[1264] The user inputs the virtual model's physical information (for example, height 170 cm, weight 60 kg) and background information (for example, studio background) through an interface provided within the system.

[1265] The terminal transmits this information to the server.

[1266] 4. Example of video creation using generative AI model

[1267] The server combines the generated three-dimensional silhouette with specified model information and uses a generative AI model such as GANs or NeRF to generate a video of the virtual model wearing the product.

[1268] The server provides an interface to present the generated multiple animation versions to the user.

[1269] 5. Example of user emotion recognition using emotion engine

[1270] The device uses a camera and microphone to capture the user's facial expressions and voice while watching videos, and also records click patterns.

[1271] The device transmits the captured data in real time to a server, where an emotion engine built into the server analyzes the data.

[1272] The server adjusts the presentation order of the videos based on the analysis results.

[1273] 6. Video presentation and selection examples

[1274] The server presents the generated video in the adjusted order to the user terminal.

[1275] The user selects the most suitable video from the displayed videos.

[1276] 7. Example of generating and saving the final video

[1277] The server renders the user's selected video in high resolution and generates the final video file.

[1278] The server stores the generated high-resolution video in cloud storage or on a dedicated server and provides a link for users to download it.

[1279] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present optimal videos according to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[1280] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1281] The flow of this system's program processing

[1282] Step 1:

[1283] Users upload product photos using the upload interface on the system. The input is the product photo selected by the user (front, back, left, right, oblique, etc.), and the output is a photo file sent to the server via the terminal.

[1284] Step 2:

[1285] The terminal divides the received product photos into packets and securely sends them to the server. The input is the photo data uploaded by the user, and the output is the divided photo data sent to the server. Specific operations include data packetization and encrypted communication.

[1286] Step 3:

[1287] The server saves the received photo data and organizes it in a folder for analysis. The input is the photo data sent from the terminal, and the output is the saved photo file and product information recorded in the database.

[1288] Step 4:

[1289] The server analyzes product photos using an image processing library such as OpenCV. The input is the saved product photo, and the output is the extracted feature point data. Specific operations include image reading and feature point detection.

[1290] Step 5:

[1291] The server matches feature points between multiple photos and reconstructs the three-dimensional shape of the product using Structure-from-Motion technology. The input is feature point data extracted from multiple photos, and the output is a three-dimensional silhouette (3D model) of the product. Specific operations include matching corresponding feature points, estimating camera position, and generating a point cloud.

[1292] Step 6:

[1293] The user specifies model physique information and background information through the interface. The input is the model and background information entered by the user, and the output is the model information and background information sent from the terminal.

[1294] Step 7:

[1295] The terminal transmits the specified model information and background information to the server. The input is the model information and background information input by the user, and the output is the model information and background information transmitted to the server.

[1296] Step 8:

[1297] The server integrates the model information and the 3D silhouette and uses a generative AI model to generate a video of the virtual model wearing the product. The input is the model information and the 3D silhouette, and the output is multiple generated videos. Specific operations include inputting data into the generative AI model and simulating the movement of the product.

[1298] Step 9:

[1299] The server generates an interface for proposing the generated videos to the user, where the input is the generated videos and the output is the video suggestion interface displayed on the user terminal.

[1300] Step 10:

[1301] The device collects emotional data using the user's camera and microphone and sends it to a server. The input is the user's facial expressions, voice, and click patterns, and the output is the emotional data sent to the server. Specific operations include facial expression recognition and voice analysis.

[1302] Step 11:

[1303] The server analyzes the data collected by the emotion engine and adjusts the video presentation order. The input is the transmitted emotion data, and the output is the adjusted video presentation order. Specific operations include data analysis and rearrangement of the presentation order.

[1304] Step 12:

[1305] The user selects the most suitable video from the displayed videos. The input is the video presented by the server, and the output is the selected video information.

[1306] Step 13:

[1307] The terminal transmits the user's selection information to the server, where the input is the video information selected by the user and the output is the selection information transmitted to the server.

[1308] Step 14:

[1309] The server renders the selected video in high resolution and generates and saves the final video file. The input is the selected video information, and the output is a high-resolution video file. Specific operations include high-resolution rendering and video file saving.

[1310] Step 15:

[1311] The server generates a link that allows the user to download the generated high-resolution video and provides it to the user. The input is the generated high-resolution video, and the output is the download link.

[1312] (Application example 2)

[1313] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1314] In today's online shopping environment, users cannot actually touch the product, making it difficult to accurately grasp its appearance and feel. There are also limited ways to maximize the product's appeal and effectively communicate it. Furthermore, there is a lack of ways to recognize users' purchasing intentions and interests and make optimal suggestions based on them, creating a need to improve the user experience.

[1315] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1316] In this invention, the server includes means for receiving and saving product photos, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user, means for saving and providing a video selected by the user as a final high-resolution video, means for recognizing and analyzing the user's emotions, means for adjusting the presentation order of videos that are likely to be of most interest based on the analyzed emotional information, and means for providing an interface for the user to specify the model's physique information and background when displaying the generated videos. This allows users to accurately grasp the appearance and feel of the product in real time, and makes it possible to suggest optimal videos according to their individual interests and emotions.

[1317] "Product Photos" are images taken or provided by users that capture the exterior of a product from multiple angles.

[1318] A "three-dimensional silhouette" is a digital model that reproduces the shape and size of a product in three dimensions based on photographs of multiple products.

[1319] A "generative AI model" is an artificial intelligence model that uses deep learning techniques to generate new data from data provided by users, and includes, for example, Generative Adversarial Networks (GANs).

[1320] A "virtual model" is a digital human model created by a generative AI model based on physique information specified by the user.

[1321] "Emotion recognition" is a technology that analyzes a user's facial expressions, voice, click patterns, etc. to identify their emotional state at that time.

[1322] "High-definition video" is a video file created with a custom resolution setting that is higher than the standard resolution.

[1323] The "presentation order" is the order in which multiple generated videos and information are displayed to the user.

[1324] "Physical information" is data that represents a person's physical characteristics, such as the user's height, weight, and gender.

[1325] An "interface" is a software part that provides the screen and operating means for the user to interact with the system.

[1326] "Background" means the digital background scene or environment in which the virtual model is displayed, as selected by the user.

[1327] The system for realizing the present invention is configured using the following specific hardware and software.

[1328] Hardware:

[1329] Devices: Smartphone, camera, microphone

[1330] Server: A high-performance computer, either a cloud service or on-premise

[1331] software:

[1332] OpenCV: Image processing library

[1333] Emotion Recognition Engine: Software that recognizes user emotions in real time

[1334] Generative AI models: Deep learning techniques, such as Generative Adversarial Networks (GANs), to generate new data based on user-uploaded photos.

[1335] User Interface: The part of the software that allows the user to interact with the system.

[1336] System Details:

[1337] 1. Upload product photos

[1338] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles and upload these photos to the app, which then sends the uploaded photos to the server.

[1339] 2. Creating a three-dimensional silhouette of a product

[1340] The server analyzes the received product photos and uses image processing libraries such as OpenCV to identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product.

[1341] 3. Specifying the model and background

[1342] The user specifies his / her physical information (e.g., height, weight) and background image through an interface provided within the system. The terminal transmits this information to the server.

[1343] 4. Video Generation

[1344] Based on the three-dimensional silhouette and the specified model information, the server uses a generative AI model to generate a video in which the virtual model wears the product, taking into account the background information specified by the user.

[1345] 5. Emotional Engine Optimization

[1346] The server uses an emotion recognition engine to analyze the user's facial expressions, voice, click patterns, etc. in real time, and adjusts the order of videos presented to them based on the user's likely interest.

[1347] 6. Video Presentation and Selection

[1348] The generated videos are presented to the user in the order adjusted by the emotion engine. The user selects the most suitable video from these videos. The selected video is then sent from the device to the server.

[1349] 7. Generate and save the final video

[1350] The server finally stores the user's selected video in high resolution and makes it available for the user to download.

[1351] Program operation description:

[1352] This system operates by combining the above hardware and software. First, the device sends user input to the server, which then analyzes and processes the received data. OpenCV is used for image analysis, and a generative AI model (such as GANs) generates the video. The emotion recognition engine collects and analyzes the user's emotional data. The user interface is the means by which the user inputs and makes selections, and also presents the final video.

[1353] Examples:

[1354] For example, if a user wants to try on a new dress, the process would be as follows:

[1355] 1. The user takes photos of the dress from multiple angles and uploads them to the app.

[1356] 2. The server generates a three-dimensional silhouette of the dress based on these photos.

[1357] 3. The user specifies their physical information (e.g., height 170 cm, weight 60 kg) and a background image (e.g., a beach scene).

[1358] 4. The server uses the generative AI model to generate multiple videos of the virtual model wearing the dress.

[1359] 5. The emotion engine analyzes the user's facial expressions and click patterns to prioritize presenting the most interesting videos to the user.

[1360] 6. Users can select their favorite videos and save and share them as final high-resolution videos.

[1361] Example prompt sentence:

[1362] 1. Take photos of your product from multiple angles and upload them to the app.

[1363] 2. Select your physical information and background.

[1364] 3. Review the videos generated by the app and choose the one you like the most.

[1365] This allows users to accurately understand the appearance and feel of a product in real time, making it possible to suggest the most suitable video based on their individual interests and emotions.

[1366] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1367] Step 1:

[1368] Users take and upload product photos

[1369] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles. These photos are then uploaded to the app. The input is the photos of multiple products taken, and the output is the photo data sent to the server. Specifically, users use the app's photo upload interface, and their device sends these photos to the server.

[1370] Step 2:

[1371] The server receives and saves the product photos.

[1372] The server receives product photos sent from the device and saves them in a folder for analysis. The input is the photo data sent from the device, and the output is a photo file stored in a storage folder on the server. The server first saves the received files in a specific directory and organizes them within the folder.

[1373] Step 3:

[1374] The server analyzes the product photo and generates a three-dimensional silhouette

[1375] The server uses image processing libraries such as OpenCV to analyze the received product photos and identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product. The input is multiple stored photos, and the output is three-dimensional silhouette data. The server then runs image analysis algorithms to detect edges and extract specific features from the photos.

[1376] Step 4:

[1377] User specifies model and background information

[1378] The user specifies the model's physical information (e.g., height, weight) and background information through an interface provided within the system. The input is the user's physical information and background image selection data, and the output is that this specified information is sent to the server. The user enters information using pull-down menus and text boxes, and the terminal sends this data to the server.

[1379] Step 5:

[1380] The server generates the video

[1381] The server integrates the 3D silhouette with the specified model information and uses a generative AI model (e.g., GANs) to generate a video in which the virtual model wears the product. At this time, it also takes into account background information specified by the user. The input is the 3D silhouette, physique information, and background information, and the output is the generated video data. Specifically, the server runs the generative AI model and generates a video in which the virtual model wears the product based on the input data.

[1382] Step 6:

[1383] The server uses an emotion engine to recognize and analyze the user's emotions in real time.

[1384] The server runs an emotion recognition engine and analyzes the user's facial expressions, voice, click patterns, etc. in real time. The input is the user's emotional data (facial expressions, voice, click patterns), and the output is analyzed emotional information. The device uses a camera and microphone to acquire user data and sends it to the server. The server sends this data to the emotion engine and returns the analysis results.

[1385] Step 7:

[1386] The server adjusts the video presentation order based on emotional information

[1387] The server adjusts the presentation order of videos that the user is most likely to be interested in based on the analyzed emotional information. The input is the analyzed emotional information, and the output is the adjusted video presentation order. The server rearranges the video presentation order based on the results of the emotion engine.

[1388] Step 8:

[1389] Presentation and selection of generated videos

[1390] The server presents the videos generated in the order adjusted by the emotion engine to the user, and the user selects the most appropriate video from the displayed videos. The input is the video presented by the server, and the output is the user's selection information. The terminal displays the video to the user and sends the user's selection to the server.

[1391] Step 9:

[1392] Generate and save the final video

[1393] The server saves the user-selected video in high resolution and generates the final video file. The input is the user-selected video, and the output is the final high-resolution video file. The server saves this in cloud storage and makes it available for user download.

[1394] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1395] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1396] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1397] [Fourth embodiment]

[1398] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1399] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1400] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1401] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1402] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1404] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1405] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1406] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1407] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1408] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1409] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1410] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1411] The system of the present invention is a technology that allows users to upload product photos and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. The program of the system of the present invention is explained below in natural language.

[1412] System program processing explanation

[1413] 1. Upload product photos

[1414] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[1415] The terminal transmits the selected photo to the server.

[1416] The server stores the received photos and organizes them into folders for analysis.

[1417] 2. Creating a three-dimensional silhouette of a product

[1418] The server analyzes the received product photos and runs image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[1419] The server uses 3D modeling technology to generate a three-dimensional silhouette of the product using the identified product shape information, and this modeling applies a structure-from-motion technology or the like.

[1420] 3. Specifying the model and background

[1421] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[1422] The terminal transmits this specification information to the server.

[1423] 4. Video creation using generative AI models

[1424] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[1425] The server generates an interface for suggesting the generated multiple videos to the user.

[1426] 5. Select a video

[1427] The user reviews the suggested videos and selects the most suitable one.

[1428] The terminal transmits the user's selection information to the server.

[1429] 6. Generate and save the final video

[1430] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[1431] The server stores the generated high-resolution video therein so that the user can download it, and provides it to the user.

[1432] Specific examples

[1433] Example 1: Creating a shirt video

[1434] 1. A user (retailer) uploads five photos of a new shirt to the system, including photos from the front, back, left and right, and one-quarter angles.

[1435] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[1436] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[1437] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[1438] 5. The user selects the most suitable video from the suggested videos.

[1439] 6. The server stores the selected video in high resolution and makes it available for download by the user.

[1440] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills, thereby effectively supporting the promotion of fashion merchandise purchases.

[1441] The processing flow will be explained below.

[1442] Step 1:

[1443] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[1444] Step 2:

[1445] The terminal transmits the selected photo to the server.

[1446] Step 3:

[1447] The server stores the received photo files and organizes them for further processing.

[1448] Step 4:

[1449] The server analyzes the stored product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[1450] Step 5:

[1451] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[1452] Step 6:

[1453] The user uses the system interface to specify the model's physical information (e.g., height, weight) and background information.

[1454] Step 7:

[1455] The terminal transmits the information specified by the user (physique information and background information of the model) to the server.

[1456] Step 8:

[1457] The server integrates the 3D silhouette with the user-specified information, uses a generative AI model to dress the virtual model in the product, and generates a video against the specified background, using generative technologies such as GANs and NeRF.

[1458] Step 9:

[1459] The server generates an interface for suggesting the generated multiple videos to the user.

[1460] Step 10:

[1461] The terminal displays the suggested videos to the user.

[1462] Step 11:

[1463] The user selects the most suitable video from the displayed videos.

[1464] Step 12:

[1465] The terminal transmits the user's selection information to the server.

[1466] Step 13:

[1467] The server processes the user's selected video in high resolution and generates the final video file.

[1468] Step 14:

[1469] The server stores and provides the generated high-resolution video for users to download.

[1470] Step 15:

[1471] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[1472] Example 1

[1473] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1474] Traditional online shopping has the problem that customers cannot actually try on products, which can cause anxiety about the size and fit of the product. Furthermore, product photos and simple descriptions alone make it difficult to fully check the product's details and actual appearance, which can lead to hesitation in purchasing. Furthermore, the need for specialized knowledge and skills to take product photos and create videos places a significant burden on product sellers.

[1475] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1476] In this invention, the server includes a means for receiving and saving product photos, a means for analyzing the photos to generate a three-dimensional silhouette of the product, a means for generating a video of a virtual model wearing the product using a generative AI model based on the generated silhouette, a means for presenting multiple videos to the user, and a means for saving and providing the video selected by the user as a final high-resolution video. This allows users to check the appearance and fit of products in detail without actually trying them on, eliminating any anxiety about purchasing. Furthermore, product sellers can create high-quality product promotional videos without specialized knowledge of photography or editing, leading to increased sales.

[1477] "Product photos" are image data of products uploaded by users.

[1478] "Means for storing" refers to the system's function of storing received product photos in a database or storage.

[1479] The "means for analyzing" refers to a function of the system that executes an image processing algorithm to identify the contours and shape of the product using a product photograph and generate a three-dimensional silhouette.

[1480] A "three-dimensional silhouette" is three-dimensional shape information of a product generated from a product photo.

[1481] A "generative AI model" is an algorithm that uses deep learning technology to generate new data based on input data, and in this invention is used to generate videos of virtual models.

[1482] A "virtual model" is a model of a person or object that is digitally generated using a generative AI model.

[1483] The "means for generating video" is a system function that creates video data based on the generated three-dimensional silhouette and virtual model.

[1484] The "presentation means" is a system function that displays the generated multiple videos so that the user can check them.

[1485] "High-resolution video" refers to high-quality video data that can display every detail clearly.

[1486] The system of the present invention is a technology that allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Specific embodiments of the system of the present invention are described below.

[1487] 1. Upload product photos:

[1488] Users take photos of the items they want to sell using a digital camera or smartphone.

[1489] The user logs into the system's web interface and selects a product photo in the upload window.

[1490] The terminal transmits the selected photo data to the server using an HTTP request.

[1491] The server stores the received photos in a database and organizes them into folders for analysis, using a cloud storage service such as Amazon S3.

[1492] 2. Generate a 3D silhouette of the product:

[1493] The server analyzes the stored photos using an image processing algorithm (e.g., OpenCV).

[1494] The server uses the analysis results to identify the contours and shape of the product, using edge detection and segmentation techniques in the process.

[1495] The server uses Structure-from-Motion (SfM) technology to generate a three-dimensional silhouette of the product based on the identified shape information.

[1496] The server saves the generated three-dimensional silhouette data as a 3D model file in preparation for the next process.

[1497] 3. Specify the model and background:

[1498] The user inputs the desired model's physical information (e.g., height, weight, etc.) and background setting (e.g., studio background or outdoor background) through the system interface.

[1499] The terminal sends the setting information entered by the user to the server in JSON format.

[1500] The server saves the received setting information in a file and prepares for the next video generation process based on this information.

[1501] 4. Video creation using generative AI models:

[1502] The server integrates the three-dimensional silhouette data with the model's physique information and background settings specified by the user.

[1503] The server uses deep learning libraries (e.g., TensorFlow and PyTorch) to run generative AI models (e.g., Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF)).

[1504] The server inputs a prompt statement (e.g., "Generate a video of the person wearing a shirt and place it against a specified background") into the generative AI model, causing it to generate multiple videos.

[1505] The server displays the generated videos on a web interface and asks the user for confirmation.

[1506] 5. Video Selection:

[1507] The user can view the suggested videos on a web interface and select the one they like best.

[1508] The terminal captures the user's selection actions and transmits the information to the server.

[1509] The server processes the selected video data again at high resolution and saves it as the final video.

[1510] 6. Generate and save the final video:

[1511] The server encodes the selected video in high resolution and generates the final video file, using libraries such as FFmpeg.

[1512] The server stores the generated high-resolution video in cloud storage such as Amazon S3 and generates a link that allows users to download it.

[1513] The server notifies the user of the download link and provides it through a web interface.

[1514] Example: Creating a shirt video

[1515] 1. A user (retailer) uploads five photos of a new shirt to the system: front, back, left, right, and right-angle views.

[1516] 2. The server analyzes the five received photos, identifies the outline of the shirt using OpenCV, and generates a three-dimensional silhouette of the shirt using SfM technology.

[1517] 3. The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background) into the system's interface.

[1518] 4. The server uses the generative AI model to generate multiple videos based on the specified criteria and suggests those videos to the user on a web interface.

[1519] 5. The user selects the most suitable video from the suggested videos and sends that information to the server.

[1520] 6. The server regenerates the selected video in high resolution, stores it on Amazon S3, and then provides a link for the user to download it.

[1521] An example of a prompt sentence for this invention is, "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background." Using this prompt sentence, the generative AI model can generate the optimal video based on the required information.

[1522] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1523] Step 1: Upload your product photos

[1524] The user takes photos of the product they want to sell (for example, photos taken from the front, back, left and right, and diagonal angles), which are used as input data.

[1525] The user logs into the system's web interface, selects a photo of the product, and clicks the upload button.

[1526] The terminal transmits the selected photo data to the server using an HTTP request.

[1527] The server stores the received photo data using a cloud storage service such as Amazon S3.

[1528] Output: Image data stored in a folder for analysis, as the photos are saved on the server.

[1529] Step 2: Generate a 3D silhouette of the product

[1530] The server runs image processing algorithms such as OpenCV on the stored photos to analyze the contours and shapes of the products, using edge detection and segmentation techniques. This analysis serves as input data.

[1531] Based on the analyzed data, the server generates a 3D silhouette of the product using Structure-from-Motion (SfM) technology, a process that reconstructs a three-dimensional shape from multiple photographs.

[1532] Output: Three-dimensional silhouette data (3D model file) of the generated product.

[1533] Step 3: Specifying the model and background

[1534] The user inputs the model's physical information (e.g., height 170 cm, weight 60 kg) and background information (e.g., studio background or outdoors) on the system interface. This becomes the input data.

[1535] The terminal converts the information entered by the user into JSON format and sends it to the server using an HTTP request.

[1536] The server saves the received JSON data to a file and uses this information when running the generative AI model.

[1537] Output: A JSON file with the saved model's physique and background information.

[1538] Step 4: Creating videos with generative AI models

[1539] The server takes in and integrates the three-dimensional silhouette data with the model and background information specified by the user. This integration becomes the input data.

[1540] The server starts running generative AI models (Generative Adversarial Networks (GANs) or Neural Radiance Fields (NeRF)) using deep learning libraries (e.g., TensorFlow or PyTorch).

[1541] The server inputs a prompt statement (e.g., "Please generate a video of a model who is 170 cm tall and weighs 60 kg wearing a new spring shirt. Please use a bright studio background.") into the generative AI model and generates multiple videos.

[1542] Output: Multiple generated video files.

[1543] Step 5: Select a video

[1544] The user reviews the proposed videos on the system's web interface and selects the most suitable one, which becomes the input data.

[1545] The device captures information about the user's selected video, converts it into JSON format, and sends it to the server in an HTTP request.

[1546] The server reprocesses the selected video data in its final high resolution and prepares it for storage.

[1547] Output: The video data selected by the user.

[1548] Step 6: Generate and save the final video

[1549] The server uses a library such as FFmpeg to encode the selected video in high resolution and generate the final video file, which becomes the input data.

[1550] The server stores the high-resolution video in cloud storage such as Amazon S3 and generates a link that users can use to download it.

[1551] The server notifies the user of the download link through a web interface.

[1552] Output: A link to a high-resolution video file that users can download.

[1553] (Application example 1)

[1554] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1555] Traditional try-on experiences in brick-and-mortar stores required customers to physically try on the products to check them, which was time-consuming and labor-intensive. Also, trying on multiple products required customers to use a fitting room and take the products in and out, which was inefficient for both customers and stores. Furthermore, if a particular size or model was not in stock at the store, customers were unable to try on that product. To solve these problems, new technology was needed to enable virtual try-on in real time.

[1556] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1557] In this invention, the server includes means for receiving and storing product images, means for analyzing the received product images to generate a three-dimensional skeleton of the product, and means for generating a motion video in which the product is applied to a virtual model using a generative AI model based on the generated three-dimensional skeleton. This allows users to obtain product images in a physical store using a mobile device or visual aid device and have a virtual try-on experience in real time. This eliminates the hassle of trying on products and allows users to efficiently try on many products, thereby increasing customer purchasing motivation beyond the constraints of stock and size.

[1558] "Product images" are photos or image data of products taken or scanned by users.

[1559] A "three-dimensional skeleton" is three-dimensional shape information of a product generated from two-dimensional image data.

[1560] A "generative AI model" is an algorithm or model that uses artificial intelligence technology to generate virtual try-on videos from image data.

[1561] "Motion video" is a video that shows a virtual model moving while wearing the product.

[1562] A "mobile terminal" is a portable information processing terminal such as a smartphone or tablet.

[1563] A "visual aid" is a device that assists with visual information, such as smart glasses or a head-mounted display.

[1564] The "user interface" refers to a screen or input device that allows the user to interact with the system, and includes functions for specifying body type information and background.

[1565] "High-definition video" refers to video data that has high resolution and quality and contains detailed information.

[1566] "Real-time" is a term that refers to processing occurring immediately, with little time delay.

[1567] "Virtual try-on experience" means using digital technology to provide the experience of trying on products without actually trying them on.

[1568] "Inventory" refers to the quantity and variety of products stored in a store or warehouse.

[1569] This technology allows users to upload product images, and generates a moving video of a virtual model wearing the product based on a three-dimensional skeleton generated from the image. To realize this technology, the following steps and system configuration are required.

[1570] Uploading product images

[1571] Users take or scan product images in physical stores using a mobile device or visual aid. The device sends the captured image to a server, which then stores it. Image processing libraries such as OpenCV are used to identify the position and shape of the product image.

[1572] Three-dimensional skeleton generation

[1573] The server analyzes the stored product images and generates a three-dimensional skeleton of the product, using Structure-from-Motion technology to extract the product's three-dimensional shape information from multiple image data.

[1574] Generation of motion images

[1575] Next, the server uses a generative AI model (e.g., GANs or NeRF) based on the generated three-dimensional skeleton to generate a moving image of the product applied to the virtual model. The user can specify the model's physical information (e.g., height, weight) and background information through the system's user interface.

[1576] Presentation and selection of motion images

[1577] The server generates an interface for presenting the generated motion images to the user, who can then review the motion images and select the most suitable one.

[1578] Final high-definition video generation and storage

[1579] The server stores the video selected by the user as high-definition video and allows the user to download it, providing a real-time virtual try-on experience and a comfortable shopping experience.

[1580] Hardware and software used

[1581] Hardware: mobile devices, smart glasses, head-mounted displays, servers

[1582] Software: OpenCV, Structure-from-Motion library, GANs, NeRF

[1583] Specific example explanation

[1584] As an example, consider the case where a user wants to virtually try on a shirt in a physical store.

[1585] 1. The user (customer) takes five photos of a shirt with their smartphone in a physical store (from the front, back, left and right, and diagonal).

[1586] 2. The device sends the captured photo to the server.

[1587] 3. The server analyzes the received photo and generates a three-dimensional skeleton of the shirt.

[1588] 4. The user specifies the model's physical information (e.g., height 170 cm, weight 60 kg) and the background (e.g., studio background).

[1589] 5. The server uses the generative AI model to generate multiple motion videos of the shirt being worn by the specified virtual model and suggests them to the user.

[1590] 6. The user selects the most suitable video from the suggested videos.

[1591] 7. The server stores the selected footage as high definition footage and makes it available for user download.

[1592] Prompt Sentence Examples

[1593] Photo path = upload_photo('Shirt photo taken in store.jpg')

[1594] Silhouette = generate_3d_silhouette(photopath)

[1595] User information = {'height': 170, 'weight': 60} The user enters their body type information.

[1596] background = 'store background' Use the store background as an example

[1597] Video = create_virtual_model(silhouette, user information, background)

[1598] save_video(video, 'Virtual Try-On Shirt 2023.mp4')

[1599] By following the above-described procedure, the present invention can generate a three-dimensional skeleton from a product image and provide a virtual try-on experience in real time.

[1600] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1601] Step 1:

[1602] Uploading product images

[1603] A user takes product images in a physical store using a mobile device or visual aid. The device sends the captured images to a server, which then stores the received images. Specifically, the user takes photos from multiple angles and the device uploads the image data to the server. The input is the captured product image, and the output is the image data stored on the server.

[1604] Step 2:

[1605] Three-dimensional skeleton generation

[1606] The server analyzes the stored product images and generates a three-dimensional skeleton of the product. Image processing libraries such as OpenCV are used for the analysis, and Structure-from-Motion technology is applied to extract the product's three-dimensional shape information from multiple image data. The input is the stored product image, and the output is three-dimensional skeleton data. Specifically, an image processing algorithm is executed on the server to measure the shape and size of the product.

[1607] Step 3:

[1608] Specifying the model and background

[1609] The user uses a user interface within the system to specify the model's physical information (e.g., height, weight) and background information. The input is the physical and background information entered by the user, and the output is a dataset containing that information. Specifically, the user enters the required information using drop-down menus and input forms in the user interface, and the information is sent to the server.

[1610] Step 4:

[1611] Virtual model generation

[1612] The server integrates the generated 3D skeleton with the model information specified by the user and uses a generative AI model (e.g., GANs or NeRF) to generate a moving image of the product applied to the virtual model. The input is the 3D skeleton data, model information, and background information, and the output is multiple moving images. Specifically, the generative AI model simulates the product's movement and executes the process of applying it to the virtual model.

[1613] Step 5:

[1614] Presentation and selection of motion images

[1615] The server generates an interface to present the generated multiple motion videos to the user. The user reviews the presented motion videos and selects the most suitable one. The input at this time is the generated motion video, and the output is the motion video selected by the user. Specifically, the user watches multiple videos on the interface and performs an operation to select the best one.

[1616] Step 6:

[1617] Final high-definition video generation and storage

[1618] The server saves the video selected by the user as high-definition video and makes it available for download. The input is the motion video selected by the user, and the output is the final high-definition video data. Specifically, the selected video is regenerated in high resolution on the server and provided to the user. During this process, the final video data is saved in a dedicated folder on the server.

[1619] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1620] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. Furthermore, by combining this with an emotion engine that recognizes the user's emotions, the present invention enables more effective video generation and presentation.

[1621] System program processing explanation

[1622] 1. Upload product photos

[1623] The user uploads photos of products (e.g., shirts, pants, etc.) to the system through the terminal using the upload interface on the system. Multiple photos can also be selected.

[1624] The terminal transmits the selected photo to the server.

[1625] The server stores the received photos and organizes them into folders for analysis.

[1626] 2. Creating a three-dimensional silhouette of a product

[1627] The server analyzes the received product photos and runs image processing algorithms to identify the product's shape and size, using image processing libraries such as OpenCV.

[1628] The server integrates the shape information from multiple photos to generate a logical 3D silhouette of the product, sometimes using Structure-from-Motion technology.

[1629] 3. Specifying the model and background

[1630] The user specifies the model's physical information (for example, height and weight) and background information through an interface provided within the system.

[1631] The terminal transmits this specification information to the server.

[1632] 4. Video creation using generative AI models

[1633] The server integrates the 3D silhouette with the user's model information and uses a generative AI model (e.g., GANs or NeRF) to generate a video of the virtual model wearing the product, taking into account the user's specified background information.

[1634] The server generates an interface for suggesting the generated multiple videos to the user.

[1635] 5. Recognition of user emotions by emotion engine

[1636] The server runs an emotion engine that recognizes the user's emotions in real time during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc.

[1637] The device uses the user's camera and microphone to send emotion data to the emotion engine.

[1638] Based on the analyzed emotional information, the emotion engine identifies the videos that the user is most likely to be interested in and changes the order in which they are presented.

[1639] 6. Video Presentation and Selection

[1640] The server presents the video generated in the order adjusted by the emotion engine to the user.

[1641] The terminal displays the suggested videos to the user, and the user selects the most suitable video from the displayed videos.

[1642] The terminal transmits the user's selection information to the server.

[1643] 7. Generate and save the final video

[1644] The server saves the video selected by the user as a high-resolution video and generates the final video file.

[1645] The server stores and provides the generated high-resolution video for users to download.

[1646] Specific examples

[1647] Example 1: Creating a shirt video

[1648] 1. A user (retailer) uploads five photos of a new shirt to the system, including front, back, left and right views, and three-quarter views.

[1649] 2. The server analyzes the five received photos and generates a three-dimensional silhouette of the shirt.

[1650] 3. The user specifies a model image (e.g., a man with a height of 170 cm and a weight of 60 kg) and a background (e.g., a studio background).

[1651] 4. The server uses the generative AI model to generate multiple videos of the shirt being worn by the specified model and suggests them to the user.

[1652] 5. The server runs an emotion engine during the video selection process, analyzing the user's facial expressions, voice, click patterns, etc., and adjusts the presentation order.

[1653] 6. The user selects the most suitable video from the suggested videos.

[1654] 7. The server stores the selected video in high resolution and makes it available for download by the user.

[1655] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present optimal videos according to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[1656] The processing flow will be explained below.

[1657] Step 1:

[1658] The user uses the system's upload interface to select multiple photos of the product and upload these photos through the terminal.

[1659] Step 2:

[1660] The terminal transmits the selected photo file to the server.

[1661] Step 3:

[1662] The server stores the received photo files and organizes them into folders for further processing.

[1663] Step 4:

[1664] The server opens the stored product photos and uses image processing algorithms (e.g., libraries such as OpenCV) to identify the product's shape and size.

[1665] Step 5:

[1666] The server then integrates the shape information from the analyzed photos to generate a three-dimensional silhouette of the product, sometimes using Structure-from-Motion technology.

[1667] Step 6:

[1668] The user specifies the model's physical information (e.g., height, weight) and background information through an interface within the system.

[1669] Step 7:

[1670] The terminal transmits the physique information and background information input by the user to the server.

[1671] Step 8:

[1672] The server uses a generative AI model (e.g., GANs or NeRF) to generate a video of a virtual model wearing the product based on the three-dimensional silhouette and the physique and background information specified by the user.

[1673] Step 9:

[1674] The server generates an interface for suggesting the generated multiple videos to the user.

[1675] Step 10:

[1676] The server runs an emotion engine to acquire emotion data in real time from the user's camera and microphone while displaying the suggested video.

[1677] Step 11:

[1678] The terminal inputs the user's facial expressions, voice, click patterns, etc. into an emotion engine and analyzes the emotion data.

[1679] Step 12:

[1680] The emotion engine identifies the user's emotional state based on the analyzed emotion information, identifies the videos that the user is most interested in, and changes the presentation order of those videos.

[1681] Step 13:

[1682] The server presents the multiple videos to the user in an order adjusted by the emotion engine.

[1683] Step 14:

[1684] The terminal displays the videos to the user in the adjusted order, and the user selects the most suitable video from the suggested videos.

[1685] Step 15:

[1686] The terminal transmits the user's selection information to the server.

[1687] Step 16:

[1688] The server processes the video selected by the user as a high-definition video and generates the final video file.

[1689] Step 17:

[1690] The server stores and provides the generated high-resolution video for users to download.

[1691] Step 18:

[1692] Users can then use their devices to download the final high-resolution video from the server for playback or sharing.

[1693] Example 2

[1694] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1695] Conventional systems have had difficulty generating a three-dimensional silhouette using product photos and creating videos in which a virtual model wears the product. Furthermore, the order in which the generated videos are presented does not take into account the user's interests or emotions, which makes it difficult to efficiently select the most suitable video. Furthermore, the interface for users to specify the model's physique and background information is inadequate, making it difficult to customize the system to meet individual needs.

[1696] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1697] In this invention, the server includes means for receiving and saving product photos via a user terminal, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user terminal, means for saving and providing a video selected by the user based on selection information transmitted from the user terminal as a final high-resolution video, and means for adjusting the video presentation order, which includes an emotion engine that collects and analyzes user emotion data using the camera and microphone of the user terminal. This significantly improves the user experience, enabling optimal video presentation and customized video generation according to individual needs.

[1698] "User terminal" refers to any device that a user uses to connect to the Internet and send and receive data.

[1699] A "server" is a computer system that provides data processing and storage functions over a network.

[1700] "Product Photos" refers to product image data uploaded by users.

[1701] "Three-dimensional silhouette" refers to a three-dimensional shape model of a product generated based on a product photo.

[1702] A "generative AI model" is a model that uses artificial intelligence algorithms to generate data to perform a specific task (in this case, video generation).

[1703] "Virtual Model" refers to a digital sculpture or animation of a digitally generated figure or form that simulates a tangible product or scene.

[1704] "High-definition video" refers to video files produced at high-quality image resolution.

[1705] An "emotion engine" refers to technology and algorithms that analyze a user's facial expressions, voice, behavioral patterns, etc. to recognize emotions in real time.

[1706] "Means for adjusting the presentation order" refers to a function that rearranges the order in which videos are displayed based on the analysis results of the emotion engine.

[1707] MODE FOR CARRYING OUT THE INVENTION

[1708] The system of the present invention allows users to upload product photos, and generates a video of a virtual model wearing the product based on a three-dimensional silhouette generated from the photo. By combining this with an emotion engine that recognizes the user's emotions, the system presents the most appropriate video to the user.

[1709] The system of the present invention includes the following hardware and software.

[1710] User terminal: This can be a smartphone, tablet, PC, or other device. The user terminal connects to the Internet and sends and receives data via the system interface.

[1711] Server: A computer system that provides data processing and storage functions over a network. The server receives and analyzes photo data, generates 3D silhouettes, generates videos using generative AI models, runs the emotion engine, and displays and stores videos.

[1712] Image processing libraries: Image processing libraries such as OpenCV are used to analyze product photos and generate 3D silhouettes.

[1713] Generative AI models: Generative Adversarial Networks (GANs) and Neural Radiance Fields (NeRF) are used to generate videos of virtual models wearing products.

[1714] Emotion engine: Analyzes the user's facial expressions, voice, and click patterns to adjust the order in which videos are presented.

[1715] Specific Examples

[1716] 1. Product photo upload example

[1717] A user (e.g., a representative from an online retailer) uses the system's interface to upload five photos of a new shirt: photos taken from the front, back, left and right, and an angle.

[1718] The terminal divides the uploaded photo into packets and sends them to the server.

[1719] 2. Product photo analysis and three-dimensional silhouette generation example

[1720] The server analyzes the received product photos and extracts feature points from each photo using OpenCV.

[1721] The server uses Structure-from-Motion technology to reconstruct the three-dimensional shape of the product based on the extracted feature points and generate a three-dimensional silhouette.

[1722] 3. Example of specifying a model and background

[1723] The user inputs the virtual model's physical information (for example, height 170 cm, weight 60 kg) and background information (for example, studio background) through an interface provided within the system.

[1724] The terminal transmits this information to the server.

[1725] 4. Example of video creation using generative AI model

[1726] The server combines the generated three-dimensional silhouette with specified model information and uses a generative AI model such as GANs or NeRF to generate a video of the virtual model wearing the product.

[1727] The server provides an interface to present the generated multiple animation versions to the user.

[1728] 5. Example of user emotion recognition using emotion engine

[1729] The device uses a camera and microphone to capture the user's facial expressions and voice while watching videos, and also records click patterns.

[1730] The device transmits the captured data in real time to a server, where an emotion engine built into the server analyzes the data.

[1731] The server adjusts the presentation order of the videos based on the analysis results.

[1732] 6. Video presentation and selection examples

[1733] The server presents the generated video in the adjusted order to the user terminal.

[1734] The user selects the most suitable video from the displayed videos.

[1735] 7. Example of generating and saving the final video

[1736] The server renders the user's selected video in high resolution and generates the final video file.

[1737] The server stores the generated high-resolution video in cloud storage or on a dedicated server and provides a link for users to download it.

[1738] In this way, the system of the present invention provides users with a means to easily and efficiently create and publish videos of models wearing products without requiring them to have photography or video editing skills. Furthermore, by combining it with an emotion engine, it is possible to present videos that are optimally tailored to the user's interests and emotions, thereby increasing the effectiveness of promoting purchases.

[1739] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1740] The flow of this system's program processing

[1741] Step 1:

[1742] Users upload product photos using the upload interface on the system. The input is the product photo selected by the user (front, back, left, right, oblique, etc.), and the output is a photo file sent to the server via the terminal.

[1743] Step 2:

[1744] The terminal divides the received product photos into packets and securely sends them to the server. The input is the photo data uploaded by the user, and the output is the divided photo data sent to the server. Specific operations include data packetization and encrypted communication.

[1745] Step 3:

[1746] The server saves the received photo data and organizes it in a folder for analysis. The input is the photo data sent from the terminal, and the output is the saved photo file and product information recorded in the database.

[1747] Step 4:

[1748] The server analyzes product photos using an image processing library such as OpenCV. The input is the saved product photo, and the output is the extracted feature point data. Specific operations include image reading and feature point detection.

[1749] Step 5:

[1750] The server matches feature points between multiple photos and reconstructs the three-dimensional shape of the product using Structure-from-Motion technology. The input is feature point data extracted from multiple photos, and the output is a three-dimensional silhouette (3D model) of the product. Specific operations include matching corresponding feature points, estimating camera position, and generating a point cloud.

[1751] Step 6:

[1752] The user specifies model physique information and background information through the interface. The input is the model and background information entered by the user, and the output is the model information and background information sent from the terminal.

[1753] Step 7:

[1754] The terminal transmits the specified model information and background information to the server. The input is the model information and background information input by the user, and the output is the model information and background information transmitted to the server.

[1755] Step 8:

[1756] The server integrates the model information and the 3D silhouette and uses a generative AI model to generate a video of the virtual model wearing the product. The input is the model information and the 3D silhouette, and the output is multiple generated videos. Specific operations include inputting data into the generative AI model and simulating the movement of the product.

[1757] Step 9:

[1758] The server generates an interface for proposing the generated videos to the user, where the input is the generated videos and the output is the video suggestion interface displayed on the user terminal.

[1759] Step 10:

[1760] The device collects emotional data using the user's camera and microphone and sends it to a server. The input is the user's facial expressions, voice, and click patterns, and the output is the emotional data sent to the server. Specific operations include facial expression recognition and voice analysis.

[1761] Step 11:

[1762] The server analyzes the data collected by the emotion engine and adjusts the video presentation order. The input is the transmitted emotion data, and the output is the adjusted video presentation order. Specific operations include data analysis and rearrangement of the presentation order.

[1763] Step 12:

[1764] The user selects the most suitable video from the displayed videos. The input is the video presented by the server, and the output is the selected video information.

[1765] Step 13:

[1766] The terminal transmits the user's selection information to the server, where the input is the video information selected by the user and the output is the selection information transmitted to the server.

[1767] Step 14:

[1768] The server renders the selected video in high resolution and generates and saves the final video file. The input is the selected video information, and the output is a high-resolution video file. Specific operations include high-resolution rendering and video file saving.

[1769] Step 15:

[1770] The server generates a link that allows the user to download the generated high-resolution video and provides it to the user. The input is the generated high-resolution video, and the output is the download link.

[1771] (Application example 2)

[1772] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1773] In today's online shopping environment, users cannot actually touch the product, making it difficult to accurately grasp its appearance and feel. There are also limited ways to maximize the product's appeal and effectively communicate it. Furthermore, there is a lack of ways to recognize users' purchasing intentions and interests and make optimal suggestions based on them, creating a need to improve the user experience.

[1774] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1775] In this invention, the server includes means for receiving and saving product photos, means for analyzing the received product photos to generate a three-dimensional silhouette of the product, means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette, means for presenting multiple generated videos to the user, means for saving and providing a video selected by the user as a final high-resolution video, means for recognizing and analyzing the user's emotions, means for adjusting the presentation order of videos that are likely to be of most interest based on the analyzed emotional information, and means for providing an interface for the user to specify the model's physique information and background when displaying the generated videos. This allows users to accurately grasp the appearance and feel of the product in real time, and makes it possible to suggest optimal videos according to their individual interests and emotions.

[1776] "Product Photos" are images taken or provided by users that capture the exterior of a product from multiple angles.

[1777] A "three-dimensional silhouette" is a digital model that reproduces the shape and size of a product in three dimensions based on photographs of multiple products.

[1778] A "generative AI model" is an artificial intelligence model that uses deep learning techniques to generate new data from data provided by users, and includes, for example, Generative Adversarial Networks (GANs).

[1779] A "virtual model" is a digital human model created by a generative AI model based on physique information specified by the user.

[1780] "Emotion recognition" is a technology that analyzes a user's facial expressions, voice, click patterns, etc. to identify their emotional state at that time.

[1781] "High-definition video" is a video file created with a custom resolution setting that is higher than the standard resolution.

[1782] The "presentation order" is the order in which multiple generated videos and information are displayed to the user.

[1783] "Physical information" is data that represents a person's physical characteristics, such as the user's height, weight, and gender.

[1784] An "interface" is a software part that provides the screen and operating means for the user to interact with the system.

[1785] "Background" means the digital background scene or environment in which the virtual model is displayed, as selected by the user.

[1786] The system for realizing the present invention is configured using the following specific hardware and software.

[1787] Hardware:

[1788] Devices: Smartphone, camera, microphone

[1789] Server: A high-performance computer, either a cloud service or on-premise

[1790] software:

[1791] OpenCV: Image processing library

[1792] Emotion Recognition Engine: Software that recognizes user emotions in real time

[1793] Generative AI models: Deep learning techniques, such as Generative Adversarial Networks (GANs), to generate new data based on user-uploaded photos.

[1794] User Interface: The part of the software that allows the user to interact with the system.

[1795] System Details:

[1796] 1. Upload product photos

[1797] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles and upload these photos to the app, which then sends the uploaded photos to the server.

[1798] 2. Creating a three-dimensional silhouette of a product

[1799] The server analyzes the received product photos and uses image processing libraries such as OpenCV to identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product.

[1800] 3. Specifying the model and background

[1801] The user specifies his / her physical information (e.g., height, weight) and background image through an interface provided within the system. The terminal transmits this information to the server.

[1802] 4. Video Generation

[1803] Based on the three-dimensional silhouette and the specified model information, the server uses a generative AI model to generate a video in which the virtual model wears the product, taking into account the background information specified by the user.

[1804] 5. Emotional Engine Optimization

[1805] The server uses an emotion recognition engine to analyze the user's facial expressions, voice, click patterns, etc. in real time, and adjusts the order of videos presented to them based on the user's likely interest.

[1806] 6. Video Presentation and Selection

[1807] The generated videos are presented to the user in the order adjusted by the emotion engine. The user selects the most suitable video from these videos. The selected video is then sent from the device to the server.

[1808] 7. Generate and save the final video

[1809] The server finally stores the user's selected video in high resolution and makes it available for the user to download.

[1810] Program operation description:

[1811] This system operates by combining the above hardware and software. First, the device sends user input to the server, which then analyzes and processes the received data. OpenCV is used for image analysis, and a generative AI model (such as GANs) generates the video. The emotion recognition engine collects and analyzes the user's emotional data. The user interface is the means by which the user inputs and makes selections, and also presents the final video.

[1812] Examples:

[1813] For example, if a user wants to try on a new dress, the process would be as follows:

[1814] 1. The user takes photos of the dress from multiple angles and uploads them to the app.

[1815] 2. The server generates a three-dimensional silhouette of the dress based on these photos.

[1816] 3. The user specifies their physical information (e.g., height 170 cm, weight 60 kg) and a background image (e.g., a beach scene).

[1817] 4. The server uses the generative AI model to generate multiple videos of the virtual model wearing the dress.

[1818] 5. The emotion engine analyzes the user's facial expressions and click patterns to prioritize presenting the most interesting videos to the user.

[1819] 6. Users can select their favorite videos and save and share them as final high-resolution videos.

[1820] Example prompt sentence:

[1821] 1. Take photos of your product from multiple angles and upload them to the app.

[1822] 2. Select your physical information and background.

[1823] 3. Review the videos generated by the app and choose the one you like the most.

[1824] This allows users to accurately understand the appearance and feel of a product in real time, making it possible to suggest the most suitable video based on their individual interests and emotions.

[1825] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1826] Step 1:

[1827] Users take and upload product photos

[1828] Users use their smartphone camera to take photos of products (e.g., dresses, shoes) from multiple angles. These photos are then uploaded to the app. The input is the photos of multiple products taken, and the output is the photo data sent to the server. Specifically, users use the app's photo upload interface, and their device sends these photos to the server.

[1829] Step 2:

[1830] The server receives and saves the product photos.

[1831] The server receives product photos sent from the device and saves them in a folder for analysis. The input is the photo data sent from the device, and the output is a photo file stored in a storage folder on the server. The server first saves the received files in a specific directory and organizes them within the folder.

[1832] Step 3:

[1833] The server analyzes the product photo and generates a three-dimensional silhouette

[1834] The server uses image processing libraries such as OpenCV to analyze the received product photos and identify the product's shape and size. It then integrates the shape information obtained from multiple photos to generate a three-dimensional silhouette of the product. The input is multiple stored photos, and the output is three-dimensional silhouette data. The server then runs image analysis algorithms to detect edges and extract specific features from the photos.

[1835] Step 4:

[1836] User specifies model and background information

[1837] The user specifies the model's physical information (e.g., height, weight) and background information through an interface provided within the system. The input is the user's physical information and background image selection data, and the output is that this specified information is sent to the server. The user enters information using pull-down menus and text boxes, and the terminal sends this data to the server.

[1838] Step 5:

[1839] The server generates the video

[1840] The server integrates the 3D silhouette with the specified model information and uses a generative AI model (e.g., GANs) to generate a video in which the virtual model wears the product. At this time, it also takes into account background information specified by the user. The input is the 3D silhouette, physique information, and background information, and the output is the generated video data. Specifically, the server runs the generative AI model and generates a video in which the virtual model wears the product based on the input data.

[1841] Step 6:

[1842] The server uses an emotion engine to recognize and analyze the user's emotions in real time.

[1843] The server runs an emotion recognition engine and analyzes the user's facial expressions, voice, click patterns, etc. in real time. The input is the user's emotional data (facial expressions, voice, click patterns), and the output is analyzed emotional information. The device uses a camera and microphone to acquire user data and sends it to the server. The server sends this data to the emotion engine and returns the analysis results.

[1844] Step 7:

[1845] The server adjusts the video presentation order based on emotional information

[1846] The server adjusts the presentation order of videos that the user is most likely to be interested in based on the analyzed emotional information. The input is the analyzed emotional information, and the output is the adjusted video presentation order. The server rearranges the video presentation order based on the results of the emotion engine.

[1847] Step 8:

[1848] Presentation and selection of generated videos

[1849] The server presents the videos generated in the order adjusted by the emotion engine to the user, and the user selects the most appropriate video from the displayed videos. The input is the video presented by the server, and the output is the user's selection information. The terminal displays the video to the user and sends the user's selection to the server.

[1850] Step 9:

[1851] Generate and save the final video

[1852] The server saves the user-selected video in high resolution and generates the final video file. The input is the user-selected video, and the output is the final high-resolution video file. The server saves this in cloud storage and makes it available for user download.

[1853] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1854] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1855] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1856] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1857] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1858] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1859] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1860] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1861] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1862] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1863] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1864] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1865] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1866] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1867] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1868] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1869] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1870] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1871] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1872] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1873] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1874] The following is further disclosed regarding the above embodiment.

[1875] (Claim 1)

[1876] a means for receiving and storing product photos;

[1877] A means for analyzing the received product photo to generate a three-dimensional silhouette of the product;

[1878] A means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette;

[1879] means for presenting the plurality of generated videos to a user;

[1880] The system includes means for storing and providing the user-selected video as a final high-resolution video.

[1881] (Claim 2)

[1882] The system of claim 1, further comprising means for identifying the position and shape of a product from multiple photographs when analyzing product photographs.

[1883] (Claim 3)

[1884] 2. The system according to claim 1, further comprising means for providing an interface for a user to specify physique information and background of a model when displaying the generated video.

[1885] "Example 1"

[1886] (Claim 1)

[1887] a means for receiving and storing product photos;

[1888] A means for analyzing the received product photo to generate a three-dimensional silhouette of the product;

[1889] A means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette;

[1890] means for presenting the plurality of generated videos to a user;

[1891] The system includes means for storing and providing the user-selected video as a final high-resolution video.

[1892] (Claim 2)

[1893] 10. The system according to claim 1, further comprising means for analyzing a plurality of photographs to identify product contours and generate a three-dimensional silhouette.

[1894] (Claim 3)

[1895] The system of claim 1 includes a means for a user to specify the model's physique information and background information, and to generate a video using a generative AI model based on that information.

[1896] "Application Example 1"

[1897] (Claim 1)

[1898] means for receiving and storing product images;

[1899] A means for analyzing the received product image and generating a three-dimensional skeleton of the product;

[1900] A means for generating a motion video in which the product is applied to a virtual model using a generative AI model based on the generated three-dimensional skeleton;

[1901] means for presenting the plurality of generated motion images to a user;

[1902] A means for saving and providing the user-selected motion video as a final high-definition video;

[1903] The system includes a means for users to acquire product images in a physical store using a mobile device or visual aid and to have a virtual try-on experience in real time.

[1904] (Claim 2)

[1905] The system of claim 1, further comprising means for identifying the position and shape of the product from multiple images when analyzing the product images.

[1906] (Claim 3)

[1907] 2. The system according to claim 1, further comprising means for providing a user interface that allows a user to specify physical information and background of a model when the generated motion video is displayed.

[1908] "Example 2: Combining Emotion Engines"

[1909] (Claim 1)

[1910] means for receiving and storing product photos via a user terminal;

[1911] A means for analyzing the received product photo to generate a three-dimensional silhouette of the product;

[1912] A means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette;

[1913] means for presenting the plurality of generated videos to a user terminal;

[1914] a means for storing and providing a final high-resolution video selected by the user based on selection information transmitted from the user terminal;

[1915] A system including a means for adjusting the order in which videos are presented, and an emotion engine that uses the camera and microphone of the user's terminal to collect and analyze user emotion data.

[1916] (Claim 2)

[1917] The system of claim 1, further comprising means for identifying the position and shape of a product from multiple photographs when analyzing product photographs.

[1918] (Claim 3)

[1919] 2. The system according to claim 1, further comprising means for providing an interface for a user to specify physique information and background of a model when displaying the generated video.

[1920] "Application example 2 when combining emotion engines"

[1921] (Claim 1)

[1922] a means for receiving and storing product photos;

[1923] A means for analyzing the received product photo to generate a three-dimensional silhouette of the product;

[1924] A means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette;

[1925] means for presenting the plurality of generated videos to a user;

[1926] means for storing and providing the user-selected video as a final high-resolution video;

[1927] means for recognizing and analyzing user emotions;

[1928] The system includes a means for adjusting the order in which videos that are deemed most interesting are presented based on the analyzed emotional information.

[1929] (Claim 2)

[1930] The system of claim 1, further comprising means for identifying the position and shape of a product from multiple photographs when analyzing product photographs.

[1931] (Claim 3)

[1932] 2. The system according to claim 1, further comprising means for providing an interface for a user to specify physique information and background of a model when displaying the generated video. [Explanation of symbols]

[1933] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for receiving and storing product photos; A means for analyzing the received product photo to generate a three-dimensional silhouette of the product; A means for generating a video in which a virtual model wears the product using a generative AI model based on the generated three-dimensional silhouette; means for presenting the plurality of generated videos to a user; The system includes means for storing and providing the user-selected video as a final high-resolution video.

2. The system according to claim 1, further comprising means for identifying the position and shape of the product from a plurality of photographs when analyzing the product photographs.

3. 2. The system according to claim 1, further comprising means for providing an interface for a user to specify physique information and background of a model when the generated animation is displayed.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A