system

The system simplifies the creation of AR content by converting user-drawn illustrations into 3D shapes with animation, addressing the complexity of existing AR technology and enabling easy, interactive, and personalized AR experiences.

JP2026074880APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The creation of augmented reality (AR) content targeting three-dimensional shapes requires specialized techniques and is difficult for ordinary users due to complex operations and limited technology for directly obtaining three-dimensional shapes from illustrations, hindering the widespread use of AR technology.

Method used

A system that extracts features from user-drawn illustrations, automatically converts them into three-dimensional shapes, adds animation, and dynamically displays them in the user's environment using augmented reality, enhancing visual appeal and positioning based on location information.

Benefits of technology

Enables ordinary users to easily create and enjoy artistic and dynamic AR content by converting illustrations into 3D shapes with animation, providing an interactive and personalized AR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074880000001_ABST
    Figure 2026074880000001_ABST
Patent Text Reader

Abstract

This system provides an environment where even ordinary users without specialized knowledge can easily create and enjoy artistic and dynamic AR content. [Solution] A system comprising means for receiving an illustration drawn by a user, means for extracting features from the received illustration and automatically generating a 3D shape, means for adding animation to the generated 3D shape, and means for displaying the animated 3D shape as augmented reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The creation of augmented reality (AR) content targeting three-dimensional shapes has hitherto required specialized techniques and know-how and has been difficult for ordinary users to utilize easily. Also, existing AR content creation platforms have complex operations, and users have had to follow many procedures, thus requiring time and effort. Furthermore, the technology for directly obtaining three-dimensional shapes from illustrations and dynamically expressing them according to individual needs is limited. Due to these problems, the spread of augmented reality technology has not progressed, and its use has been limited in general.

Means for Solving the Problems

[0005] This invention provides a system that extracts features from user-drawn illustrations and automatically converts them into three-dimensional shapes. Furthermore, animation is automatically added to these three-dimensional shapes, and they are dynamically displayed in the user's surrounding environment using augmented reality technology. The system also enhances the visual appeal of the illustrations through beautification processing. In addition, it analyzes location information and positions the three-dimensional shapes in a way that is appropriate to the user's actual environment. As a result, even ordinary users without specialized knowledge can easily create and enjoy artistic and dynamic AR content.

[0006] A "user" is an individual or group that operates the system and attempts to generate a 3D shape based on their own illustration.

[0007] An "illustration" is an image or painting created by a user, from which features are extracted from a two-dimensional representation and converted into a three-dimensional shape.

[0008] "Receiving" refers to the process where illustrations created by a user are sent to a server via their device, and that data is retrieved on the server side.

[0009] "Feature extraction" refers to analyzing important data such as shape, color, and contours from an illustration to obtain the information necessary to generate a 3D model.

[0010] A "3D shape" is a three-dimensional model automatically generated by a computer based on the characteristics of an illustration.

[0011] "Automatic generation" refers to the process where a program uses AI technology to create a 3D shape without manual human intervention.

[0012] "Adding animation" refers to setting movement and actions for the generated 3D shape, allowing its changes to be displayed over time.

[0013] Augmented reality is a technology that overlays digital content onto the real world environment, displaying virtual information in real space.

[0014] "Beautifying" an illustration involves adjusting its color tone and sharpness to make it visually more appealing.

[0015] "Analyzing location information" means receiving information about the user's current location and surrounding environment, and then performing data processing based on that information to appropriately position AR content for display. [Brief explanation of the drawing]

[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Modes for Carrying Out the Invention

[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained. [[ID=​​​​​​​​​In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0024] [First Embodiment]

[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0037] The system implementing this invention receives illustrations drawn by a user, converts them into three-dimensional shapes and animated content, and displays them as augmented reality. This system is realized through the coordinated operation of the user's terminal, a central server, and a device equipped with AR capabilities.

[0038] First, users create illustrations on their devices using a dedicated application. These illustrations can be drawn directly using digital painting tools, or they can be imported by taking a picture of a drawing on paper.

[0039] Next, the device sends this illustration data to the server. The server analyzes the features of the illustration and extracts its shape and color information. Based on this information, the server uses AI technology to automatically generate a 3D shape. The server then adds pre-prepared animations to the generated 3D model. Finally, it adjusts the illustration's color tone and sharpness to make it visually appealing.

[0040] Furthermore, to enable augmented reality display, the server acquires the user's location information and calculates the optimal placement of 3D shapes for the AR device. The constructed AR content is then sent back to the user's device and finally displayed in the user's real space via the AR device.

[0041] As a concrete example, consider using an illustration of an animal drawn by a child. The user draws an illustration of an animal on their device and sends it to the server. The server analyzes the outline and colors of the received illustration and generates a 3D model of the animal. This model is then given animations such as walking and jumping. The completed AR content appears before the user's eyes through their AR glasses, just as if it were real. This allows the user to observe the animal moving around in their own physical environment.

[0042] This system provides a way for ordinary users to easily create 3D content and enjoy interactive experiences through AR technology.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user draws illustrations on their device. They launch a dedicated app and create illustrations directly using digital painting tools, or they take a picture of a drawing on paper with the device's camera and import it into the app.

[0046] Step 2:

[0047] The terminal sends the completed illustration data to the server. The transmission format is a standard image format such as JPEG or PNG, and data compression is performed to improve communication efficiency.

[0048] Step 3:

[0049] The server receives the transmitted illustration. After receiving it, it analyzes the shape and color of the illustration and extracts feature data to create a 3D shape.

[0050] Step 4:

[0051] The server automatically generates the 3D shape of the illustration using AI technology based on the analysis data. This generation process is carried out by combining existing 3D model libraries and generation algorithms.

[0052] Step 5:

[0053] The server adds animation to the generated 3D shape. The animation uses a pre-configured template and is adjusted to make the model move naturally.

[0054] Step 6:

[0055] The server automatically adjusts the color tone and sharpness and applies beautification processing to improve the visual quality of the illustrations. This enhances their final appearance.

[0056] Step 7:

[0057] The server acquires the user's location information and analyzes the 3D shape at a position and scale suitable for AR display. Based on the location analysis, the placement is adjusted to match the user's real-world environment.

[0058] Step 8:

[0059] The server sends the created AR content data to the user's device.

[0060] Step 9:

[0061] The device displays the received AR content through a device specified by the user, such as AR glasses or a smartphone. This allows the user to enjoy an augmented reality experience in the real world.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Currently, there are limited ways for users to easily convert their own illustrations into 3D shapes or animated content, and then extend and experience them in real space. Furthermore, there is a lack of features to improve visual effects in a way that is suitable for the real world, making it difficult for users to realize their ideal visual experience.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for receiving an image drawn by a user using a generation device, means for extracting features of the received image using an analysis device and automatically generating a three-dimensional shape, and means for adding movement to the generated three-dimensional shape. This makes it possible for users to easily turn their own illustrations into three-dimensional objects and experience them extended into real space with movement.

[0067] "User" refers to an individual or group that operates a specific electronic device or system.

[0068] "Generating device" refers to a computer system or device used by a user to create and edit visual information.

[0069] "Image" refers to a visual representation created by a user through a generation device.

[0070] "Means of receiving" refers to a device or system that has the function of acquiring information from an external source via data communication.

[0071] "Analyzing features" refers to performing computational processing to identify useful patterns, shapes, colors, and other elements from data.

[0072] "Three-dimensional shape" refers to information that describes the shape of an object that exists in three-dimensional space.

[0073] "Adding action" refers to setting a series of actions, such as movement or deformation, for a three-dimensional shape.

[0074] Augmented reality refers to a technology that integrates and displays digital content within a real-world physical environment.

[0075] The "visual environment" refers to the physical and digital spaces that a user can perceive through their vision.

[0076] "Automatically adjusting color and image clarity" refers to algorithmic processing that appropriately corrects the color tone and resolution of digital images.

[0077] "Improving visual effects" refers to techniques used to enhance the beauty and appearance of displayed content.

[0078] "Analyzing location information" refers to processing and understanding data based on the user's geographical location.

[0079] "Appropriate placement" refers to arranging digital content in a reasonable manner within a physical or virtual space.

[0080] This invention is a system that converts user-drawn images into three-dimensional shapes and animated content, and displays them in a visual environment as augmented reality. It is realized through the cooperation of the user, terminal, and server.

[0081] Users create images using digital painting tools with a dedicated generator. They can also digitize images drawn on paper by taking a picture of it with the device's camera. Once the image is complete, the device sends the data to the server. The data is encrypted during transmission to ensure secure reception.

[0082] The server analyzes the received image data using an analysis device. This process utilizes AI technology to extract image contours and color information, and identify necessary features. Libraries such as OpenCV and TENSORFLOW® are used for image analysis. Based on the analysis results, the server automatically generates a 3D shape. 3D modeling tools such as Blender and Unity are used for 3D shape generation.

[0083] The generated 3D shapes are given motion using motion capture data, making them appear to move. The server also automatically adjusts the colors and image clarity to improve the appearance. Tools such as Adobe Color and Photoshop are used in this process.

[0084] Furthermore, the server analyzes location information and appropriately positions augmented reality content based on the device's physical location. It uses ARKit or ARCore to perform calculations for placing 3D models in real space.

[0085] Ultimately, the device displays augmented reality content in the visual environment via an AR device, based on data received from the server. For example, an image of an animal drawn by the user can be rendered in 3D and displayed in the living room with animations.

[0086] As an example of a prompt, it is possible to input instructions to the server such as, "Convert the illustration of a sea creature drawn by the user into a 3D model and make it swim in the AR space." This allows the user to experience their imagination in three dimensions in the real world.

[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0088] Step 1:

[0089] The user activates the generator and creates an image using digital painting tools. Input includes brush strokes and color selections from the drawing tools. Once the user finishes drawing, they save it as a digital file (e.g., PNG, JPEG). The output is a digital image file.

[0090] Step 2:

[0091] The terminal sends image files created by the user to the server. The input is a digital image file, and the terminal encrypts the data using an appropriate protocol for transmission. The output is sent to the server as encrypted image data.

[0092] Step 3:

[0093] The server analyzes the received image data. The input is encrypted digital image data, which is decrypted, and AI technology (e.g., image recognition models) is used to extract contour and color features. Specifically, image processing is performed using OpenCV or TensorFlow. The output is feature data for generating 3D shapes.

[0094] Step 4:

[0095] The server automatically generates a 3D shape based on extracted feature data. The input is feature data. Using the generated AI model, a 3D model is created using tools such as Blender or Unity. Textures are also applied during this process. The output is a 3D model of the three-dimensional shape.

[0096] Step 5:

[0097] The server adds motion to the generated 3D model. The input is the generated 3D model, to which motion capture data is applied to add movement. The output is the 3D model with motion.

[0098] Step 6:

[0099] The server adjusts color and image clarity to improve visual effects. The input is a 3D model with animations, which is adjusted using Adobe Color and Photoshop. The output is a visually adjusted 3D model.

[0100] Step 7:

[0101] The server analyzes the user's location information and calculates the optimal placement in the augmented reality environment. The inputs are location information and a pre-calibrated 3D model. Placement optimization is performed using ARKit or ARCore. The output is the pre-calibrated 3D model.

[0102] Step 8:

[0103] The terminal dynamically sends 3D model data with pre-calculated placement, received from the server, to the display device. The input is the pre-calculated 3D model. It is displayed overlaid on the real world via an AR device. The output is 3D content existing in an augmented reality environment.

[0104] (Application Example 1)

[0105] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0106] While technologies existed to convert user-generated illustrations into 3D shapes, systems capable of real-time customization to meet user needs were limited. Furthermore, there was a lack of easy ways for users to test their designs in a virtual space, making an intuitive and interactive experience difficult. Therefore, there was a need to realize a consistent process from illustration creation to design confirmation.

[0107] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0108] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a 3D shape, and means for adding animation to the generated 3D shape. This allows the user to instantly try out their designs in a virtual space and interactively check their appearance.

[0109] "Means for receiving user-generated illustrations" refers to the process of receiving digital images created by users from their devices and transmitting them to the system.

[0110] "A method for extracting features from illustrations and automatically generating 3D shapes" refers to a technology that analyzes shape and color information from received image data and generates three-dimensional objects based on that information.

[0111] "Methods for adding animation to 3D shapes" refers to the process of adding movement to generated three-dimensional shapes to realize visually dynamic expressions.

[0112] "Means of displaying as augmented reality" refers to technologies that overlay virtual objects onto real-world environments, providing users with a new visual experience.

[0113] "A method of placing a design in a virtual environment and testing its appearance" refers to a method of placing a user's design in a digital space and simulating its look and impression.

[0114] "Means of providing an interface" refers to a mechanism that supports users in intuitively operating the system through a screen and utilizing the desired functions.

[0115] The system for implementing this invention allows users to convert illustrations they have drawn into 3D shapes in a virtual environment and then test their appearance. Users first create illustrations using digital painting tools on their smartphones or tablets. If necessary, they can also digitize and import illustrations drawn on paper.

[0116] The user's device sends the created illustration to the server. The server receives the illustration and first uses a Python script to perform shape analysis using the OpenCV library. This analysis extracts shape and color information, and based on this data, it automatically generates a 3D shape using the Blender API.

[0117] The generated 3D model is animated using Unity's AR Foundation and ARKit / ARCore, and then displayed on the user's device in an augmented reality environment. During this process, the server obtains location information from the user's device and calculates the optimal placement of the 3D shape. This entire process is designed to provide a smooth interactive experience across the entire system.

[0118] As a concrete example, users can draw illustrations of clothing items based on their own designs, and these illustrations are instantly placed in a virtual environment, allowing for visual confirmation and feedback on the designs. This system enables users to enjoy a real-time customization experience utilizing the digital space.

[0119] An example of a prompt message might be, "Design a stylish T-shirt, and it will be virtually rendered in 3D. Please share your design ideas. For example, a beach-themed resort design." This prompt message stimulates the user's creativity and promotes further design innovation.

[0120] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0121] Step 1:

[0122] The user's device launches a digital painting tool, and the user draws a T-shirt design.

[0123] Input: User-drawn design data

[0124] Output: Digitized illustration file

[0125] Operation: Users use the application's in-app painting tools to add lines and colors directly to the digital canvas and complete their designs.

[0126] Step 2:

[0127] Send an illustration of the terminal to the server.

[0128] Input: Digitized illustration files

[0129] Output: Notification that the transfer of illustration data to the server is complete.

[0130] Operation: Data is sent from the terminal to the server using a secure protocol. Once the transmission is complete, the server notifies the client that it has received the data.

[0131] Step 3:

[0132] The server analyzes the illustration data, extracts features, and generates a 3D shape.

[0133] Input: Received illustration data

[0134] Output: 3D shape data

[0135] Operation: The received data is analyzed using OpenCV in a Python script to extract shapes and patterns. Then, a 3D model based on these features is generated using the Blender API.

[0136] Step 4:

[0137] The server generates a 3D shape and then adds animation to it.

[0138] Input: 3D shape data

[0139] Output: Animated 3D shape data

[0140] Operation: To apply preset animations to shape data, the Unity engine's animation tools are used to add movement and deformation to the generated 3D model.

[0141] Step 5:

[0142] The server calculates the optimal augmented reality placement based on location information and sends the data to the user's device.

[0143] Input: Animated 3D shape data, user position information

[0144] Output: 3D shape data with placement completed on the terminal

[0145] Operation: The server analyzes location information from the terminal, calculates the optimal placement of the model at the user's current location, and then sends the 3D model data to the device.

[0146] Step 6:

[0147] The user's device uses AR functionality to display animated 3D shapes in the real environment.

[0148] Input: Completed 3D shape data

[0149] Output: Display of 3D shapes within the user's field of view.

[0150] Operation: The device utilizes Unity's AR Foundation to display virtual objects as if they were integrated into the real world through the device's camera, providing the user with an interactive experience.

[0151] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0152] The system implementing this invention generates a 3D shape based on an illustration drawn by the user, displays augmented reality content with added animation, and integrates an emotion engine that recognizes the user's emotions. This system is realized through the cooperation of the user's terminal, a central server, and a device with AR capabilities.

[0153] First, the user creates an illustration within the application on their device. This illustration is imported into the device via a digital painting tool or external image input. Next, this illustration is sent to a server, which extracts the necessary features from the received illustration and automatically generates a 3D shape. At the same time, the user's emotions are also analyzed in real time by an emotion engine using camera and sensor data.

[0154] The server adds emotion-appropriate effects and actions to the generated 3D shapes and animations based on the recognized user's emotional information. For example, if the emotion engine detects "joy," the 3D shape may be corrected with brighter colors, or joyful movements may be added to the animation.

[0155] Furthermore, the system automatically applies a beautification process to improve the visual quality of the illustrations. This beautification process also adjusts the color tone and sharpness based on the analysis results of the emotion engine, ensuring that the visuals are appropriate to the user's emotional state. The completed AR content is then optimally positioned based on the user's location information and displayed on the AR device through the terminal.

[0156] For example, if a user uses the system while feeling down, the emotion engine will recognize that emotion, and encouraging messages will appear on the generated 3D shape, or warm colors will be added to the background. This allows users to receive a personalized AR experience that is tailored to their emotions.

[0157] In this way, this system supports the creation of creative and interactive AR content, enabling users to experience things that resonate with their emotions.

[0158] The following describes the processing flow.

[0159] Step 1:

[0160] Users create illustrations using applications on their devices. They select a digital painting tool, choose colors and brush types, and begin designing. They freely express themselves using their fingers or a pen on the screen of their smartphone or tablet.

[0161] Step 2:

[0162] The drawn illustration is sent to the server by the device. During transmission, the illustration data is compressed and converted to a format suitable for smooth communication (e.g., JPEG or PNG).

[0163] Step 3:

[0164] The server receives the illustration. The server uses image analysis technology to extract the color, shape, and texture features contained in the illustration.

[0165] Step 4:

[0166] The server automatically generates a 3D shape based on the extracted features. Using AI technology, the most suitable 3D model is created according to the characteristics of the illustration.

[0167] Step 5:

[0168] The server uses camera and sensor data provided by the user to operate the emotion engine. It analyzes the user's emotions from their facial expressions and voice.

[0169] Step 6:

[0170] Based on the results recognized by the emotion engine, the server adds animations and effects to the 3D shape. For example, if the user appears happy, brighter colors and bouncy movements are added to the 3D model.

[0171] Step 7:

[0172] The server automatically adjusts the color tone and sharpness of the illustrations, performing beautification processing. Hues and brightness are applied according to the emotion, adjusting them to enhance their visual appeal.

[0173] Step 8:

[0174] The server checks the user's location and calculates coordinates to optimally position the 3D shapes in augmented reality. This ensures that the placement is natural and easy to view in the user's real-world environment.

[0175] Step 9:

[0176] The completed AR content is sent from the server to the device. The device then interacts with an AR device (e.g., AR glasses) to display a 3D shape superimposed onto the real world. The user can then observe and experience this.

[0177] (Example 2)

[0178] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0179] This invention aims to realize an augmented reality experience that reflects the user's emotional state by generating three-dimensional data based on a two-dimensional image drawn by the user and by adding actions to the generated data. Conventional technologies have struggled to generate interactive content that takes into account the individual emotions of users, posing challenges in providing a personalized experience. Furthermore, the placement and display of augmented reality content have not been optimized, sometimes detracting from immersion depending on the usage scenario.

[0180] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0181] In this invention, the server includes means for receiving images created by the user, means for extracting characteristics from the received images and automatically generating three-dimensional data, means for analyzing the user's state, and means for adding actions and effects to the generated three-dimensional data according to the analyzed user state. This enables an interactive and personalized augmented reality experience based on the user's emotions.

[0182] "User-generated images" refer to visual information generated or imported by users using digital tools.

[0183] "Extracting characteristics" refers to the process of identifying and analyzing important elements and features from received image data.

[0184] "Three-dimensional data" refers to a dataset that defines three-dimensional shapes generated from two-dimensional information.

[0185] "Adding motion" refers to the process of adding animation or movement attributes to three-dimensional data.

[0186] "Analyzing the user's state" means recognizing and digitizing the user's emotions and actions through sensors and algorithms.

[0187] "Adding effects" means applying saturation, motion, and other visual changes to 3D data and animations based on the results of user sentiment analysis.

[0188] "Displaying as augmented reality" means overlaying digital data onto a real-world environment for a visual presentation.

[0189] "Analyzing location data" means analyzing the user's current location and location information to determine the optimal placement of digital content.

[0190] This invention is a system for realizing an interactive and personalized augmented reality experience using user-generated images. Users create images using digital painting applications. Specifically, general-purpose painting software or image processing applications may be used. This image is drawn on the user's device, and the device sends the image data to the server via secure communication. The server extracts the characteristics of the received image using an image processing library and generates three-dimensional data.

[0191] Image processing and machine learning libraries such as OpenCV and TensorFlow are used in this process. Furthermore, the server acquires data from the terminal's camera and microphone in real time to recognize the user's state and performs analysis using an emotion engine. Commonly available emotion recognition software is used for this analysis. Based on the analysis results, actions and effects corresponding to the user's emotions are added to the generated three-dimensional data.

[0192] Additional effects are reflected, for example, in the color scheme and movement style. Ultimately, the server uses the user's location information to position the generated augmented reality content optimally and displays it on the AR device through the terminal.

[0193] For example, if a user draws an illustration expressing their mood, the system can provide a colorful, dynamic, three-dimensional animation that reflects that emotion. An example of a prompt is as follows: "Generate interactive AR content to cheer up a depressed user. The background should be in calming colors and include an encouraging message." In this way, the present invention provides a creative and emotionally responsive augmented reality experience.

[0194] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0195] Step 1:

[0196] The user draws an image using a digital painting application on their device. The user can freely edit elements such as color and shape. The created image is saved on the device and sent to the server via a communication module. The input is the image drawn by the user, and the output is the digital image data sent to the server.

[0197] Step 2:

[0198] The server extracts features from the received image data. The server uses OpenCV and TensorFlow to perform image analysis, including edge detection and color distribution analysis. This analysis extracts important image features as data. The input is the received digital image data, and the output is a dataset of the extracted features.

[0199] Step 3:

[0200] The server generates three-dimensional data based on the extracted feature data. Three-dimensional modeling software such as Blender or Unity is used to create appropriate vertex information and mesh structures. A shape generation algorithm is applied during this process. The input is a dataset of extracted features, and the output is the generated three-dimensional model data.

[0201] Step 4:

[0202] The device uses a camera and microphone to capture the user's state and sends the data to a server. The input consists of images of the user's facial expressions and audio data, while the output is the dataset sent to the server.

[0203] Step 5:

[0204] The server analyzes the transmitted user state data. Using an emotion recognition engine, it estimates the user's emotions in real time. The input is a dataset of transmitted user state data, and the output is the analyzed emotion information.

[0205] Step 6:

[0206] The server adds actions and effects to the 3D model based on the user's analyzed sentiment information. This includes adjusting color tones and adding animations. The input is the generated 3D model data and analyzed sentiment information, and the output is the updated 3D content.

[0207] Step 7:

[0208] The server appropriately positions the updated 3D content based on the user's location information and sends it to the terminal. The terminal displays the content using an AR device, providing the user with an augmented reality experience. The input is the updated 3D content data and the user's location information, and the output is the augmented reality content displayed on the AR device.

[0209] (Application Example 2)

[0210] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0211] In physical stores, there is a need for methods that allow customers to visually enjoy selecting products based on their own drawings. However, conventional technologies have made it difficult to provide such an interactive and personalized experience in real time. Furthermore, dynamic adjustments of shapes and colors in response to customer emotions have not been considered, making it impossible to provide a deeper user experience.

[0212] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0213] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a three-dimensional shape, means for adding animation to the generated three-dimensional shape, and means for recognizing the user's emotions and adjusting the three-dimensional shape and animation according to those emotions. This allows customers to experience product selection using illustrations they have drawn themselves in a physical store using augmented reality, and further enjoy personalized visual effects that respond to their emotions.

[0214] A "user" is someone who uses the system to create illustrations or to visually experience augmented reality content.

[0215] An "illustration" refers to a two-dimensional image or drawing created or input by a user.

[0216] "Features" refer to important patterns and shapes extracted from the received illustration, and are essential elements for generating three-dimensional shapes.

[0217] A "3D shape" is a three-dimensional digital object that is automatically generated based on the characteristics of an illustration.

[0218] "Animation" refers to visual effects that involve movement or change added to three-dimensional shapes.

[0219] "Emotions" refer to information that indicates the user's psychological state, and the system uses this information to provide a personalized experience by recognizing it.

[0220] Augmented reality is a visual experience that overlays digital information onto images of the real world.

[0221] This system is realized through the cooperation of the user, terminal, and server. First, the user creates an illustration using a terminal with a dedicated application installed. Digital painting tools and image input functions can be used to create the illustration. This illustration is then sent from the terminal to the server.

[0222] The server extracts key features from received illustrations and automatically generates 3D shapes based on them. This generation process utilizes machine learning libraries such as TensorFlow. Furthermore, the server is equipped with an emotion engine that analyzes the user's emotions in real time via cameras and sensors. Based on this emotion information, personalized effects are added to the 3D shapes and animations.

[0223] On the user's device, the completed 3D shape is displayed as augmented reality. The displayed AR content reflects the analysis results of an emotion recognition module linked with OpenCV, and is accompanied by visual effects that match the user's emotions. Once the content is displayed, the user can experience a real-time, changing product display in a physical store through devices such as smart glasses.

[0224] For example, if a user draws an illustration of their pet at a pet supply store, a 3D model of the pet is created based on that sketch and displayed on smart glasses. If the user shows a joyful expression upon seeing the pet, the system recognizes that emotion and adds an animation that makes the pet appear to move.

[0225] Examples of prompts to input into a generative AI model are as follows:

[0226] "At a pet supply store, users are drawing illustrations of pets. Please generate 3D models based on these sketches and add animations that express the pet's excitement when the user smiles."

[0227] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0228] Step 1:

[0229] The user creates an illustration using their device. They draw the illustration using the device's digital painting tool and save it as a digital image. The input is the illustration drawn by the user on the screen, and the output is a digital image file.

[0230] Step 2:

[0231] The terminal sends the created illustration to the server. The input is a digital image file, which is transferred to the server's receiving port. The output is image data stored on the server.

[0232] Step 3:

[0233] The server extracts features from received illustrations and generates 3D shapes. Image processing algorithms are used to extract important shape data, and 3D modeling is performed based on this data. The input is received image data, and the output is 3D shape data.

[0234] Step 4:

[0235] The server analyzes the user's emotions using cameras and sensors. The emotion analysis engine processes the image and sensor data to estimate the emotional state from the user's facial expressions and movements. The input is the user's facial expressions and movement data, and the output is the estimated emotion information.

[0236] Step 5:

[0237] The server adds animation and effects to 3D shapes based on emotional information. It utilizes a generative AI model to apply emotion-dependent actions and color changes to the model. Input is 3D shape data and emotional information, and output is animated 3D shape data.

[0238] Step 6:

[0239] The device displays animated 3D shapes as augmented reality. The data is transferred to an AR device and overlaid onto a display in a physical store, integrating it into the user's field of view. The input is animated 3D shape data, and the output is augmented reality imagery within the user's field of view.

[0240] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0241] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0242] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0243] [Second Embodiment]

[0244] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0245] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0246] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0247] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0248] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0249] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0250] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0251] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0252] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0253] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0254] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0255] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0256] The system implementing this invention receives illustrations drawn by a user, converts them into three-dimensional shapes and animated content, and displays them as augmented reality. This system is realized through the coordinated operation of the user's terminal, a central server, and a device equipped with AR capabilities.

[0257] First, users create illustrations on their devices using a dedicated application. These illustrations can be drawn directly using digital painting tools, or they can be imported by taking a picture of a drawing on paper.

[0258] Next, the device sends this illustration data to the server. The server analyzes the features of the illustration and extracts its shape and color information. Based on this information, the server uses AI technology to automatically generate a 3D shape. The server then adds pre-prepared animations to the generated 3D model. Finally, it adjusts the illustration's color tone and sharpness to make it visually appealing.

[0259] Furthermore, to enable augmented reality display, the server acquires the user's location information and calculates the optimal placement of 3D shapes for the AR device. The constructed AR content is then sent back to the user's device and finally displayed in the user's real space via the AR device.

[0260] As a concrete example, consider using an illustration of an animal drawn by a child. The user draws an illustration of an animal on their device and sends it to the server. The server analyzes the outline and colors of the received illustration and generates a 3D model of the animal. This model is then given animations such as walking and jumping. The completed AR content appears before the user's eyes through their AR glasses, just as if it were real. This allows the user to observe the animal moving around in their own physical environment.

[0261] This system provides a way for ordinary users to easily create 3D content and enjoy interactive experiences through AR technology.

[0262] The following describes the processing flow.

[0263] Step 1:

[0264] The user draws illustrations on their device. They launch a dedicated app and create illustrations directly using digital painting tools, or they take a picture of a drawing on paper with the device's camera and import it into the app.

[0265] Step 2:

[0266] The terminal sends the completed illustration data to the server. The transmission format is a standard image format such as JPEG or PNG, and data compression is performed to improve communication efficiency.

[0267] Step 3:

[0268] The server receives the transmitted illustration. After receiving it, it analyzes the shape and color of the illustration and extracts feature data to create a 3D shape.

[0269] Step 4:

[0270] The server automatically generates the 3D shape of the illustration using AI technology based on the analysis data. This generation process is carried out by combining existing 3D model libraries and generation algorithms.

[0271] Step 5:

[0272] The server adds animation to the generated 3D shape. The animation uses a pre-configured template and is adjusted to make the model move naturally.

[0273] Step 6:

[0274] The server automatically adjusts the color tone and sharpness and applies beautification processing to improve the visual quality of the illustrations. This enhances their final appearance.

[0275] Step 7:

[0276] The server acquires the user's location information and analyzes the 3D shape at a position and scale suitable for AR display. Based on the location analysis, the placement is adjusted to match the user's real-world environment.

[0277] Step 8:

[0278] The server sends the created AR content data to the user's device.

[0279] Step 9:

[0280] The terminal displays the received AR content through a device specified by the user, such as AR glasses or a smartphone. As a result, the user can enjoy an extended experience in the real world.

[0281] (Example 1)

[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0283] Currently, there are limited ways for users to easily convert their drawn illustrations into three-dimensional shapes or animated content and expand and experience it in the real space. In addition, there is a lack of functions to improve visual effects in a form suitable for the real world, and it is difficult for users to realize the ideal visual experience, which is an issue.

[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0285] In this invention, the server includes means for receiving an image drawn by a user using a generation device, means for extracting the features of the received image by an analysis device and automatically generating a three-dimensional shape, and means for imparting an operation to the generated three-dimensional shape. As a result, the user can easily turn their illustration into a three-dimensional form and expand and experience it with operations in the real space.

[0286] "User" refers to an individual or group that operates a specific electronic device or system.

[0287] "Generation device" refers to a computer system or device used by a user to create and edit visual information.

[0288] "Image" refers to a visual representation created by a user through a generation device.

[0289] "Means of receiving" refers to a device or system that has the function of acquiring information from an external source via data communication.

[0290] "Analyzing features" refers to performing computational processing to identify useful patterns, shapes, colors, and other elements from data.

[0291] "Three-dimensional shape" refers to information that describes the shape of an object that exists in three-dimensional space.

[0292] "Adding action" refers to setting a series of actions, such as movement or deformation, for a three-dimensional shape.

[0293] Augmented reality refers to a technology that integrates and displays digital content within a real-world physical environment.

[0294] The "visual environment" refers to the physical and digital spaces that a user can perceive through their vision.

[0295] "Automatically adjusting color and image clarity" refers to algorithmic processing that appropriately corrects the color tone and resolution of digital images.

[0296] "Improving visual effects" refers to techniques used to enhance the beauty and appearance of displayed content.

[0297] "Analyzing location information" refers to processing and understanding data based on the user's geographical location.

[0298] "Appropriate placement" refers to arranging digital content in a reasonable manner within a physical or virtual space.

[0299] This invention is a system that converts user-drawn images into three-dimensional shapes and animated content, and displays them in a visual environment as augmented reality. It is realized through the cooperation of the user, terminal, and server.

[0300] Users create images using digital painting tools with a dedicated generator. They can also digitize images drawn on paper by taking a picture of it with the device's camera. Once the image is complete, the device sends the data to the server. The data is encrypted during transmission to ensure secure reception.

[0301] The server analyzes the received image data using an analysis device. This process utilizes AI technology to extract image contours and color information, and identify necessary features. Libraries such as OpenCV and TensorFlow are used for image analysis. Based on the analysis results, the server automatically generates a 3D shape. 3D modeling tools such as Blender and Unity are used for 3D shape generation.

[0302] The generated 3D shapes are given motion using motion capture data, making them appear to move. The server also automatically adjusts the colors and image clarity to improve the appearance. Tools such as Adobe Color and Photoshop are used in this process.

[0303] Furthermore, the server analyzes location information and appropriately positions augmented reality content based on the device's physical location. It uses ARKit or ARCore to perform calculations for placing 3D models in real space.

[0304] Ultimately, the device displays augmented reality content in the visual environment via an AR device, based on data received from the server. For example, an image of an animal drawn by the user can be rendered in 3D and displayed in the living room with animations.

[0305] As an example of a prompt sentence, it is possible to input an instruction such as "Convert the illustration of the marine creature drawn by the user into a 3D model and make it swim in the AR space." into the server. By doing so, the user can vividly experience their imagination in the real space.

[0306] The flow of the specific process in Example 1 will be described with reference to FIG. 11.

[0307] Step 1:

[0308] The user activates the generation device and creates an image using a digital painting tool. The inputs include brush strokes and color selections from the drawing tool. When the user finishes drawing the picture, it is saved in a digital file format (e.g., PNG, JPEG). The output is a digital format image file.

[0309] Step 2:

[0310] The terminal sends the image file created by the user to the server. The input is a digital format image file, and to send this, the terminal encrypts the data using an appropriate protocol. The output is sent to the server as encrypted image data.

[0311] Step 3:

[0312] The server analyzes the received image data. The input is encrypted digital image data, which is decrypted and the features of the outline and color are extracted using AI technology (e.g., image recognition model). Specifically, image processing is performed using OpenCV or TensorFlow. The output is feature data for generating a three-dimensional shape.

[0313] Step 4:

[0314] The server automatically generates a 3D shape based on extracted feature data. The input is feature data. Using the generated AI model, a 3D model is created using tools such as Blender or Unity. Textures are also applied during this process. The output is a 3D model of the three-dimensional shape.

[0315] Step 5:

[0316] The server adds motion to the generated 3D model. The input is the generated 3D model, to which motion capture data is applied to add movement. The output is the 3D model with motion.

[0317] Step 6:

[0318] The server adjusts color and image clarity to improve visual effects. The input is a 3D model with animations, which is adjusted using Adobe Color and Photoshop. The output is a visually adjusted 3D model.

[0319] Step 7:

[0320] The server analyzes the user's location information and calculates the optimal placement in the augmented reality environment. The inputs are location information and a pre-calibrated 3D model. Placement optimization is performed using ARKit or ARCore. The output is the pre-calibrated 3D model.

[0321] Step 8:

[0322] The terminal dynamically sends 3D model data with pre-calculated placement, received from the server, to the display device. The input is the pre-calculated 3D model. It is displayed overlaid on the real world via an AR device. The output is 3D content existing in an augmented reality environment.

[0323] (Application Example 1)

[0324] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0325] While technologies existed to convert user-generated illustrations into 3D shapes, systems capable of real-time customization to meet user needs were limited. Furthermore, there was a lack of easy ways for users to test their designs in a virtual space, making an intuitive and interactive experience difficult. Therefore, there was a need to realize a consistent process from illustration creation to design confirmation.

[0326] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0327] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a 3D shape, and means for adding animation to the generated 3D shape. This allows the user to instantly try out their designs in a virtual space and interactively check their appearance.

[0328] "Means for receiving user-generated illustrations" refers to the process of receiving digital images created by users from their devices and transmitting them to the system.

[0329] "A method for extracting features from illustrations and automatically generating 3D shapes" refers to a technology that analyzes shape and color information from received image data and generates three-dimensional objects based on that information.

[0330] "Methods for adding animation to 3D shapes" refers to the process of adding movement to generated three-dimensional shapes to realize visually dynamic expressions.

[0331] "Means of displaying as augmented reality" refers to technologies that overlay virtual objects onto real-world environments, providing users with a new visual experience.

[0332] "A method of placing a design in a virtual environment and testing its appearance" refers to a method of placing a user's design in a digital space and simulating its look and impression.

[0333] "Means of providing an interface" refers to a mechanism that supports users in intuitively operating the system through a screen and utilizing the desired functions.

[0334] The system for implementing this invention allows users to convert illustrations they have drawn into 3D shapes in a virtual environment and then test their appearance. Users first create illustrations using digital painting tools on their smartphones or tablets. If necessary, they can also digitize and import illustrations drawn on paper.

[0335] The user's device sends the created illustration to the server. The server receives the illustration and first uses a Python script to perform shape analysis using the OpenCV library. This analysis extracts shape and color information, and based on this data, it automatically generates a 3D shape using the Blender API.

[0336] The generated 3D model is animated using Unity's AR Foundation and ARKit / ARCore, and then displayed on the user's device in an augmented reality environment. During this process, the server obtains location information from the user's device and calculates the optimal placement of the 3D shape. This entire process is designed to provide a smooth interactive experience across the entire system.

[0337] As a concrete example, users can draw illustrations of clothing items based on their own designs, and these illustrations are instantly placed in a virtual environment, allowing for visual confirmation and feedback on the designs. This system enables users to enjoy a real-time customization experience utilizing the digital space.

[0338] An example of a prompt message might be, "Design a stylish T-shirt, and it will be virtually rendered in 3D. Please share your design ideas. For example, a beach-themed resort design." This prompt message stimulates the user's creativity and promotes further design innovation.

[0339] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0340] Step 1:

[0341] The user's device launches a digital painting tool, and the user draws a T-shirt design.

[0342] Input: User-drawn design data

[0343] Output: Digitized illustration file

[0344] Operation: Users use the application's in-app painting tools to add lines and colors directly to the digital canvas and complete their designs.

[0345] Step 2:

[0346] Send an illustration of the terminal to the server.

[0347] Input: Digitized illustration files

[0348] Output: Notification that the transfer of illustration data to the server is complete.

[0349] Operation: Data is sent from the terminal to the server using a secure protocol. Once the transmission is complete, the server notifies the client that it has received the data.

[0350] Step 3:

[0351] The server analyzes the illustration data, extracts features, and generates a 3D shape.

[0352] Input: Received illustration data

[0353] Output: 3D shape data

[0354] Operation: The received data is analyzed using OpenCV in a Python script to extract shapes and patterns. Then, a 3D model based on these features is generated using the Blender API.

[0355] Step 4:

[0356] The server generates a 3D shape and then adds animation to it.

[0357] Input: 3D shape data

[0358] Output: Animated 3D shape data

[0359] Operation: To apply preset animations to shape data, the Unity engine's animation tools are used to add movement and deformation to the generated 3D model.

[0360] Step 5:

[0361] The server calculates the optimal augmented reality placement based on location information and sends the data to the user's device.

[0362] Input: Animated 3D shape data, user position information

[0363] Output: 3D shape data with placement completed on the terminal

[0364] Operation: The server analyzes location information from the terminal, calculates the optimal placement of the model at the user's current location, and then sends the 3D model data to the device.

[0365] Step 6:

[0366] The user's device uses AR functionality to display animated 3D shapes in the real environment.

[0367] Input: Completed 3D shape data

[0368] Output: Display of 3D shapes within the user's field of view.

[0369] Operation: The device utilizes Unity's AR Foundation to display virtual objects as if they were integrated into the real world through the device's camera, providing the user with an interactive experience.

[0370] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0371] The system implementing this invention generates a 3D shape based on an illustration drawn by the user, displays augmented reality content with added animation, and integrates an emotion engine that recognizes the user's emotions. This system is realized through the cooperation of the user's terminal, a central server, and a device with AR capabilities.

[0372] First, the user creates an illustration within the application on their device. This illustration is imported into the device via a digital painting tool or external image input. Next, this illustration is sent to a server, which extracts the necessary features from the received illustration and automatically generates a 3D shape. At the same time, the user's emotions are also analyzed in real time by an emotion engine using camera and sensor data.

[0373] The server adds emotion-appropriate effects and actions to the generated 3D shapes and animations based on the recognized user's emotional information. For example, if the emotion engine detects "joy," the 3D shape may be corrected with brighter colors, or joyful movements may be added to the animation.

[0374] Furthermore, the system automatically applies a beautification process to improve the visual quality of the illustrations. This beautification process also adjusts the color tone and sharpness based on the analysis results of the emotion engine, ensuring that the visuals are appropriate to the user's emotional state. The completed AR content is then optimally positioned based on the user's location information and displayed on the AR device through the terminal.

[0375] For example, if a user uses the system while feeling down, the emotion engine will recognize that emotion, and encouraging messages will appear on the generated 3D shape, or warm colors will be added to the background. This allows users to receive a personalized AR experience that is tailored to their emotions.

[0376] In this way, this system supports the creation of creative and interactive AR content, enabling users to experience things that resonate with their emotions.

[0377] The following describes the processing flow.

[0378] Step 1:

[0379] Users create illustrations using applications on their devices. They select a digital painting tool, choose colors and brush types, and begin designing. They freely express themselves using their fingers or a pen on the screen of their smartphone or tablet.

[0380] Step 2:

[0381] The drawn illustration is sent to the server by the device. During transmission, the illustration data is compressed and converted to a format suitable for smooth communication (e.g., JPEG or PNG).

[0382] Step 3:

[0383] The server receives the illustration. The server uses image analysis technology to extract the color, shape, and texture features contained in the illustration.

[0384] Step 4:

[0385] The server automatically generates a 3D shape based on the extracted features. Using AI technology, the most suitable 3D model is created according to the characteristics of the illustration.

[0386] Step 5:

[0387] The server uses camera and sensor data provided by the user to operate the emotion engine. It analyzes the user's emotions from their facial expressions and voice.

[0388] Step 6:

[0389] Based on the results recognized by the emotion engine, the server adds animations and effects to the 3D shape. For example, if the user appears happy, brighter colors and bouncy movements are added to the 3D model.

[0390] Step 7:

[0391] The server automatically adjusts the color tone and sharpness of the illustrations, performing beautification processing. Hues and brightness are applied according to the emotion, adjusting them to enhance their visual appeal.

[0392] Step 8:

[0393] The server checks the user's location and calculates coordinates to optimally position the 3D shapes in augmented reality. This ensures that the placement is natural and easy to view in the user's real-world environment.

[0394] Step 9:

[0395] The completed AR content is sent from the server to the device. The device then interacts with an AR device (e.g., AR glasses) to display a 3D shape superimposed onto the real world. The user can then observe and experience this.

[0396] (Example 2)

[0397] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0398] This invention aims to realize an augmented reality experience that reflects the user's emotional state by generating three-dimensional data based on a two-dimensional image drawn by the user and by adding actions to the generated data. Conventional technologies have struggled to generate interactive content that takes into account the individual emotions of users, posing challenges in providing a personalized experience. Furthermore, the placement and display of augmented reality content have not been optimized, sometimes detracting from immersion depending on the usage scenario.

[0399] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0400] In this invention, the server includes means for receiving images created by the user, means for extracting characteristics from the received images and automatically generating three-dimensional data, means for analyzing the user's state, and means for adding actions and effects to the generated three-dimensional data according to the analyzed user state. This enables an interactive and personalized augmented reality experience based on the user's emotions.

[0401] "User-generated images" refer to visual information generated or imported by users using digital tools.

[0402] "Extracting characteristics" refers to the process of identifying and analyzing important elements and features from received image data.

[0403] "Three-dimensional data" refers to a dataset that defines three-dimensional shapes generated from two-dimensional information.

[0404] "Adding motion" refers to the process of adding animation or movement attributes to three-dimensional data.

[0405] "Analyzing the user's state" means recognizing and digitizing the user's emotions and actions through sensors and algorithms.

[0406] "Adding effects" means applying saturation, motion, and other visual changes to 3D data and animations based on the results of user sentiment analysis.

[0407] "Displaying as augmented reality" means overlaying digital data onto a real-world environment for a visual presentation.

[0408] "Analyzing location data" means analyzing the user's current location and location information to determine the optimal placement of digital content.

[0409] This invention is a system for realizing an interactive and personalized augmented reality experience using user-generated images. Users create images using digital painting applications. Specifically, general-purpose painting software or image processing applications may be used. This image is drawn on the user's device, and the device sends the image data to the server via secure communication. The server extracts the characteristics of the received image using an image processing library and generates three-dimensional data.

[0410] Image processing and machine learning libraries such as OpenCV and TensorFlow are used in this process. Furthermore, the server acquires data from the terminal's camera and microphone in real time to recognize the user's state and performs analysis using an emotion engine. Commonly available emotion recognition software is used for this analysis. Based on the analysis results, actions and effects corresponding to the user's emotions are added to the generated three-dimensional data.

[0411] Additional effects are reflected, for example, in the color scheme and movement style. Ultimately, the server uses the user's location information to position the generated augmented reality content optimally and displays it on the AR device through the terminal.

[0412] For example, if a user draws an illustration expressing their mood, the system can provide a colorful, dynamic, three-dimensional animation that reflects that emotion. An example of a prompt is as follows: "Generate interactive AR content to cheer up a depressed user. The background should be in calming colors and include an encouraging message." In this way, the present invention provides a creative and emotionally responsive augmented reality experience.

[0413] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0414] Step 1:

[0415] The user draws an image using a digital painting application on their device. The user can freely edit elements such as color and shape. The created image is saved on the device and sent to the server via a communication module. The input is the image drawn by the user, and the output is the digital image data sent to the server.

[0416] Step 2:

[0417] The server extracts features from the received image data. The server uses OpenCV and TensorFlow to perform image analysis, including edge detection and color distribution analysis. This analysis extracts important image features as data. The input is the received digital image data, and the output is a dataset of the extracted features.

[0418] Step 3:

[0419] The server generates three-dimensional data based on the extracted feature data. Three-dimensional modeling software such as Blender or Unity is used to create appropriate vertex information and mesh structures. A shape generation algorithm is applied during this process. The input is a dataset of extracted features, and the output is the generated three-dimensional model data.

[0420] Step 4:

[0421] The device uses a camera and microphone to capture the user's state and sends the data to a server. The input consists of images of the user's facial expressions and audio data, while the output is the dataset sent to the server.

[0422] Step 5:

[0423] The server analyzes the transmitted user state data. Using an emotion recognition engine, it estimates the user's emotions in real time. The input is a dataset of transmitted user state data, and the output is the analyzed emotion information.

[0424] Step 6:

[0425] The server adds actions and effects to the 3D model based on the user's analyzed sentiment information. This includes adjusting color tones and adding animations. The input is the generated 3D model data and analyzed sentiment information, and the output is the updated 3D content.

[0426] Step 7:

[0427] The server appropriately positions the updated 3D content based on the user's location information and sends it to the terminal. The terminal displays the content using an AR device, providing the user with an augmented reality experience. The input is the updated 3D content data and the user's location information, and the output is the augmented reality content displayed on the AR device.

[0428] (Application Example 2)

[0429] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0430] In physical stores, there is a need for methods that allow customers to visually enjoy selecting products based on their own drawings. However, conventional technologies have made it difficult to provide such an interactive and personalized experience in real time. Furthermore, dynamic adjustments of shapes and colors in response to customer emotions have not been considered, making it impossible to provide a deeper user experience.

[0431] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0432] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a three-dimensional shape, means for adding animation to the generated three-dimensional shape, and means for recognizing the user's emotions and adjusting the three-dimensional shape and animation according to those emotions. This allows customers to experience product selection using illustrations they have drawn themselves in a physical store using augmented reality, and further enjoy personalized visual effects that respond to their emotions.

[0433] A "user" is someone who uses the system to create illustrations or to visually experience augmented reality content.

[0434] An "illustration" refers to a two-dimensional image or drawing created or input by a user.

[0435] "Features" refer to important patterns and shapes extracted from the received illustration, and are essential elements for generating three-dimensional shapes.

[0436] A "3D shape" is a three-dimensional digital object that is automatically generated based on the characteristics of an illustration.

[0437] "Animation" refers to visual effects that involve movement or change added to three-dimensional shapes.

[0438] "Emotions" refer to information that indicates the user's psychological state, and the system uses this information to provide a personalized experience by recognizing it.

[0439] Augmented reality is a visual experience that overlays digital information onto images of the real world.

[0440] This system is realized through the cooperation of the user, terminal, and server. First, the user creates an illustration using a terminal with a dedicated application installed. Digital painting tools and image input functions can be used to create the illustration. This illustration is then sent from the terminal to the server.

[0441] The server extracts key features from received illustrations and automatically generates 3D shapes based on them. This generation process utilizes machine learning libraries such as TensorFlow. Furthermore, the server is equipped with an emotion engine that analyzes the user's emotions in real time via cameras and sensors. Based on this emotion information, personalized effects are added to the 3D shapes and animations.

[0442] On the user's device, the completed 3D shape is displayed as augmented reality. The displayed AR content reflects the analysis results of an emotion recognition module linked with OpenCV, and is accompanied by visual effects that match the user's emotions. Once the content is displayed, the user can experience a real-time, changing product display in a physical store through devices such as smart glasses.

[0443] For example, if a user draws an illustration of their pet at a pet supply store, a 3D model of the pet is created based on that sketch and displayed on smart glasses. If the user shows a joyful expression upon seeing the pet, the system recognizes that emotion and adds an animation that makes the pet appear to move.

[0444] Examples of prompts to input into a generative AI model are as follows:

[0445] "At a pet supply store, users are drawing illustrations of pets. Please generate 3D models based on these sketches and add animations that express the pet's excitement when the user smiles."

[0446] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0447] Step 1:

[0448] The user creates an illustration using their device. They draw the illustration using the device's digital painting tool and save it as a digital image. The input is the illustration drawn by the user on the screen, and the output is a digital image file.

[0449] Step 2:

[0450] The terminal sends the created illustration to the server. The input is a digital image file, which is transferred to the server's receiving port. The output is image data stored on the server.

[0451] Step 3:

[0452] The server extracts features from received illustrations and generates 3D shapes. Image processing algorithms are used to extract important shape data, and 3D modeling is performed based on this data. The input is received image data, and the output is 3D shape data.

[0453] Step 4:

[0454] The server analyzes the user's emotions using cameras and sensors. The emotion analysis engine processes the image and sensor data to estimate the emotional state from the user's facial expressions and movements. The input is the user's facial expressions and movement data, and the output is the estimated emotion information.

[0455] Step 5:

[0456] The server adds animation and effects to 3D shapes based on emotional information. It utilizes a generative AI model to apply emotion-dependent actions and color changes to the model. Input is 3D shape data and emotional information, and output is animated 3D shape data.

[0457] Step 6:

[0458] The device displays animated 3D shapes as augmented reality. The data is transferred to an AR device and overlaid onto a display in a physical store, integrating it into the user's field of view. The input is animated 3D shape data, and the output is augmented reality imagery within the user's field of view.

[0459] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0460] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0461] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0462] [Third Embodiment]

[0463] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0464] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0465] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0466] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0467] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0468] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0469] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0470] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0471] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0472] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0473] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0474] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0475] The system implementing this invention receives illustrations drawn by a user, converts them into three-dimensional shapes and animated content, and displays them as augmented reality. This system is realized through the coordinated operation of the user's terminal, a central server, and a device equipped with AR capabilities.

[0476] First, users create illustrations on their devices using a dedicated application. These illustrations can be drawn directly using digital painting tools, or they can be imported by taking a picture of a drawing on paper.

[0477] Next, the device sends this illustration data to the server. The server analyzes the features of the illustration and extracts its shape and color information. Based on this information, the server uses AI technology to automatically generate a 3D shape. The server then adds pre-prepared animations to the generated 3D model. Finally, it adjusts the illustration's color tone and sharpness to make it visually appealing.

[0478] Furthermore, to enable augmented reality display, the server acquires the user's location information and calculates the optimal placement of 3D shapes for the AR device. The constructed AR content is then sent back to the user's device and finally displayed in the user's real space via the AR device.

[0479] As a concrete example, consider using an illustration of an animal drawn by a child. The user draws an illustration of an animal on their device and sends it to the server. The server analyzes the outline and colors of the received illustration and generates a 3D model of the animal. This model is then given animations such as walking and jumping. The completed AR content appears before the user's eyes through their AR glasses, just as if it were real. This allows the user to observe the animal moving around in their own physical environment.

[0480] This system provides a way for ordinary users to easily create 3D content and enjoy interactive experiences through AR technology.

[0481] The following describes the processing flow.

[0482] Step 1:

[0483] The user draws illustrations on their device. They launch a dedicated app and create illustrations directly using digital painting tools, or they take a picture of a drawing on paper with the device's camera and import it into the app.

[0484] Step 2:

[0485] The terminal sends the completed illustration data to the server. The transmission format is a standard image format such as JPEG or PNG, and data compression is performed to improve communication efficiency.

[0486] Step 3:

[0487] The server receives the transmitted illustration. After receiving it, it analyzes the shape and color of the illustration and extracts feature data to create a 3D shape.

[0488] Step 4:

[0489] The server automatically generates the 3D shape of the illustration using AI technology based on the analysis data. This generation process is carried out by combining existing 3D model libraries and generation algorithms.

[0490] Step 5:

[0491] The server adds animation to the generated 3D shape. The animation uses a pre-configured template and is adjusted to make the model move naturally.

[0492] Step 6:

[0493] The server automatically adjusts the color tone and sharpness and applies beautification processing to improve the visual quality of the illustrations. This enhances their final appearance.

[0494] Step 7:

[0495] The server acquires the user's location information and analyzes the 3D shape at a position and scale suitable for AR display. Based on the location analysis, the placement is adjusted to match the user's real-world environment.

[0496] Step 8:

[0497] The server sends the created AR content data to the user's device.

[0498] Step 9:

[0499] The device displays the received AR content through a device specified by the user, such as AR glasses or a smartphone. This allows the user to enjoy an augmented reality experience in the real world.

[0500] (Example 1)

[0501] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] Currently, there are limited ways for users to easily convert their own illustrations into 3D shapes or animated content, and then extend and experience them in real space. Furthermore, there is a lack of features to improve visual effects in a way that is suitable for the real world, making it difficult for users to realize their ideal visual experience.

[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0504] In this invention, the server includes means for receiving an image drawn by a user using a generation device, means for extracting features of the received image using an analysis device and automatically generating a three-dimensional shape, and means for adding movement to the generated three-dimensional shape. This makes it possible for users to easily turn their own illustrations into three-dimensional objects and experience them extended into real space with movement.

[0505] "User" refers to an individual or group that operates a specific electronic device or system.

[0506] "Generating device" refers to a computer system or device used by a user to create and edit visual information.

[0507] "Image" refers to a visual representation created by a user through a generation device.

[0508] "Means of receiving" refers to a device or system that has the function of acquiring information from an external source via data communication.

[0509] "Analyzing features" refers to performing computational processing to identify useful patterns, shapes, colors, and other elements from data.

[0510] "Three-dimensional shape" refers to information that describes the shape of an object that exists in three-dimensional space.

[0511] "Adding action" refers to setting a series of actions, such as movement or deformation, for a three-dimensional shape.

[0512] Augmented reality refers to a technology that integrates and displays digital content within a real-world physical environment.

[0513] The "visual environment" refers to the physical and digital spaces that a user can perceive through their vision.

[0514] "Automatically adjusting color and image clarity" refers to algorithmic processing that appropriately corrects the color tone and resolution of digital images.

[0515] "Improving visual effects" refers to techniques used to enhance the beauty and appearance of displayed content.

[0516] "Analyzing location information" refers to processing and understanding data based on the user's geographical location.

[0517] "Appropriate placement" refers to arranging digital content in a reasonable manner within a physical or virtual space.

[0518] This invention is a system that converts user-drawn images into three-dimensional shapes and animated content, and displays them in a visual environment as augmented reality. It is realized through the cooperation of the user, terminal, and server.

[0519] Users create images using digital painting tools with a dedicated generator. They can also digitize images drawn on paper by taking a picture of it with the device's camera. Once the image is complete, the device sends the data to the server. The data is encrypted during transmission to ensure secure reception.

[0520] The server analyzes the received image data using an analysis device. This process utilizes AI technology to extract image contours and color information, and identify necessary features. Libraries such as OpenCV and TensorFlow are used for image analysis. Based on the analysis results, the server automatically generates a 3D shape. 3D modeling tools such as Blender and Unity are used for 3D shape generation.

[0521] The generated 3D shapes are given motion using motion capture data, making them appear to move. The server also automatically adjusts the colors and image clarity to improve the appearance. Tools such as Adobe Color and Photoshop are used in this process.

[0522] Furthermore, the server analyzes location information and appropriately positions augmented reality content based on the device's physical location. It uses ARKit or ARCore to perform calculations for placing 3D models in real space.

[0523] Ultimately, the device displays augmented reality content in the visual environment via an AR device, based on data received from the server. For example, an image of an animal drawn by the user can be rendered in 3D and displayed in the living room with animations.

[0524] As an example of a prompt, it is possible to input instructions to the server such as, "Convert the illustration of a sea creature drawn by the user into a 3D model and make it swim in the AR space." This allows the user to experience their imagination in three dimensions in the real world.

[0525] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0526] Step 1:

[0527] The user activates the generator and creates an image using digital painting tools. Input includes brush strokes and color selections from the drawing tools. Once the user finishes drawing, they save it as a digital file (e.g., PNG, JPEG). The output is a digital image file.

[0528] Step 2:

[0529] The terminal sends image files created by the user to the server. The input is a digital image file, and the terminal encrypts the data using an appropriate protocol for transmission. The output is sent to the server as encrypted image data.

[0530] Step 3:

[0531] The server analyzes the received image data. The input is encrypted digital image data, which is decrypted, and AI technology (e.g., image recognition models) is used to extract contour and color features. Specifically, image processing is performed using OpenCV or TensorFlow. The output is feature data for generating 3D shapes.

[0532] Step 4:

[0533] The server automatically generates a 3D shape based on extracted feature data. The input is feature data. Using the generated AI model, a 3D model is created using tools such as Blender or Unity. Textures are also applied during this process. The output is a 3D model of the three-dimensional shape.

[0534] Step 5:

[0535] The server adds motion to the generated 3D model. The input is the generated 3D model, to which motion capture data is applied to add movement. The output is the 3D model with motion.

[0536] Step 6:

[0537] The server adjusts color and image clarity to improve visual effects. The input is a 3D model with animations, which is adjusted using Adobe Color and Photoshop. The output is a visually adjusted 3D model.

[0538] Step 7:

[0539] The server analyzes the user's location information and calculates the optimal placement in the augmented reality environment. The inputs are location information and a pre-calibrated 3D model. Placement optimization is performed using ARKit or ARCore. The output is the pre-calibrated 3D model.

[0540] Step 8:

[0541] The terminal dynamically sends 3D model data with pre-calculated placement, received from the server, to the display device. The input is the pre-calculated 3D model. It is displayed overlaid on the real world via an AR device. The output is 3D content existing in an augmented reality environment.

[0542] (Application Example 1)

[0543] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0544] While technologies existed to convert user-generated illustrations into 3D shapes, systems capable of real-time customization to meet user needs were limited. Furthermore, there was a lack of easy ways for users to test their designs in a virtual space, making an intuitive and interactive experience difficult. Therefore, there was a need to realize a consistent process from illustration creation to design confirmation.

[0545] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0546] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a 3D shape, and means for adding animation to the generated 3D shape. This allows the user to instantly try out their designs in a virtual space and interactively check their appearance.

[0547] "Means for receiving user-generated illustrations" refers to the process of receiving digital images created by users from their devices and transmitting them to the system.

[0548] "A method for extracting features from illustrations and automatically generating 3D shapes" refers to a technology that analyzes shape and color information from received image data and generates three-dimensional objects based on that information.

[0549] "Methods for adding animation to 3D shapes" refers to the process of adding movement to generated three-dimensional shapes to realize visually dynamic expressions.

[0550] "Means of displaying as augmented reality" refers to technologies that overlay virtual objects onto real-world environments, providing users with a new visual experience.

[0551] "A method of placing a design in a virtual environment and testing its appearance" refers to a method of placing a user's design in a digital space and simulating its look and impression.

[0552] "Means of providing an interface" refers to a mechanism that supports users in intuitively operating the system through a screen and utilizing the desired functions.

[0553] The system for implementing this invention allows users to convert illustrations they have drawn into 3D shapes in a virtual environment and then test their appearance. Users first create illustrations using digital painting tools on their smartphones or tablets. If necessary, they can also digitize and import illustrations drawn on paper.

[0554] The user's device sends the created illustration to the server. The server receives the illustration and first uses a Python script to perform shape analysis using the OpenCV library. This analysis extracts shape and color information, and based on this data, it automatically generates a 3D shape using the Blender API.

[0555] The generated 3D model is animated using Unity's AR Foundation and ARKit / ARCore, and then displayed on the user's device in an augmented reality environment. During this process, the server obtains location information from the user's device and calculates the optimal placement of the 3D shape. This entire process is designed to provide a smooth interactive experience across the entire system.

[0556] As a concrete example, users can draw illustrations of clothing items based on their own designs, and these illustrations are instantly placed in a virtual environment, allowing for visual confirmation and feedback on the designs. This system enables users to enjoy a real-time customization experience utilizing the digital space.

[0557] An example of a prompt message might be, "Design a stylish T-shirt, and it will be virtually rendered in 3D. Please share your design ideas. For example, a beach-themed resort design." This prompt message stimulates the user's creativity and promotes further design innovation.

[0558] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0559] Step 1:

[0560] The user's device launches a digital painting tool, and the user draws a T-shirt design.

[0561] Input: User-drawn design data

[0562] Output: Digitized illustration file

[0563] Operation: Users use the application's in-app painting tools to add lines and colors directly to the digital canvas and complete their designs.

[0564] Step 2:

[0565] Send an illustration of the terminal to the server.

[0566] Input: Digitized illustration files

[0567] Output: Notification that the transfer of illustration data to the server is complete.

[0568] Operation: Data is sent from the terminal to the server using a secure protocol. Once the transmission is complete, the server notifies the client that it has received the data.

[0569] Step 3:

[0570] The server analyzes the illustration data, extracts features, and generates a 3D shape.

[0571] Input: Received illustration data

[0572] Output: 3D shape data

[0573] Operation: The received data is analyzed using OpenCV in a Python script to extract shapes and patterns. Then, a 3D model based on these features is generated using the Blender API.

[0574] Step 4:

[0575] The server generates a 3D shape and then adds animation to it.

[0576] Input: 3D shape data

[0577] Output: Animated 3D shape data

[0578] Operation: To apply preset animations to shape data, the Unity engine's animation tools are used to add movement and deformation to the generated 3D model.

[0579] Step 5:

[0580] The server calculates the optimal augmented reality placement based on location information and sends the data to the user's device.

[0581] Input: Animated 3D shape data, user position information

[0582] Output: 3D shape data with placement completed on the terminal

[0583] Operation: The server analyzes location information from the terminal, calculates the optimal placement of the model at the user's current location, and then sends the 3D model data to the device.

[0584] Step 6:

[0585] The user's device uses AR functionality to display animated 3D shapes in the real environment.

[0586] Input: Completed 3D shape data

[0587] Output: Display of 3D shapes within the user's field of view.

[0588] Operation: The device utilizes Unity's AR Foundation to display virtual objects as if they were integrated into the real world through the device's camera, providing the user with an interactive experience.

[0589] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0590] The system implementing this invention generates a 3D shape based on an illustration drawn by the user, displays augmented reality content with added animation, and integrates an emotion engine that recognizes the user's emotions. This system is realized through the cooperation of the user's terminal, a central server, and a device with AR capabilities.

[0591] First, the user creates an illustration within the application on their device. This illustration is imported into the device via a digital painting tool or external image input. Next, this illustration is sent to a server, which extracts the necessary features from the received illustration and automatically generates a 3D shape. At the same time, the user's emotions are also analyzed in real time by an emotion engine using camera and sensor data.

[0592] The server adds emotion-appropriate effects and actions to the generated 3D shapes and animations based on the recognized user's emotional information. For example, if the emotion engine detects "joy," the 3D shape may be corrected with brighter colors, or joyful movements may be added to the animation.

[0593] Furthermore, the system automatically applies a beautification process to improve the visual quality of the illustrations. This beautification process also adjusts the color tone and sharpness based on the analysis results of the emotion engine, ensuring that the visuals are appropriate to the user's emotional state. The completed AR content is then optimally positioned based on the user's location information and displayed on the AR device through the terminal.

[0594] For example, if a user uses the system while feeling down, the emotion engine will recognize that emotion, and encouraging messages will appear on the generated 3D shape, or warm colors will be added to the background. This allows users to receive a personalized AR experience that is tailored to their emotions.

[0595] In this way, this system supports the creation of creative and interactive AR content, enabling users to experience things that resonate with their emotions.

[0596] The following describes the processing flow.

[0597] Step 1:

[0598] Users create illustrations using applications on their devices. They select a digital painting tool, choose colors and brush types, and begin designing. They freely express themselves using their fingers or a pen on the screen of their smartphone or tablet.

[0599] Step 2:

[0600] The drawn illustration is sent to the server by the device. During transmission, the illustration data is compressed and converted to a format suitable for smooth communication (e.g., JPEG or PNG).

[0601] Step 3:

[0602] The server receives the illustration. The server uses image analysis technology to extract the color, shape, and texture features contained in the illustration.

[0603] Step 4:

[0604] The server automatically generates a 3D shape based on the extracted features. Using AI technology, the most suitable 3D model is created according to the characteristics of the illustration.

[0605] Step 5:

[0606] The server uses camera and sensor data provided by the user to operate the emotion engine. It analyzes the user's emotions from their facial expressions and voice.

[0607] Step 6:

[0608] Based on the results recognized by the emotion engine, the server adds animations and effects to the 3D shape. For example, if the user appears happy, brighter colors and bouncy movements are added to the 3D model.

[0609] Step 7:

[0610] The server automatically adjusts the color tone and sharpness of the illustrations, performing beautification processing. Hues and brightness are applied according to the emotion, adjusting them to enhance their visual appeal.

[0611] Step 8:

[0612] The server checks the user's location and calculates coordinates to optimally position the 3D shapes in augmented reality. This ensures that the placement is natural and easy to view in the user's real-world environment.

[0613] Step 9:

[0614] The completed AR content is sent from the server to the device. The device then interacts with an AR device (e.g., AR glasses) to display a 3D shape superimposed onto the real world. The user can then observe and experience this.

[0615] (Example 2)

[0616] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0617] This invention aims to realize an augmented reality experience that reflects the user's emotional state by generating three-dimensional data based on a two-dimensional image drawn by the user and by adding actions to the generated data. Conventional technologies have struggled to generate interactive content that takes into account the individual emotions of users, posing challenges in providing a personalized experience. Furthermore, the placement and display of augmented reality content have not been optimized, sometimes detracting from immersion depending on the usage scenario.

[0618] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0619] In this invention, the server includes means for receiving images created by the user, means for extracting characteristics from the received images and automatically generating three-dimensional data, means for analyzing the user's state, and means for adding actions and effects to the generated three-dimensional data according to the analyzed user state. This enables an interactive and personalized augmented reality experience based on the user's emotions.

[0620] "User-generated images" refer to visual information generated or imported by users using digital tools.

[0621] "Extracting characteristics" refers to the process of identifying and analyzing important elements and features from received image data.

[0622] "Three-dimensional data" refers to a dataset that defines three-dimensional shapes generated from two-dimensional information.

[0623] "Adding motion" refers to the process of adding animation or movement attributes to three-dimensional data.

[0624] "Analyzing the user's state" means recognizing and digitizing the user's emotions and actions through sensors and algorithms.

[0625] "Adding effects" means applying saturation, motion, and other visual changes to 3D data and animations based on the results of user sentiment analysis.

[0626] "Displaying as augmented reality" means overlaying digital data onto a real-world environment for a visual presentation.

[0627] "Analyzing location data" means analyzing the user's current location and location information to determine the optimal placement of digital content.

[0628] This invention is a system for realizing an interactive and personalized augmented reality experience using user-generated images. Users create images using digital painting applications. Specifically, general-purpose painting software or image processing applications may be used. This image is drawn on the user's device, and the device sends the image data to the server via secure communication. The server extracts the characteristics of the received image using an image processing library and generates three-dimensional data.

[0629] Image processing and machine learning libraries such as OpenCV and TensorFlow are used in this process. Furthermore, the server acquires data from the terminal's camera and microphone in real time to recognize the user's state and performs analysis using an emotion engine. Commonly available emotion recognition software is used for this analysis. Based on the analysis results, actions and effects corresponding to the user's emotions are added to the generated three-dimensional data.

[0630] Additional effects are reflected, for example, in the color scheme and movement style. Ultimately, the server uses the user's location information to position the generated augmented reality content optimally and displays it on the AR device through the terminal.

[0631] For example, if a user draws an illustration expressing their mood, the system can provide a colorful, dynamic, three-dimensional animation that reflects that emotion. An example of a prompt is as follows: "Generate interactive AR content to cheer up a depressed user. The background should be in calming colors and include an encouraging message." In this way, the present invention provides a creative and emotionally responsive augmented reality experience.

[0632] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0633] Step 1:

[0634] The user draws an image using a digital painting application on their device. The user can freely edit elements such as color and shape. The created image is saved on the device and sent to the server via a communication module. The input is the image drawn by the user, and the output is the digital image data sent to the server.

[0635] Step 2:

[0636] The server extracts features from the received image data. The server uses OpenCV and TensorFlow to perform image analysis, including edge detection and color distribution analysis. This analysis extracts important image features as data. The input is the received digital image data, and the output is a dataset of the extracted features.

[0637] Step 3:

[0638] The server generates three-dimensional data based on the extracted feature data. Three-dimensional modeling software such as Blender or Unity is used to create appropriate vertex information and mesh structures. A shape generation algorithm is applied during this process. The input is a dataset of extracted features, and the output is the generated three-dimensional model data.

[0639] Step 4:

[0640] The device uses a camera and microphone to capture the user's state and sends the data to a server. The input consists of images of the user's facial expressions and audio data, while the output is the dataset sent to the server.

[0641] Step 5:

[0642] The server analyzes the transmitted user state data. Using an emotion recognition engine, it estimates the user's emotions in real time. The input is a dataset of transmitted user state data, and the output is the analyzed emotion information.

[0643] Step 6:

[0644] The server adds actions and effects to the 3D model based on the user's analyzed sentiment information. This includes adjusting color tones and adding animations. The input is the generated 3D model data and analyzed sentiment information, and the output is the updated 3D content.

[0645] Step 7:

[0646] The server appropriately positions the updated 3D content based on the user's location information and sends it to the terminal. The terminal displays the content using an AR device, providing the user with an augmented reality experience. The input is the updated 3D content data and the user's location information, and the output is the augmented reality content displayed on the AR device.

[0647] (Application Example 2)

[0648] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0649] In physical stores, there is a need for methods that allow customers to visually enjoy selecting products based on their own drawings. However, conventional technologies have made it difficult to provide such an interactive and personalized experience in real time. Furthermore, dynamic adjustments of shapes and colors in response to customer emotions have not been considered, making it impossible to provide a deeper user experience.

[0650] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0651] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a three-dimensional shape, means for adding animation to the generated three-dimensional shape, and means for recognizing the user's emotions and adjusting the three-dimensional shape and animation according to those emotions. This allows customers to experience product selection using illustrations they have drawn themselves in a physical store using augmented reality, and further enjoy personalized visual effects that respond to their emotions.

[0652] A "user" is someone who uses the system to create illustrations or to visually experience augmented reality content.

[0653] An "illustration" refers to a two-dimensional image or drawing created or input by a user.

[0654] "Features" refer to important patterns and shapes extracted from the received illustration, and are essential elements for generating three-dimensional shapes.

[0655] A "3D shape" is a three-dimensional digital object that is automatically generated based on the characteristics of an illustration.

[0656] "Animation" refers to visual effects that involve movement or change added to three-dimensional shapes.

[0657] "Emotions" refer to information that indicates the user's psychological state, and the system uses this information to provide a personalized experience by recognizing it.

[0658] Augmented reality is a visual experience that overlays digital information onto images of the real world.

[0659] This system is realized through the cooperation of the user, terminal, and server. First, the user creates an illustration using a terminal with a dedicated application installed. Digital painting tools and image input functions can be used to create the illustration. This illustration is then sent from the terminal to the server.

[0660] The server extracts key features from received illustrations and automatically generates 3D shapes based on them. This generation process utilizes machine learning libraries such as TensorFlow. Furthermore, the server is equipped with an emotion engine that analyzes the user's emotions in real time via cameras and sensors. Based on this emotion information, personalized effects are added to the 3D shapes and animations.

[0661] On the user's device, the completed 3D shape is displayed as augmented reality. The displayed AR content reflects the analysis results of an emotion recognition module linked with OpenCV, and is accompanied by visual effects that match the user's emotions. Once the content is displayed, the user can experience a real-time, changing product display in a physical store through devices such as smart glasses.

[0662] For example, if a user draws an illustration of their pet at a pet supply store, a 3D model of the pet is created based on that sketch and displayed on smart glasses. If the user shows a joyful expression upon seeing the pet, the system recognizes that emotion and adds an animation that makes the pet appear to move.

[0663] Examples of prompts to input into a generative AI model are as follows:

[0664] "At a pet supply store, users are drawing illustrations of pets. Please generate 3D models based on these sketches and add animations that express the pet's excitement when the user smiles."

[0665] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0666] Step 1:

[0667] The user creates an illustration using their device. They draw the illustration using the device's digital painting tool and save it as a digital image. The input is the illustration drawn by the user on the screen, and the output is a digital image file.

[0668] Step 2:

[0669] The terminal sends the created illustration to the server. The input is a digital image file, which is transferred to the server's receiving port. The output is image data stored on the server.

[0670] Step 3:

[0671] The server extracts features from received illustrations and generates 3D shapes. Image processing algorithms are used to extract important shape data, and 3D modeling is performed based on this data. The input is received image data, and the output is 3D shape data.

[0672] Step 4:

[0673] The server analyzes the user's emotions using cameras and sensors. The emotion analysis engine processes the image and sensor data to estimate the emotional state from the user's facial expressions and movements. The input is the user's facial expressions and movement data, and the output is the estimated emotion information.

[0674] Step 5:

[0675] The server adds animation and effects to 3D shapes based on emotional information. It utilizes a generative AI model to apply emotion-dependent actions and color changes to the model. Input is 3D shape data and emotional information, and output is animated 3D shape data.

[0676] Step 6:

[0677] The device displays animated 3D shapes as augmented reality. The data is transferred to an AR device and overlaid onto a display in a physical store, integrating it into the user's field of view. The input is animated 3D shape data, and the output is augmented reality imagery within the user's field of view.

[0678] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0679] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0680] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0681] [Fourth Embodiment]

[0682] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0683] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0684] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0685] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0686] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0687] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0688] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0689] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0690] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0691] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0692] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0693] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0694] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0695] The system implementing this invention receives illustrations drawn by a user, converts them into three-dimensional shapes and animated content, and displays them as augmented reality. This system is realized through the coordinated operation of the user's terminal, a central server, and a device equipped with AR capabilities.

[0696] First, users create illustrations on their devices using a dedicated application. These illustrations can be drawn directly using digital painting tools, or they can be imported by taking a picture of a drawing on paper.

[0697] Next, the device sends this illustration data to the server. The server analyzes the features of the illustration and extracts its shape and color information. Based on this information, the server uses AI technology to automatically generate a 3D shape. The server then adds pre-prepared animations to the generated 3D model. Finally, it adjusts the illustration's color tone and sharpness to make it visually appealing.

[0698] Furthermore, to enable augmented reality display, the server acquires the user's location information and calculates the optimal placement of 3D shapes for the AR device. The constructed AR content is then sent back to the user's device and finally displayed in the user's real space via the AR device.

[0699] As a concrete example, consider using an illustration of an animal drawn by a child. The user draws an illustration of an animal on their device and sends it to the server. The server analyzes the outline and colors of the received illustration and generates a 3D model of the animal. This model is then given animations such as walking and jumping. The completed AR content appears before the user's eyes through their AR glasses, just as if it were real. This allows the user to observe the animal moving around in their own physical environment.

[0700] This system provides a way for ordinary users to easily create 3D content and enjoy interactive experiences through AR technology.

[0701] The following describes the processing flow.

[0702] Step 1:

[0703] The user draws illustrations on their device. They launch a dedicated app and create illustrations directly using digital painting tools, or they take a picture of a drawing on paper with the device's camera and import it into the app.

[0704] Step 2:

[0705] The terminal sends the completed illustration data to the server. The transmission format is a standard image format such as JPEG or PNG, and data compression is performed to improve communication efficiency.

[0706] Step 3:

[0707] The server receives the transmitted illustration. After receiving it, it analyzes the shape and color of the illustration and extracts feature data to create a 3D shape.

[0708] Step 4:

[0709] The server automatically generates the 3D shape of the illustration using AI technology based on the analysis data. This generation process is carried out by combining existing 3D model libraries and generation algorithms.

[0710] Step 5:

[0711] The server adds animation to the generated 3D shape. The animation uses a pre-configured template and is adjusted to make the model move naturally.

[0712] Step 6:

[0713] The server automatically adjusts the color tone and sharpness and applies beautification processing to improve the visual quality of the illustrations. This enhances their final appearance.

[0714] Step 7:

[0715] The server acquires the user's location information and analyzes the 3D shape at a position and scale suitable for AR display. Based on the location analysis, the placement is adjusted to match the user's real-world environment.

[0716] Step 8:

[0717] The server sends the created AR content data to the user's device.

[0718] Step 9:

[0719] The device displays the received AR content through a device specified by the user, such as AR glasses or a smartphone. This allows the user to enjoy an augmented reality experience in the real world.

[0720] (Example 1)

[0721] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0722] Currently, there are limited ways for users to easily convert their own illustrations into 3D shapes or animated content, and then extend and experience them in real space. Furthermore, there is a lack of features to improve visual effects in a way that is suitable for the real world, making it difficult for users to realize their ideal visual experience.

[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0724] In this invention, the server includes means for receiving an image drawn by a user using a generation device, means for extracting features of the received image using an analysis device and automatically generating a three-dimensional shape, and means for adding movement to the generated three-dimensional shape. This makes it possible for users to easily turn their own illustrations into three-dimensional objects and experience them extended into real space with movement.

[0725] "User" refers to an individual or group that operates a specific electronic device or system.

[0726] "Generating device" refers to a computer system or device used by a user to create and edit visual information.

[0727] "Image" refers to a visual representation created by a user through a generation device.

[0728] "Means of receiving" refers to a device or system that has the function of acquiring information from an external source via data communication.

[0729] "Analyzing features" refers to performing computational processing to identify useful patterns, shapes, colors, and other elements from data.

[0730] "Three-dimensional shape" refers to information that describes the shape of an object that exists in three-dimensional space.

[0731] "Adding action" refers to setting a series of actions, such as movement or deformation, for a three-dimensional shape.

[0732] Augmented reality refers to a technology that integrates and displays digital content within a real-world physical environment.

[0733] The "visual environment" refers to the physical and digital spaces that a user can perceive through their vision.

[0734] "Automatically adjusting color and image clarity" refers to algorithmic processing that appropriately corrects the color tone and resolution of digital images.

[0735] "Improving visual effects" refers to techniques used to enhance the beauty and appearance of displayed content.

[0736] "Analyzing location information" refers to processing and understanding data based on the user's geographical location.

[0737] "Appropriate placement" refers to arranging digital content in a reasonable manner within a physical or virtual space.

[0738] This invention is a system that converts user-drawn images into three-dimensional shapes and animated content, and displays them in a visual environment as augmented reality. It is realized through the cooperation of the user, terminal, and server.

[0739] Users create images using digital painting tools with a dedicated generator. They can also digitize images drawn on paper by taking a picture of it with the device's camera. Once the image is complete, the device sends the data to the server. The data is encrypted during transmission to ensure secure reception.

[0740] The server analyzes the received image data using an analysis device. This process utilizes AI technology to extract image contours and color information, and identify necessary features. Libraries such as OpenCV and TensorFlow are used for image analysis. Based on the analysis results, the server automatically generates a 3D shape. 3D modeling tools such as Blender and Unity are used for 3D shape generation.

[0741] The generated 3D shapes are given motion using motion capture data, making them appear to move. The server also automatically adjusts the colors and image clarity to improve the appearance. Tools such as Adobe Color and Photoshop are used in this process.

[0742] Furthermore, the server analyzes location information and appropriately positions augmented reality content based on the device's physical location. It uses ARKit or ARCore to perform calculations for placing 3D models in real space.

[0743] Ultimately, the device displays augmented reality content in the visual environment via an AR device, based on data received from the server. For example, an image of an animal drawn by the user can be rendered in 3D and displayed in the living room with animations.

[0744] As an example of a prompt, it is possible to input instructions to the server such as, "Convert the illustration of a sea creature drawn by the user into a 3D model and make it swim in the AR space." This allows the user to experience their imagination in three dimensions in the real world.

[0745] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0746] Step 1:

[0747] The user activates the generator and creates an image using digital painting tools. Input includes brush strokes and color selections from the drawing tools. Once the user finishes drawing, they save it as a digital file (e.g., PNG, JPEG). The output is a digital image file.

[0748] Step 2:

[0749] The terminal sends image files created by the user to the server. The input is a digital image file, and the terminal encrypts the data using an appropriate protocol for transmission. The output is sent to the server as encrypted image data.

[0750] Step 3:

[0751] The server analyzes the received image data. The input is encrypted digital image data, which is decrypted, and AI technology (e.g., image recognition models) is used to extract contour and color features. Specifically, image processing is performed using OpenCV or TensorFlow. The output is feature data for generating 3D shapes.

[0752] Step 4:

[0753] The server automatically generates a 3D shape based on extracted feature data. The input is feature data. Using the generated AI model, a 3D model is created using tools such as Blender or Unity. Textures are also applied during this process. The output is a 3D model of the three-dimensional shape.

[0754] Step 5:

[0755] The server adds motion to the generated 3D model. The input is the generated 3D model, to which motion capture data is applied to add movement. The output is the 3D model with motion.

[0756] Step 6:

[0757] The server adjusts color and image clarity to improve visual effects. The input is a 3D model with animations, which is adjusted using Adobe Color and Photoshop. The output is a visually adjusted 3D model.

[0758] Step 7:

[0759] The server analyzes the user's location information and calculates the optimal placement in the augmented reality environment. The inputs are location information and a pre-calibrated 3D model. Placement optimization is performed using ARKit or ARCore. The output is the pre-calibrated 3D model.

[0760] Step 8:

[0761] The terminal dynamically sends 3D model data with pre-calculated placement, received from the server, to the display device. The input is the pre-calculated 3D model. It is displayed overlaid on the real world via an AR device. The output is 3D content existing in an augmented reality environment.

[0762] (Application Example 1)

[0763] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0764] While technologies existed to convert user-generated illustrations into 3D shapes, systems capable of real-time customization to meet user needs were limited. Furthermore, there was a lack of easy ways for users to test their designs in a virtual space, making an intuitive and interactive experience difficult. Therefore, there was a need to realize a consistent process from illustration creation to design confirmation.

[0765] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0766] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a 3D shape, and means for adding animation to the generated 3D shape. This allows the user to instantly try out their designs in a virtual space and interactively check their appearance.

[0767] "Means for receiving user-generated illustrations" refers to the process of receiving digital images created by users from their devices and transmitting them to the system.

[0768] "A method for extracting features from illustrations and automatically generating 3D shapes" refers to a technology that analyzes shape and color information from received image data and generates three-dimensional objects based on that information.

[0769] "Methods for adding animation to 3D shapes" refers to the process of adding movement to generated three-dimensional shapes to realize visually dynamic expressions.

[0770] "Means of displaying as augmented reality" refers to technologies that overlay virtual objects onto real-world environments, providing users with a new visual experience.

[0771] "A method of placing a design in a virtual environment and testing its appearance" refers to a method of placing a user's design in a digital space and simulating its look and impression.

[0772] "Means of providing an interface" refers to a mechanism that supports users in intuitively operating the system through a screen and utilizing the desired functions.

[0773] The system for implementing this invention allows users to convert illustrations they have drawn into 3D shapes in a virtual environment and then test their appearance. Users first create illustrations using digital painting tools on their smartphones or tablets. If necessary, they can also digitize and import illustrations drawn on paper.

[0774] The user's device sends the created illustration to the server. The server receives the illustration and first uses a Python script to perform shape analysis using the OpenCV library. This analysis extracts shape and color information, and based on this data, it automatically generates a 3D shape using the Blender API.

[0775] The generated 3D model is animated using Unity's AR Foundation and ARKit / ARCore, and then displayed on the user's device in an augmented reality environment. During this process, the server obtains location information from the user's device and calculates the optimal placement of the 3D shape. This entire process is designed to provide a smooth interactive experience across the entire system.

[0776] As a concrete example, users can draw illustrations of clothing items based on their own designs, and these illustrations are instantly placed in a virtual environment, allowing for visual confirmation and feedback on the designs. This system enables users to enjoy a real-time customization experience utilizing the digital space.

[0777] An example of a prompt message might be, "Design a stylish T-shirt, and it will be virtually rendered in 3D. Please share your design ideas. For example, a beach-themed resort design." This prompt message stimulates the user's creativity and promotes further design innovation.

[0778] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0779] Step 1:

[0780] The user's device launches a digital painting tool, and the user draws a T-shirt design.

[0781] Input: User-drawn design data

[0782] Output: Digitized illustration file

[0783] Operation: Users use the application's in-app painting tools to add lines and colors directly to the digital canvas and complete their designs.

[0784] Step 2:

[0785] Send an illustration of the terminal to the server.

[0786] Input: Digitized illustration files

[0787] Output: Notification that the transfer of illustration data to the server is complete.

[0788] Operation: Data is sent from the terminal to the server using a secure protocol. Once the transmission is complete, the server notifies the client that it has received the data.

[0789] Step 3:

[0790] The server analyzes the illustration data, extracts features, and generates a 3D shape.

[0791] Input: Received illustration data

[0792] Output: 3D shape data

[0793] Operation: The received data is analyzed using OpenCV in a Python script to extract shapes and patterns. Then, a 3D model based on these features is generated using the Blender API.

[0794] Step 4:

[0795] The server generates a 3D shape and then adds animation to it.

[0796] Input: 3D shape data

[0797] Output: Animated 3D shape data

[0798] Operation: To apply preset animations to shape data, the Unity engine's animation tools are used to add movement and deformation to the generated 3D model.

[0799] Step 5:

[0800] The server calculates the optimal augmented reality placement based on location information and sends the data to the user's device.

[0801] Input: Animated 3D shape data, user position information

[0802] Output: 3D shape data with placement completed on the terminal

[0803] Operation: The server analyzes location information from the terminal, calculates the optimal placement of the model at the user's current location, and then sends the 3D model data to the device.

[0804] Step 6:

[0805] The user's device uses AR functionality to display animated 3D shapes in the real environment.

[0806] Input: Completed 3D shape data

[0807] Output: Display of 3D shapes within the user's field of view.

[0808] Operation: The device utilizes Unity's AR Foundation to display virtual objects as if they were integrated into the real world through the device's camera, providing the user with an interactive experience.

[0809] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0810] The system implementing this invention generates a 3D shape based on an illustration drawn by the user, displays augmented reality content with added animation, and integrates an emotion engine that recognizes the user's emotions. This system is realized through the cooperation of the user's terminal, a central server, and a device with AR capabilities.

[0811] First, the user creates an illustration within the application on their device. This illustration is imported into the device via a digital painting tool or external image input. Next, this illustration is sent to a server, which extracts the necessary features from the received illustration and automatically generates a 3D shape. At the same time, the user's emotions are also analyzed in real time by an emotion engine using camera and sensor data.

[0812] The server adds emotion-appropriate effects and actions to the generated 3D shapes and animations based on the recognized user's emotional information. For example, if the emotion engine detects "joy," the 3D shape may be corrected with brighter colors, or joyful movements may be added to the animation.

[0813] Furthermore, the system automatically applies a beautification process to improve the visual quality of the illustrations. This beautification process also adjusts the color tone and sharpness based on the analysis results of the emotion engine, ensuring that the visuals are appropriate to the user's emotional state. The completed AR content is then optimally positioned based on the user's location information and displayed on the AR device through the terminal.

[0814] For example, if a user uses the system while feeling down, the emotion engine will recognize that emotion, and encouraging messages will appear on the generated 3D shape, or warm colors will be added to the background. This allows users to receive a personalized AR experience that is tailored to their emotions.

[0815] In this way, this system supports the creation of creative and interactive AR content, enabling users to experience things that resonate with their emotions.

[0816] The following describes the processing flow.

[0817] Step 1:

[0818] Users create illustrations using applications on their devices. They select a digital painting tool, choose colors and brush types, and begin designing. They freely express themselves using their fingers or a pen on the screen of their smartphone or tablet.

[0819] Step 2:

[0820] The drawn illustration is sent to the server by the device. During transmission, the illustration data is compressed and converted to a format suitable for smooth communication (e.g., JPEG or PNG).

[0821] Step 3:

[0822] The server receives the illustration. The server uses image analysis technology to extract the color, shape, and texture features contained in the illustration.

[0823] Step 4:

[0824] The server automatically generates a 3D shape based on the extracted features. Using AI technology, the most suitable 3D model is created according to the characteristics of the illustration.

[0825] Step 5:

[0826] The server uses camera and sensor data provided by the user to operate the emotion engine. It analyzes the user's emotions from their facial expressions and voice.

[0827] Step 6:

[0828] Based on the results recognized by the emotion engine, the server adds animations and effects to the 3D shape. For example, if the user appears happy, brighter colors and bouncy movements are added to the 3D model.

[0829] Step 7:

[0830] The server automatically adjusts the color tone and sharpness of the illustrations, performing beautification processing. Hues and brightness are applied according to the emotion, adjusting them to enhance their visual appeal.

[0831] Step 8:

[0832] The server checks the user's location and calculates coordinates to optimally position the 3D shapes in augmented reality. This ensures that the placement is natural and easy to view in the user's real-world environment.

[0833] Step 9:

[0834] The completed AR content is sent from the server to the device. The device then interacts with an AR device (e.g., AR glasses) to display a 3D shape superimposed onto the real world. The user can then observe and experience this.

[0835] (Example 2)

[0836] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0837] This invention aims to realize an augmented reality experience that reflects the user's emotional state by generating three-dimensional data based on a two-dimensional image drawn by the user and by adding actions to the generated data. Conventional technologies have struggled to generate interactive content that takes into account the individual emotions of users, posing challenges in providing a personalized experience. Furthermore, the placement and display of augmented reality content have not been optimized, sometimes detracting from immersion depending on the usage scenario.

[0838] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0839] In this invention, the server includes means for receiving images created by the user, means for extracting characteristics from the received images and automatically generating three-dimensional data, means for analyzing the user's state, and means for adding actions and effects to the generated three-dimensional data according to the analyzed user state. This enables an interactive and personalized augmented reality experience based on the user's emotions.

[0840] "User-generated images" refer to visual information generated or imported by users using digital tools.

[0841] "Extracting characteristics" refers to the process of identifying and analyzing important elements and features from received image data.

[0842] "Three-dimensional data" refers to a dataset that defines three-dimensional shapes generated from two-dimensional information.

[0843] "Adding motion" refers to the process of adding animation or movement attributes to three-dimensional data.

[0844] "Analyzing the user's state" means recognizing and digitizing the user's emotions and actions through sensors and algorithms.

[0845] "Adding effects" means applying saturation, motion, and other visual changes to 3D data and animations based on the results of user sentiment analysis.

[0846] "Displaying as augmented reality" means overlaying digital data onto a real-world environment for a visual presentation.

[0847] "Analyzing location data" means analyzing the user's current location and location information to determine the optimal placement of digital content.

[0848] This invention is a system for realizing an interactive and personalized augmented reality experience using user-generated images. Users create images using digital painting applications. Specifically, general-purpose painting software or image processing applications may be used. This image is drawn on the user's device, and the device sends the image data to the server via secure communication. The server extracts the characteristics of the received image using an image processing library and generates three-dimensional data.

[0849] Image processing and machine learning libraries such as OpenCV and TensorFlow are used in this process. Furthermore, the server acquires data from the terminal's camera and microphone in real time to recognize the user's state and performs analysis using an emotion engine. Commonly available emotion recognition software is used for this analysis. Based on the analysis results, actions and effects corresponding to the user's emotions are added to the generated three-dimensional data.

[0850] Additional effects are reflected, for example, in the color scheme and movement style. Ultimately, the server uses the user's location information to position the generated augmented reality content optimally and displays it on the AR device through the terminal.

[0851] For example, if a user draws an illustration expressing their mood, the system can provide a colorful, dynamic, three-dimensional animation that reflects that emotion. An example of a prompt is as follows: "Generate interactive AR content to cheer up a depressed user. The background should be in calming colors and include an encouraging message." In this way, the present invention provides a creative and emotionally responsive augmented reality experience.

[0852] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0853] Step 1:

[0854] The user draws an image using a digital painting application on their device. The user can freely edit elements such as color and shape. The created image is saved on the device and sent to the server via a communication module. The input is the image drawn by the user, and the output is the digital image data sent to the server.

[0855] Step 2:

[0856] The server extracts features from the received image data. The server uses OpenCV and TensorFlow to perform image analysis, including edge detection and color distribution analysis. This analysis extracts important image features as data. The input is the received digital image data, and the output is a dataset of the extracted features.

[0857] Step 3:

[0858] The server generates three-dimensional data based on the extracted feature data. Three-dimensional modeling software such as Blender or Unity is used to create appropriate vertex information and mesh structures. A shape generation algorithm is applied during this process. The input is a dataset of extracted features, and the output is the generated three-dimensional model data.

[0859] Step 4:

[0860] The device uses a camera and microphone to capture the user's state and sends the data to a server. The input consists of images of the user's facial expressions and audio data, while the output is the dataset sent to the server.

[0861] Step 5:

[0862] The server analyzes the transmitted user state data. Using an emotion recognition engine, it estimates the user's emotions in real time. The input is a dataset of transmitted user state data, and the output is the analyzed emotion information.

[0863] Step 6:

[0864] The server adds actions and effects to the 3D model based on the user's analyzed sentiment information. This includes adjusting color tones and adding animations. The input is the generated 3D model data and analyzed sentiment information, and the output is the updated 3D content.

[0865] Step 7:

[0866] The server appropriately positions the updated 3D content based on the user's location information and sends it to the terminal. The terminal displays the content using an AR device, providing the user with an augmented reality experience. The input is the updated 3D content data and the user's location information, and the output is the augmented reality content displayed on the AR device.

[0867] (Application Example 2)

[0868] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0869] In physical stores, there is a need for methods that allow customers to visually enjoy selecting products based on their own drawings. However, conventional technologies have made it difficult to provide such an interactive and personalized experience in real time. Furthermore, dynamic adjustments of shapes and colors in response to customer emotions have not been considered, making it impossible to provide a deeper user experience.

[0870] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0871] In this invention, the server includes means for receiving illustrations drawn by the user, means for extracting features from the received illustrations and automatically generating a three-dimensional shape, means for adding animation to the generated three-dimensional shape, and means for recognizing the user's emotions and adjusting the three-dimensional shape and animation according to those emotions. This allows customers to experience product selection using illustrations they have drawn themselves in a physical store using augmented reality, and further enjoy personalized visual effects that respond to their emotions.

[0872] A "user" is someone who uses the system to create illustrations or to visually experience augmented reality content.

[0873] An "illustration" refers to a two-dimensional image or drawing created or input by a user.

[0874] "Features" refer to important patterns and shapes extracted from the received illustration, and are essential elements for generating three-dimensional shapes.

[0875] A "3D shape" is a three-dimensional digital object that is automatically generated based on the characteristics of an illustration.

[0876] "Animation" refers to visual effects that involve movement or change added to three-dimensional shapes.

[0877] "Emotions" refer to information that indicates the user's psychological state, and the system uses this information to provide a personalized experience by recognizing it.

[0878] Augmented reality is a visual experience that overlays digital information onto images of the real world.

[0879] This system is realized through the cooperation of the user, terminal, and server. First, the user creates an illustration using a terminal with a dedicated application installed. Digital painting tools and image input functions can be used to create the illustration. This illustration is then sent from the terminal to the server.

[0880] The server extracts key features from received illustrations and automatically generates 3D shapes based on them. This generation process utilizes machine learning libraries such as TensorFlow. Furthermore, the server is equipped with an emotion engine that analyzes the user's emotions in real time via cameras and sensors. Based on this emotion information, personalized effects are added to the 3D shapes and animations.

[0881] On the user's device, the completed 3D shape is displayed as augmented reality. The displayed AR content reflects the analysis results of an emotion recognition module linked with OpenCV, and is accompanied by visual effects that match the user's emotions. Once the content is displayed, the user can experience a real-time, changing product display in a physical store through devices such as smart glasses.

[0882] For example, if a user draws an illustration of their pet at a pet supply store, a 3D model of the pet is created based on that sketch and displayed on smart glasses. If the user shows a joyful expression upon seeing the pet, the system recognizes that emotion and adds an animation that makes the pet appear to move.

[0883] Examples of prompts to input into a generative AI model are as follows:

[0884] "At a pet supply store, users are drawing illustrations of pets. Please generate 3D models based on these sketches and add animations that express the pet's excitement when the user smiles."

[0885] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0886] Step 1:

[0887] The user creates an illustration using their device. They draw the illustration using the device's digital painting tool and save it as a digital image. The input is the illustration drawn by the user on the screen, and the output is a digital image file.

[0888] Step 2:

[0889] The terminal sends the created illustration to the server. The input is a digital image file, which is transferred to the server's receiving port. The output is image data stored on the server.

[0890] Step 3:

[0891] The server extracts features from received illustrations and generates 3D shapes. Image processing algorithms are used to extract important shape data, and 3D modeling is performed based on this data. The input is received image data, and the output is 3D shape data.

[0892] Step 4:

[0893] The server analyzes the user's emotions using cameras and sensors. The emotion analysis engine processes the image and sensor data to estimate the emotional state from the user's facial expressions and movements. The input is the user's facial expressions and movement data, and the output is the estimated emotion information.

[0894] Step 5:

[0895] The server adds animation and effects to 3D shapes based on emotional information. It utilizes a generative AI model to apply emotion-dependent actions and color changes to the model. Input is 3D shape data and emotional information, and output is animated 3D shape data.

[0896] Step 6:

[0897] The device displays animated 3D shapes as augmented reality. The data is transferred to an AR device and overlaid onto a display in a physical store, integrating it into the user's field of view. The input is animated 3D shape data, and the output is augmented reality imagery within the user's field of view.

[0898] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0899] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0900] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0901] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0902] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0903] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0904] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0905] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0906] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0907] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0908] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0909] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0910] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0911] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0912] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0913] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0914] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0915] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0916] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0917] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0918] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0919] The following is further disclosed regarding the embodiments described above.

[0920] (Claim 1)

[0921] A means of receiving illustrations drawn by users,

[0922] A means for extracting features from the received illustration and automatically generating a three-dimensional shape,

[0923] Means for adding animation to the generated three-dimensional shape,

[0924] A means for displaying the aforementioned animated 3D shape as augmented reality,

[0925] A system that includes this.

[0926] (Claim 2)

[0927] The system according to claim 1, comprising means for automatically adjusting and beautifying the color tone and sharpness of the aforementioned illustration.

[0928] (Claim 3)

[0929] The system according to claim 1, further comprising means for analyzing location information and appropriately arranging the three-dimensional shape of the augmented reality to be displayed.

[0930] "Example 1"

[0931] (Claim 1)

[0932] Means for receiving an image drawn by a user using a generator,

[0933] A means for extracting features from the received image using an analysis device and automatically generating a three-dimensional shape,

[0934] Means for imparting movement to the generated three-dimensional shape,

[0935] Means for displaying the aforementioned three-dimensional shape with movement as augmented reality in a visual environment,

[0936] A system that includes this.

[0937] (Claim 2)

[0938] The system according to claim 1, comprising means for automatically adjusting the color and image clarity of the aforementioned image to improve the visual effect.

[0939] (Claim 3)

[0940] The system according to claim 1, further comprising means for analyzing location information and appropriately arranging the augmented reality three-dimensional shape to be displayed.

[0941] "Application Example 1"

[0942] (Claim 1)

[0943] A means of receiving illustrations drawn by users,

[0944] A means for extracting features from the received illustration and automatically generating a three-dimensional shape,

[0945] Means for adding animation to the generated three-dimensional shape,

[0946] A means for displaying the aforementioned animated 3D shape as augmented reality,

[0947] A means of placing user-designed shapes in a virtual environment and testing their appearance,

[0948] A means to provide an interface that allows users to try out their designs in a virtual space,

[0949] A system that includes this.

[0950] (Claim 2)

[0951] The system according to claim 1, comprising means for automatically adjusting and beautifying the color tone and sharpness of the aforementioned illustration.

[0952] (Claim 3)

[0953] The system according to claim 1, further comprising means for analyzing location information and appropriately arranging the three-dimensional shape of the augmented reality to be displayed.

[0954] "Example 2 of combining an emotion engine"

[0955] (Claim 1)

[0956] A means of receiving images created by the user,

[0957] A means for extracting characteristics from the received image and automatically generating three-dimensional data,

[0958] Means for adding motion to the generated three-dimensional data,

[0959] Means for analyzing the user's state,

[0960] Means for applying effects to the three-dimensional data with motion according to the analyzed state of the user,

[0961] A means for displaying the three-dimensional data with the aforementioned effects applied as augmented reality,

[0962] A system that includes this.

[0963] (Claim 2)

[0964] The system according to claim 1, comprising means for automatically adjusting and improving the color tone and sharpness of the aforementioned image.

[0965] (Claim 3)

[0966] The system according to claim 1, further comprising means for analyzing positional data and appropriately arranging the augmented reality three-dimensional data to be displayed.

[0967] "Application example 2 when combining with an emotional engine"

[0968] (Claim 1)

[0969] A means of receiving illustrations drawn by users,

[0970] A means for extracting features from the received illustration and automatically generating a three-dimensional shape,

[0971] Means for adding animation to the generated three-dimensional shape,

[0972] A means for recognizing the user's emotions and adjusting the 3D shape and animation according to those emotions,

[0973] A means for displaying the aforementioned animated 3D shape as augmented reality,

[0974] A system that includes this.

[0975] (Claim 2)

[0976] The system according to claim 1, comprising means for automatically adjusting and beautifying the color tone and sharpness of the aforementioned illustration.

[0977] (Claim 3)

[0978] The system according to claim 1, further comprising means for analyzing location information and appropriately arranging the three-dimensional shape of the augmented reality to be displayed. [Explanation of symbols]

[0979] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving illustrations drawn by users, A means for extracting features from the received illustration and automatically generating a three-dimensional shape, Means for adding animation to the generated three-dimensional shape, A means for displaying the aforementioned animated 3D shape as augmented reality, A system that includes this.

2. The system according to claim 1, comprising means for automatically adjusting and beautifying the color tone and sharpness of the aforementioned illustration.

3. The system according to claim 1, further comprising means for analyzing location information and appropriately arranging the three-dimensional shape of the augmented reality to be displayed.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A