system

A system using AI and AR to suggest and visualize furniture based on user preferences and spatial information addresses the inefficiencies in furniture selection, ensuring a good fit and reducing the time and labor required for decision-making.

JP2026071013APending Publication Date: 2026-04-28SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-16
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Selecting furniture is time-consuming and laborious, and there is a risk that the chosen furniture may not fit the space, necessitating a more efficient and effective method for choosing optimal furniture based on user preferences and spatial information.

Method used

A system that collects user information, uses AI to suggest optimal items, employs augmented reality (AR) to virtually place items within the user's space, and integrates item evaluation information for comprehensive decision-making.

Benefits of technology

Enables users to efficiently select furniture that fits their space by providing visual confirmation and integrated information for informed decisions, saving time and effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071013000001_ABST
    Figure 2026071013000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for collecting user information, A means of recommending the most suitable items based on the user's preferences and spatial information, A means for virtually arranging and displaying items, A means of integrating and presenting evaluation information for goods, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The selection of furniture requires choosing the optimal one from many options. Particularly for an individual who wants to arrange a new living environment, there is a problem of a large burden of time and labor. Also, a problem may occur that the furniture does not fit the space after actual purchase. There is a need for a method to solve these problems and select furniture more efficiently and effectively.

Means for Solving the Problems

[0005] This invention provides a means for collecting user information and using generating AI to suggest optimal items based on the user's preferences and spatial information. It also employs augmented reality (AR) technology to virtually place the suggested items within the user's space, allowing for visual confirmation of the results. Furthermore, it solves the above-mentioned problems by integrating item evaluation information and constructing a system that comprehensively presents information useful for user decision-making.

[0006] "User" refers to the person who selects the items suggested using this system.

[0007] "Object information" refers to data related to the size, design, and other characteristics of the user's space.

[0008] "Preferences" refer to the design, color, and style tendencies that users hold based on their past actions and choices.

[0009] "Spatial information" refers to information about the physical structure and layout of the user's room or living environment.

[0010] "Items" refers to products such as furniture and interior furnishings that are proposed by this system.

[0011] "Recommendation" refers to the act of suggesting items that are suitable for the user based on collected data.

[0012] "Virtual placement" refers to using AR technology to visually position objects digitally within a real-world space.

[0013] "Display" refers to the act of visually representing information on a device screen.

[0014] "Evaluation information" refers to information about the quality and value of an item, including past user feedback, market reviews, and pricing information.

[0015] "Integration" refers to unifying data obtained from different information sources and presenting it in an easy-to-understand form for users.

[0016] "Presentation" refers to the act of providing information to users and enabling them to use it as material for selection.

Brief Description of Drawings

[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Modes for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a method for realizing a system that enables users to select furniture efficiently and effectively. The system functions through the interaction of terminals, servers, and users.

[0039] First, the device provides the user with the ability to take photos of their room and upload them to the application. At this time, the user can grant the system access to their social media accounts. The device then sends this data to the server.

[0040] The server analyzes the received image data and uses an application engine to recognize the physical layout, color scheme, and existing interior style of the room. It also analyzes data from social media to identify the user's preferences and lifestyle. This process collects user attribute information.

[0041] Next, the server uses a generative AI model to search a global furniture database based on the collected attribute and room information, and selects suitable item candidates. This AI model has the ability to recommend highly relevant furniture based on past data and current user attributes.

[0042] Subsequently, the candidate furniture generated by the server is sent to the terminal. The terminal uses AR technology to provide a function that virtually places the suggested furniture in the user's room. Through the terminal, the user can visually check the virtually placed furniture and evaluate its harmony and fit within the space.

[0043] Furthermore, the server collects and integrates sales information, user reviews, and pricing information for each piece of furniture. This information is presented visually to the user via their terminal to assist in making decisions when choosing furniture.

[0044] As a concrete example, suppose a user is looking for modern-style furniture to decorate their new home. The user takes photos of the room with their device and allows the use of social media data. Based on this information, the server recommends several sofas and tables and provides a visualization of how they would look in the room using augmented reality (AR). The user then selects and purchases the best furniture, taking into account the latest sales information. This allows the user to achieve their ideal interior design while saving time and effort.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The user takes photos of the room using their device. The photos are uploaded to the server through the application. The user can also grant the application access to their social media accounts.

[0048] Step 2:

[0049] The server processes the received photos of the room using image analysis technology to identify the room's size, color scheme, and existing interior. This extracts the room's physical characteristics as digital data.

[0050] Step 3:

[0051] The server analyzes data collected from social media to identify user preferences and lifestyles. This analysis provides attribute information such as the user's preferred style, colors, and themes.

[0052] Step 4:

[0053] The server utilizes a generative AI model to select the most suitable item candidates from a global furniture database based on collected room characteristics and user attribute information. The AI ​​model extracts furniture that meets the selection criteria.

[0054] Step 5:

[0055] The server sends information about the selected furniture candidates to the terminal. The terminal uses this information and augmented reality (AR) functionality to virtually place the furniture in the room. This allows the user to visually see how the furniture will be arranged.

[0056] Step 6:

[0057] Users operate a device to view virtually placed furniture from various angles and judge its harmony with the overall room and its suitability for the space.

[0058] Step 7:

[0059] The server collects and integrates sales information and customer reviews for each piece of furniture from online data. It then provides users with information that summarizes the advantages and disadvantages of each option.

[0060] Step 8:

[0061] Users select the most suitable furniture based on the information provided through their device and proceed with the purchase. This allows users to make decisions quickly.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Traditionally, selecting furniture and interior items has presented challenges in making optimal choices that suit individual user preferences and spaces. Furthermore, a lack of visual support for making appropriate selections without physically inspecting the items meant that the selection process was time-consuming and laborious. Additionally, centrally collecting and presenting relevant pricing and user review information was difficult.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for collecting and analyzing user image information, means for recommending optimal items using a generating AI model based on the analyzed spatial structure and color tone, means for virtually arranging and visualizing the recommended items in space, and means for collecting, integrating, and displaying sales information and evaluation information related to the items. This enables the selection of items optimized for the user, allows for visual confirmation of their arrangement, and enables efficient decision-making by comprehensively viewing related information.

[0067] "Means for collecting and analyzing user image information" refers to a system that acquires image data provided by users and processes that data to identify the layout, color tone, and style of a space.

[0068] "A method for recommending optimal items using a generative AI model" refers to a system that utilizes AI technology to present highly relevant furniture and interior items based on collected user attribute information and spatial characteristics.

[0069] "A means of virtually placing and visualizing recommended items in a space" refers to a system that uses AR technology or similar methods to simulate and display selected furniture and interior items in the user's room in real time, providing an image of how they would actually look in place.

[0070] "Means for collecting, integrating, and displaying sales information and evaluation information related to goods" refers to a system that collects and organizes market price information, user reviews, and sales information related to presented furniture and interior items, and presents them comprehensively to users.

[0071] This invention provides a system for users to efficiently select furniture and interior items. This system operates through the interaction of a terminal, a server, and the user.

[0072] The device primarily provides the functionality for users to take photos of their rooms and upload them to the application. Users can also grant the system permission to access their social media account data. The captured images are sent from the device to the server, and the social media data is similarly transferred to the server.

[0073] The server uses an image analysis engine to analyze received image data and recognize the physical layout, color scheme, and interior style of the room. For social media data, attribute information is collected by identifying user preferences and lifestyles using natural language processing techniques. In this process, for example, the TENSORFLOW® library is used for image analysis.

[0074] Next, the server uses a generative AI model to select highly relevant furniture from a global database based on the collected attribute and room information. The generative AI model uses machine learning algorithms to recommend furniture based on past data and current user attributes. At this point, the server sends the generative AI a prompt message stating, "Please suggest the most suitable furniture based on user attribute data."

[0075] Subsequently, the furniture selected by the server is sent to the terminal, which uses augmented reality (AR) technology to virtually place the furniture in the user's room. This process allows the user to visually confirm how the furniture will look. Furthermore, the server collects sales information, user reviews, and pricing information and presents it to the user as integrated information through the terminal. This enables the user to make efficient decisions.

[0076] For example, if a user wants to furnish their new home with modern-style furniture, they first take a photo of the room with their device and allow the use of their social media data. The server analyzes this information, recommends suitable sofas and tables, and provides a visualization of how they would look in the room using augmented reality (AR) technology. As a result, the user can effectively and efficiently select the interior items they desire.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The user takes a picture of their room using their device and uploads it to the application. The user can also grant permission for the application to use data from their social media account. The input consists of the captured image of the room and social media authentication information, while the output is a data file that can be sent. Specifically, the user activates the device's camera function and presses the capture button to acquire the image.

[0080] Step 2:

[0081] The device sends captured image data and data from social media to the server. The input is the data file created earlier, and the output is the server's reception status. The device sends the data using the secure HTTPS protocol and displays a progress bar to notify the user when the transmission is complete.

[0082] Step 3:

[0083] The server analyzes the received image data and retrieves the results. The input is image data received from the terminal, which is processed by the analysis engine, and the output is information about the room layout and color tone. Specifically, the server uses an image analysis engine (e.g., TensorFlow) to identify the color tone and layout of the image and stores it in JSON format.

[0084] Step 4:

[0085] The server analyzes SNS data to identify user preferences and lifestyles. The input is SNS profile information, and natural language processing techniques are used for analysis. The output is user attribute information. Specifically, the server uses a natural language processing library to analyze text data, performing keyword extraction and topic modeling.

[0086] Step 5:

[0087] The server uses a generative AI model to recommend furniture. The input consists of analyzed user attribute information and room information. The AI ​​model searches for the most suitable furniture from the big data, and the output is a list of recommended furniture. Specifically, the server sends a prompt to the AI ​​model, "Please suggest the most suitable furniture based on the user attribute data," and receives a response from the model.

[0088] Step 6:

[0089] A list of furniture selected from the server is sent to the terminal. The terminal then uses AR technology to virtually place these furniture items in the user's room. The input is furniture data from the server, and the output is an AR display visually presented to the user. Specifically, the furniture is displayed as virtual objects on the terminal's screen and can be viewed in real time by overlaying it with the camera image.

[0090] Step 7:

[0091] The server collects and integrates sales information, reviews, and pricing information related to the proposed furniture. The input is the furniture's ID information, and the output is presented to the user in a centralized information pack. Specifically, the server uses external APIs to retrieve the necessary data, organizes it, and sends it to the terminal.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In today's retail environment, consumers are required to choose the optimal product from a vast selection, which is often a time-consuming and laborious task. Furthermore, even if consumers visually examine products in a physical store, it is difficult to confirm their suitability in their actual living space beforehand. This leads to problems such as increased dissatisfaction and the risk of returns after purchase.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for utilizing visual equipment to extract user requests, means for recommending optimal products based on the user's preferences and location structure information, and means for displaying and arranging products three-dimensionally using virtual reality technology. This allows consumers to check the suitability of products in their actual living spaces in advance and make purchase decisions efficiently.

[0097] "User requirements" refer to the specific needs and criteria that users have when choosing a product.

[0098] "Visual devices" refer to augmented reality devices and wearable devices that users use to virtually inspect products.

[0099] "Location structure information" refers to data about the physical layout and design of a user's residence or commercial space.

[0100] The "optimal product" is an item that has been determined to be the most suitable based on the user's requirements and preferences.

[0101] "Virtual reality technology" refers to the technology that uses computer technology to construct a virtual world, allowing users to visually experience objects within that world.

[0102] The embodiments for carrying out the invention will now be described. To realize this invention, the system is configured as follows.

[0103] The terminal will primarily utilize smart glasses or headset devices as visual devices worn by the user. The visual devices will capture images of in-store products and transmit that data to the server. Specifically, the hardware used will be smart glasses (e.g., Microsoft® HoloLens®, Google® Glass®).

[0104] The server receives data transmitted from visual devices and analyzes the structural information of the location using an image analysis library (e.g., OpenCV). It also analyzes past purchase history and information collected from social media to identify user attributes and extract user preferences and requests.

[0105] Furthermore, the server uses a generative AI model to recommend the most suitable products from the product database based on the analyzed data. In this process, prompts are used to instruct the AI ​​model. For example, a prompt might say, "Suggest the most modern style furniture for this living room. Please select based on this image and the user's preferences."

[0106] The terminal uses information received from the server and virtual reality technology to provide a visualization of how the product would look if it were actually placed in the user's living space. This allows the user to see in real time how the product would look in their own space.

[0107] As a concrete example, when a user visiting a furniture store uses smart glasses to view a sofa, they can see a virtual view of the sofa placed in their own living room. In this way, users can confirm the suitability of their product selection before purchasing it.

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The user wears a visual device and focuses their gaze on an item of interest in the store. The visual device (smart glasses) captures the image and sends it to the server along with the user's gaze data. At this point, the input is the image data of the product and the user's gaze data, and the output is the transmission of data to the server.

[0111] Step 2:

[0112] The server analyzes the received video data using an image analysis library (e.g., OpenCV) to identify structural information of the location. This analysis extracts positional information and color tones related to the placement of products. The input is the captured video data, and the output is the spatial information obtained through the analysis.

[0113] Step 3:

[0114] The server analyzes the user's past purchase history and data from social media to identify user preferences and requests. This process uses data mining techniques to extract user attributes. The input is data about user attributes, and the output is the extracted preferences and purchasing trends.

[0115] Step 4:

[0116] The server utilizes a generative AI model to search and recommend the most suitable products from its product database based on analyzed spatial information and user preferences. Advanced recommendations are provided by giving instructions to the AI ​​using prompts. Input consists of spatial information and user preferences, while output is a list of recommended products.

[0117] Step 5:

[0118] The terminal uses recommended product information received from the server to visualize the product placement using virtual reality technology. It presents a virtual product placement to the user, making it appear as if the products are actually present in that space. The input is recommended product information, and the output is a virtual view with the products placed.

[0119] Step 6:

[0120] Users review the visualized product placement and evaluate how well the product fits into their living space. Based on this evaluation, users determine their purchase intent and make a selection. The input is visual information of the virtual product placement, and the output is the user's evaluation and decision.

[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0122] This invention provides a method for realizing a system that recognizes a user's emotional state using an emotion engine and personalizes furniture selection and presentation based on that information. This system functions through interaction between a terminal, a server, and an emotion engine.

[0123] Users can first take photos of their room using their device and upload that data to the server via the application. Furthermore, users grant access to a feature that allows the emotion engine to analyze facial expressions and voice data.

[0124] The server analyzes the physical structure and color scheme from received images of the room, and analyzes social media data to identify the user's preferences and lifestyle patterns. In addition, the emotion engine analyzes multimodal data to capture emotions, such as the user's facial expressions, voice tone, and entered text. This analysis allows the user to obtain their current emotional state and past emotional history.

[0125] The server takes data obtained from the emotion engine into account and uses a generative AI model to recommend furniture that suits the user's preferences and emotional state. For example, if the analysis determines that the user is in an emotional state of wanting to relax, it can suggest sofas in calming colors and furniture with relaxing designs.

[0126] Next, the furniture candidates selected by the server are sent to the terminal. The terminal virtually places these furniture pieces using AR technology, providing the user with a visual simulation. Through this visualization, the user can see how the selected furniture harmonizes with the overall space.

[0127] Furthermore, the server aggregates sales information and customer reviews for each piece of furniture. This information is presented to the user via their device and used as a criterion for selection. Users can then select and purchase the most suitable furniture, taking into account their own emotional state and the actual condition of their room.

[0128] For example, if a user is feeling stressed and wants to add a calming element to their living space, the emotion engine will pick up on this information and recommend relaxing interior design. In this way, users can quickly create a comfortable living space that corresponds to their emotions.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] The user takes photos of the room using their device and uploads the data to the server through the application. The user also grants access to the emotion engine and configures settings to allow input of facial expressions and voice.

[0132] Step 2:

[0133] The server analyzes the received images of the room to identify its size, color scheme, furniture arrangement, and other characteristics. This image analysis generates digital data representing the physical features of the room.

[0134] Step 3:

[0135] The server analyzes the user's social media data to obtain attribute information about the user's preferences. This includes the user's preferred design style and color palette.

[0136] Step 4:

[0137] The emotion engine analyzes the user's facial expressions and voice to understand their emotional state in real time. Furthermore, by considering their past emotional history, it provides data to make more appropriate recommendations.

[0138] Step 5:

[0139] Based on data from the emotion engine, the server utilizes a generative AI model to recommend the most suitable furniture for the user's preferences and emotional state. For example, if the emotion of wanting to relax is recognized, furniture with a calming design will be selected.

[0140] Step 6:

[0141] The furniture candidates selected by the server are sent to the terminal. The terminal uses AR technology to virtually place the furniture in the user's room, providing a realistic visual simulation.

[0142] Step 7:

[0143] Users operate a device to view virtually placed furniture. They evaluate the harmony of the furniture with the space and its placement while completing a visual challenge.

[0144] Step 8:

[0145] The server collects and organizes sales information and reviews for each furniture item. This information is then presented to the user via their terminal for final decision-making.

[0146] Step 9:

[0147] Users select the most suitable furniture based on the information presented through their device and complete the purchase process. As a result, users can choose the optimal interior design that suits their emotions.

[0148] (Example 2)

[0149] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0150] To improve the comfort of modern living spaces, personalized interior design suggestions tailored to the user's emotional state are crucial. However, conventional systems have struggled to make recommendations that take emotional states into account, failing to adequately meet individual user needs. Therefore, there is a need for a system that accurately recognizes the user's emotional state and recommends the most suitable items based on that understanding.

[0151] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0152] In this invention, the server includes means for collecting data to recognize the user's emotional state, means for recommending the most suitable items based on the emotional state and spatial information, and means for arranging the items in a virtual space and displaying them using augmented reality technology. This makes it possible to suggest personalized interiors that correspond to the user's emotional state.

[0153] "Users" refer to individuals who use the system to receive interior design suggestions based on their emotional state.

[0154] "Emotional state" refers to the psychological condition obtained by analyzing multimodal data such as the user's facial expressions, voice tone, and entered text.

[0155] "Means of data collection" refers to the technical processes and devices used to obtain information such as images, audio, and text from users and transmit it to a server.

[0156] "Spatial information" refers to information that includes physical characteristics such as structure and color scheme related to the user's living space.

[0157] "Means of recommending optimal items" refers to algorithms and technologies that propose interiors deemed optimal for the user based on analyzed emotional states and spatial information.

[0158] "Means of placing objects in a virtual space and displaying them using augmented reality technology" refers to a technology that overlays selected objects as digital data onto the user's actual space and provides a visual simulation.

[0159] "Means of integrating and presenting evaluation information" refers to technologies and processes for aggregating reviews and evaluations of selected items and presenting them to users in an easy-to-understand manner.

[0160] A "generative AI model" refers to an artificial intelligence algorithm used to analyze a user's emotional state and lifestyle patterns and provide individually optimized recommendations.

[0161] A "prompt" refers to the text input to give specific instructions to a generative AI model.

[0162] The following describes the specific operations and techniques used in embodiments for carrying out the present invention.

[0163] This system consists of a user terminal, a server that processes data, and an emotion engine that analyzes emotions.

[0164] The user first takes a picture of the room using their device. This picture is sent to the server via the device. The user also consents to the collection of data on their facial expressions and voice tone, allowing the emotion engine to analyze it.

[0165] The server analyzes the received image data using image recognition software to extract spatial information such as the physical layout and color scheme of the room. The server also collects data from social media and other sources to identify the user's lifestyle patterns and preferences.

[0166] The emotion engine uses multimodal data (facial expressions, voice, text) obtained from the user to analyze the user's current emotional state in detail. This analysis also takes into account the user's usual emotional history.

[0167] Based on all this information, the server utilizes a generated AI model to select the most suitable interior for the user. For example, if it determines that the user is seeking relaxation, it will suggest furniture in calming colors.

[0168] The device uses augmented reality (AR) technology to virtually place furniture information received from the server into the user's room. This allows the user to intuitively see how the furniture would actually look in their room.

[0169] Furthermore, the server aggregates sales information and customer reviews related to the selected furniture and presents them to the user via the terminal. This allows users to make informed decisions when purchasing interior items that resonate with their emotions.

[0170] For example, a prompt such as "Suggest furniture that will help the user relax" is input into the AI ​​model in an understandable format. This allows for personalized recommendations.

[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0172] Step 1:

[0173] The user takes a picture of the room using their device. The user's action activates the device's camera application, capturing an image of the room. This image data is saved to the device's memory and prepared for transmission.

[0174] Step 2:

[0175] The user uploads image data from their device to the server. The device sends the collected image data to the server via a secure communication protocol. The input is the captured image data, and the output is the data stored on the server.

[0176] Step 3:

[0177] The server receives and analyzes image data. Using an image recognition algorithm, it extracts spatial information such as physical structure and color tone. The input to this process is the received image data, and the output is structural and color tone information. Specifically, it identifies the location of walls and furniture through pixel analysis.

[0178] Step 4:

[0179] The user consents to the collection of emotional data through their device. With permission, the device transmits the user's facial expressions and voice data to the emotion engine. This input is real-time, multimodal data, and the output is data ready for emotion analysis.

[0180] Step 5:

[0181] The emotion engine analyzes the received data. Through facial expression analysis, voice tone analysis, and text analysis, it identifies the user's emotional state. This input is multimodal data, and the output is the user's emotional state and history. Specific operations include voice frequency analysis and facial expression change detection.

[0182] Step 6:

[0183] The server uses a generated AI model to recommend interior design suitable for the user. The input is analyzed emotional state and spatial information, and the output is a list of furniture suitable for the user. Specifically, the AI ​​combines past preference data and emotional data to select the optimal items.

[0184] Step 7:

[0185] The server sends recommended furniture information to the device. The device then uses AR technology to virtually place the received information in the room, providing the user with visual feedback. The input is furniture information, and the output is a visualization using augmented reality technology.

[0186] Step 8:

[0187] The server aggregates and displays sales information and customer reviews related to furniture. The input is a database of information on each piece of furniture, and the output is integrated evaluation information. Specifically, it uses recursive data queries to highlight highly-rated items.

[0188] Step 9:

[0189] The user selects and purchases the most suitable furniture. The user makes their purchase decision considering AR visualizations and evaluation information. The input is the presented interior information, and the output is a notification that the purchase process is complete. Specific actions include one-click purchase options and payment processing.

[0190] (Application Example 2)

[0191] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0192] Modern consumers tend to seek products quickly and accurately that suit their emotional state and individual circumstances when making purchasing decisions. However, many conventional systems have struggled to accurately grasp a user's emotional state and provide products optimized for it. Therefore, the challenge lies in realizing personalized product recommendations that respond to users' emotions and environmental information.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0194] In this invention, the server includes means for analyzing the user's emotional state, means for analyzing the user's physical environment from images, and means for recommending items based on the user's emotional state and environmental information. This makes it possible to provide personalized items based on the user's real-time emotions and environment.

[0195] "Methods for analyzing a user's emotional state" refers to technologies that identify a user's current emotional state based on data obtained from the user's facial expressions, voice, text, etc.

[0196] "Methods for analyzing a user's physical environment from images" refers to technologies that process images of a space provided by the user and extract physical elements such as the room's structure, design, and colors.

[0197] "Means for recommending products based on the user's emotional state and environmental information" refers to a technology that uses analyzed emotional state and environmental data to select and present products suitable for the user.

[0198] "Means of virtually placing the item using augmented reality technology" refers to a method of using augmented reality technology to virtually visualize an item in the user's actual environment.

[0199] "Means for integrating and presenting product evaluation information" refers to a technology that collects third-party evaluations and related information about recommended products and presents them in an easy-to-understand manner for users.

[0200] To implement this invention, the user must first use a device such as a smartphone or head-mounted display to take pictures of their room and their own face. The device has the function of uploading this data to a server.

[0201] The server processes data using an emotion analysis engine to analyze the user's emotional state from facial expressions and voice tone. Open-source software OpenCV is used for facial recognition, and a proprietary voice analysis tool is applied for voice analysis. Furthermore, the server receives images of the room and analyzes the physical structure and color tones of the environment using OpenCV and other tools.

[0202] Next, the server uses a generative AI model to select the most suitable items based on the acquired emotional and environmental data. During this process, it generates specific prompt statements, and the AI ​​model recommends items based on these prompt statements. An example of such a prompt statement might be, "A prompt to suggest the most suitable furniture and interior design for a user experiencing stress."

[0203] The selected items are sent to the device and virtually placed in the room using augmented reality technology. This process utilizes the device's AR capabilities, allowing the user to visually see how the items blend in with the real-world space.

[0204] Furthermore, the server aggregates third-party evaluation information and sales information on items and provides it to the terminal. This enables users to make more informed decisions. Thus, the present invention provides a system that enables personalized item selection based on the user's emotions and environment.

[0205] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0206] Step 1:

[0207] Users use smartphones or head-mounted displays to capture images of their room and their own faces. This provides image data that captures the user's environment and emotions. This image data serves as foundational data for subsequent analysis.

[0208] Step 2:

[0209] The device uploads the acquired image data to the server. The input consists of images of the room and the user's face, and the output is the server receiving the data. The uploaded images are transferred to the server for analysis.

[0210] Step 3:

[0211] The server provides uploaded facial images to an emotion analysis engine for facial recognition and voice analysis. Using facial images and voice data as input, it obtains emotional state data as output. This emotional data is used to select items suitable for the user using a generative AI model.

[0212] Step 4:

[0213] The server analyzes images of a room to identify the structure and color tone of the physical environment. The input is image data of the room, and the output is spatial information and color tone data of the room. Image processing using OpenCV is used to extract physical features as digital data.

[0214] Step 5:

[0215] The server provides prompt text to a generating AI model based on emotional and environmental data, recommending the most suitable items for the user. The input includes emotional state and spatial information data, and the output is item recommendation results. It generates specific prompts such as "a prompt to suggest the most suitable furniture and interior for a user experiencing stress," and the AI ​​model selects the appropriate items.

[0216] Step 6:

[0217] The selected item information is sent to the terminal and virtually placed in the room using augmented reality (AR) functionality. The input is the item recommendation result, and the output is a visual simulation using AR. Through the terminal, the user can see in real time how the items will blend in with the room.

[0218] Step 7:

[0219] The server collects third-party reviews and sales information for recommended items and transmits it to the terminal. The input is item data, and the output is integrated reviews and sales information. This allows users to obtain data to support their decision-making regarding item selection.

[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0223] [Second Embodiment]

[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0236] This invention provides a method for realizing a system that enables users to select furniture efficiently and effectively. The system functions through the interaction of terminals, servers, and users.

[0237] First, the device provides the user with the ability to take photos of their room and upload them to the application. At this time, the user can grant the system access to their social media accounts. The device then sends this data to the server.

[0238] The server analyzes the received image data and uses an application engine to recognize the physical layout, color scheme, and existing interior style of the room. It also analyzes data from social media to identify the user's preferences and lifestyle. This process collects user attribute information.

[0239] Next, the server uses a generative AI model to search a global furniture database based on the collected attribute and room information, and selects suitable item candidates. This AI model has the ability to recommend highly relevant furniture based on past data and current user attributes.

[0240] Subsequently, the candidate furniture generated by the server is sent to the terminal. The terminal uses AR technology to provide a function that virtually places the suggested furniture in the user's room. Through the terminal, the user can visually check the virtually placed furniture and evaluate its harmony and fit within the space.

[0241] Furthermore, the server collects and integrates sales information, user reviews, and pricing information for each piece of furniture. This information is presented visually to the user via their terminal to assist in making decisions when choosing furniture.

[0242] As a concrete example, suppose a user is looking for modern-style furniture to decorate their new home. The user takes photos of the room with their device and allows the use of social media data. Based on this information, the server recommends several sofas and tables and provides a visualization of how they would look in the room using augmented reality (AR). The user then selects and purchases the best furniture, taking into account the latest sales information. This allows the user to achieve their ideal interior design while saving time and effort.

[0243] The following describes the processing flow.

[0244] Step 1:

[0245] The user takes photos of the room using their device. The photos are uploaded to the server through the application. The user can also grant the application access to their social media accounts.

[0246] Step 2:

[0247] The server processes the received photos of the room using image analysis technology to identify the room's size, color scheme, and existing interior. This extracts the room's physical characteristics as digital data.

[0248] Step 3:

[0249] The server analyzes data collected from social media to identify user preferences and lifestyles. This analysis provides attribute information such as the user's preferred style, colors, and themes.

[0250] Step 4:

[0251] The server utilizes a generative AI model to select the most suitable item candidates from a global furniture database based on collected room characteristics and user attribute information. The AI ​​model extracts furniture that meets the selection criteria.

[0252] Step 5:

[0253] The server sends information about the selected furniture candidates to the terminal. The terminal uses this information and augmented reality (AR) functionality to virtually place the furniture in the room. This allows the user to visually see how the furniture will be arranged.

[0254] Step 6:

[0255] Users operate a device to view virtually placed furniture from various angles and judge its harmony with the overall room and its suitability for the space.

[0256] Step 7:

[0257] The server collects and integrates sales information and customer reviews for each piece of furniture from online data. It then provides users with information that summarizes the advantages and disadvantages of each option.

[0258] Step 8:

[0259] Users select the most suitable furniture based on the information provided through their device and proceed with the purchase. This allows users to make decisions quickly.

[0260] (Example 1)

[0261] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0262] Traditionally, selecting furniture and interior items has presented challenges in making optimal choices that suit individual user preferences and spaces. Furthermore, a lack of visual support for making appropriate selections without physically inspecting the items meant that the selection process was time-consuming and laborious. Additionally, centrally collecting and presenting relevant pricing and user review information was difficult.

[0263] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0264] In this invention, the server includes means for collecting and analyzing user image information, means for recommending optimal items using a generating AI model based on the analyzed spatial structure and color tone, means for virtually arranging and visualizing the recommended items in space, and means for collecting, integrating, and displaying sales information and evaluation information related to the items. This enables the selection of items optimized for the user, allows for visual confirmation of their arrangement, and enables efficient decision-making by comprehensively viewing related information.

[0265] "Means for collecting and analyzing user image information" refers to a system that acquires image data provided by users and processes that data to identify the layout, color tone, and style of a space.

[0266] "A method for recommending optimal items using a generative AI model" refers to a system that utilizes AI technology to present highly relevant furniture and interior items based on collected user attribute information and spatial characteristics.

[0267] "A means of virtually placing and visualizing recommended items in a space" refers to a system that uses AR technology or similar methods to simulate and display selected furniture and interior items in the user's room in real time, providing an image of how they would actually look in place.

[0268] "Means for collecting, integrating, and displaying sales information and evaluation information related to goods" refers to a system that collects and organizes market price information, user reviews, and sales information related to presented furniture and interior items, and presents them comprehensively to users.

[0269] This invention provides a system for users to efficiently select furniture and interior items. This system operates through the interaction of a terminal, a server, and the user.

[0270] The device primarily provides the functionality for users to take photos of their rooms and upload them to the application. Users can also grant the system permission to access their social media account data. The captured images are sent from the device to the server, and the social media data is similarly transferred to the server.

[0271] The server uses an image analysis engine to analyze received image data and recognize the physical layout, color scheme, and interior style of the room. For social media data, attribute information is collected by identifying user preferences and lifestyles using natural language processing techniques. In this process, for example, the TensorFlow library is used for image analysis.

[0272] Next, the server uses a generative AI model to select highly relevant furniture from a global database based on the collected attribute and room information. The generative AI model uses machine learning algorithms to recommend furniture based on past data and current user attributes. At this point, the server sends the generative AI a prompt message stating, "Please suggest the most suitable furniture based on user attribute data."

[0273] Subsequently, the furniture selected by the server is sent to the terminal, which uses augmented reality (AR) technology to virtually place the furniture in the user's room. This process allows the user to visually confirm how the furniture will look. Furthermore, the server collects sales information, user reviews, and pricing information and presents it to the user as integrated information through the terminal. This enables the user to make efficient decisions.

[0274] For example, if a user wants to furnish their new home with modern-style furniture, they first take a photo of the room with their device and allow the use of their social media data. The server analyzes this information, recommends suitable sofas and tables, and provides a visualization of how they would look in the room using augmented reality (AR) technology. As a result, the user can effectively and efficiently select the interior items they desire.

[0275] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0276] Step 1:

[0277] The user takes a picture of their room using their device and uploads it to the application. The user can also grant permission for the application to use data from their social media account. The input consists of the captured image of the room and social media authentication information, while the output is a data file that can be sent. Specifically, the user activates the device's camera function and presses the capture button to acquire the image.

[0278] Step 2:

[0279] The terminal sends the captured image data and data from the SNS to the server. The input is the data file created earlier, and the output is the reception status at the server. The terminal uses the secure HTTPS protocol to send the data and displays a progress bar to notify the user of the completion of the transmission.

[0280] Step 3:

[0281] The server analyzes the received image data and obtains the results. The input is the image data received from the terminal, data processing is performed by an analysis engine, and the output is information on the layout and color tone of the room. Specifically, the server uses an image analysis engine (e.g., TensorFlow) to identify the color tone and layout of the image and stores it in JSON format.

[0282] Step 4:

[0283] The server analyzes the SNS data to identify the user's preferences and lifestyle. The input is the profile information of the SNS, and natural language processing technology is used for the analysis. The output is user attribute information. As a specific operation, the server uses a natural language processing library to analyze the text data and perform keyword extraction and topic modeling.

[0284] Step 5:

[0285] The server uses a generative AI model to recommend furniture. The input is the analyzed user attribute information and room information. The AI model searches for the optimal furniture from the big data, and the output is a list of recommended furniture. As a specific operation, the server sends the prompt sentence "Please propose the optimal furniture based on the user attribute data" to the AI model and obtains the response from the model.

[0286] Step 6:

[0287] A list of furniture selected from the server is sent to the terminal. The terminal then uses AR technology to virtually place these furniture items in the user's room. The input is furniture data from the server, and the output is an AR display visually presented to the user. Specifically, the furniture is displayed as virtual objects on the terminal's screen and can be viewed in real time by overlaying it with the camera image.

[0288] Step 7:

[0289] The server collects and integrates sales information, reviews, and pricing information related to the proposed furniture. The input is the furniture's ID information, and the output is presented to the user in a centralized information pack. Specifically, the server uses external APIs to retrieve the necessary data, organizes it, and sends it to the terminal.

[0290] (Application Example 1)

[0291] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0292] In today's retail environment, consumers are required to choose the optimal product from a vast selection, which is often a time-consuming and laborious task. Furthermore, even if consumers visually examine products in a physical store, it is difficult to confirm their suitability in their actual living space beforehand. This leads to problems such as increased dissatisfaction and the risk of returns after purchase.

[0293] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0294] In this invention, the server includes means for utilizing visual equipment to extract user requests, means for recommending optimal products based on the user's preferences and location structure information, and means for displaying and arranging products three-dimensionally using virtual reality technology. This allows consumers to check the suitability of products in their actual living spaces in advance and make purchase decisions efficiently.

[0295] "User requirements" refer to the specific needs and criteria that users have when choosing a product.

[0296] "Visual devices" refer to augmented reality devices and wearable devices that users use to virtually inspect products.

[0297] "Location structure information" refers to data about the physical layout and design of a user's residence or commercial space.

[0298] The "optimal product" is an item that has been determined to be the most suitable based on the user's requirements and preferences.

[0299] "Virtual reality technology" refers to the technology that uses computer technology to construct a virtual world, allowing users to visually experience objects within that world.

[0300] The embodiments for carrying out the invention will now be described. To realize this invention, the system is configured as follows.

[0301] The terminal will primarily utilize smart glasses or headset devices as visual devices worn by the user. These visual devices will capture images of in-store merchandise and transmit that data to a server. Specifically, smart glasses (e.g., Microsoft HoloLens, Google Glass) will be used.

[0302] The server receives the data transmitted from the visual device and analyzes the structural information of the location using an image analysis library (e.g., OpenCV). Also, to extract the preferences and requirements of the user, it analyzes the information collected from the past purchase history and SNS, and identifies the user attributes.

[0303] Furthermore, the server uses a generative AI model to recommend the optimal products from the product database based on the analyzed data. At this time, a prompt sentence is used to instruct the AI model. For example, a prompt sentence such as "Propose modern-style furniture that best suits this living room. I would like you to select based on this image and the user's preferences" is utilized.

[0304] On the terminal, using the information received from the server, it uses virtual reality technology to provide a visualization of the case where the product is actually placed in the user's living space. Thereby, the user can confirm in real time how the product looks in their own space.

[0305] As a specific example, when a user visiting a furniture store views a sofa using smart glasses, it is possible to see a virtual view of the sofa placed in the living room of their home. In this way, the user can confirm the suitability of the product selection before purchase.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The user wears a visual device and focuses on an interesting product in the store. The visual device (smart glasses) captures the video and transmits it to the server, including the user's gaze data. The input at this point is the video data of the product and the user's gaze data, and the output is the data transmission to the server.

[0309] Step 2:

[0310] The server analyzes the received video data using an image analysis library (e.g., OpenCV) to identify structural information of the location. This analysis extracts positional information and color tones related to the placement of products. The input is the captured video data, and the output is the spatial information obtained through the analysis.

[0311] Step 3:

[0312] The server analyzes the user's past purchase history and data from social media to identify user preferences and requests. This process uses data mining techniques to extract user attributes. The input is data about user attributes, and the output is the extracted preferences and purchasing trends.

[0313] Step 4:

[0314] The server utilizes a generative AI model to search and recommend the most suitable products from its product database based on analyzed spatial information and user preferences. Advanced recommendations are provided by giving instructions to the AI ​​using prompts. Input consists of spatial information and user preferences, while output is a list of recommended products.

[0315] Step 5:

[0316] The terminal uses recommended product information received from the server to visualize the product placement using virtual reality technology. It presents a virtual product placement to the user, making it appear as if the products are actually present in that space. The input is recommended product information, and the output is a virtual view with the products placed.

[0317] Step 6:

[0318] Users review the visualized product placement and evaluate how well the product fits into their living space. Based on this evaluation, users determine their purchase intent and make a selection. The input is visual information of the virtual product placement, and the output is the user's evaluation and decision.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention provides a method for realizing a system that recognizes a user's emotional state using an emotion engine and personalizes furniture selection and presentation based on that information. This system functions through interaction between a terminal, a server, and an emotion engine.

[0321] Users can first take photos of their room using their device and upload that data to the server via the application. Furthermore, users grant access to a feature that allows the emotion engine to analyze facial expressions and voice data.

[0322] The server analyzes the physical structure and color scheme from received images of the room, and analyzes social media data to identify the user's preferences and lifestyle patterns. In addition, the emotion engine analyzes multimodal data to capture emotions, such as the user's facial expressions, voice tone, and entered text. This analysis allows the user to obtain their current emotional state and past emotional history.

[0323] The server takes data obtained from the emotion engine into account and uses a generative AI model to recommend furniture that suits the user's preferences and emotional state. For example, if the analysis determines that the user is in an emotional state of wanting to relax, it can suggest sofas in calming colors and furniture with relaxing designs.

[0324] Next, the furniture candidates selected by the server are sent to the terminal. The terminal virtually places these furniture pieces using AR technology, providing the user with a visual simulation. Through this visualization, the user can see how the selected furniture harmonizes with the overall space.

[0325] Furthermore, the server aggregates sales information and customer reviews for each piece of furniture. This information is presented to the user via their device and used as a criterion for selection. Users can then select and purchase the most suitable furniture, taking into account their own emotional state and the actual condition of their room.

[0326] For example, if a user is feeling stressed and wants to add a calming element to their living space, the emotion engine will pick up on this information and recommend relaxing interior design. In this way, users can quickly create a comfortable living space that corresponds to their emotions.

[0327] The following describes the processing flow.

[0328] Step 1:

[0329] The user takes photos of the room using their device and uploads the data to the server through the application. The user also grants access to the emotion engine and configures settings to allow input of facial expressions and voice.

[0330] Step 2:

[0331] The server analyzes the received images of the room to identify its size, color scheme, furniture arrangement, and other characteristics. This image analysis generates digital data representing the physical features of the room.

[0332] Step 3:

[0333] The server analyzes the user's social media data to obtain attribute information about the user's preferences. This includes the user's preferred design style and color palette.

[0334] Step 4:

[0335] The emotion engine analyzes the user's facial expressions and voice to understand their emotional state in real time. Furthermore, by considering their past emotional history, it provides data to make more appropriate recommendations.

[0336] Step 5:

[0337] Based on data from the emotion engine, the server utilizes a generative AI model to recommend the most suitable furniture for the user's preferences and emotional state. For example, if the emotion of wanting to relax is recognized, furniture with a calming design will be selected.

[0338] Step 6:

[0339] The furniture candidates selected by the server are sent to the terminal. The terminal uses AR technology to virtually place the furniture in the user's room, providing a realistic visual simulation.

[0340] Step 7:

[0341] Users operate a device to view virtually placed furniture. They evaluate the harmony of the furniture with the space and its placement while completing a visual challenge.

[0342] Step 8:

[0343] The server collects and organizes sales information and reviews for each furniture item. This information is then presented to the user via their terminal for final decision-making.

[0344] Step 9:

[0345] Users select the most suitable furniture based on the information presented through their device and complete the purchase process. As a result, users can choose the optimal interior design that suits their emotions.

[0346] (Example 2)

[0347] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0348] To improve the comfort of modern living spaces, personalized interior design suggestions tailored to the user's emotional state are crucial. However, conventional systems have struggled to make recommendations that take emotional states into account, failing to adequately meet individual user needs. Therefore, there is a need for a system that accurately recognizes the user's emotional state and recommends the most suitable items based on that understanding.

[0349] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0350] In this invention, the server includes means for collecting data to recognize the user's emotional state, means for recommending the most suitable items based on the emotional state and spatial information, and means for arranging the items in a virtual space and displaying them using augmented reality technology. This makes it possible to suggest personalized interiors that correspond to the user's emotional state.

[0351] "Users" refer to individuals who use the system to receive interior design suggestions based on their emotional state.

[0352] "Emotional state" refers to the psychological condition obtained by analyzing multimodal data such as the user's facial expressions, voice tone, and entered text.

[0353] "Means of data collection" refers to the technical processes and devices used to obtain information such as images, audio, and text from users and transmit it to a server.

[0354] "Spatial information" refers to information that includes physical characteristics such as structure and color scheme related to the user's living space.

[0355] "Means of recommending optimal items" refers to algorithms and technologies that propose interiors deemed optimal for the user based on analyzed emotional states and spatial information.

[0356] "Means of placing objects in a virtual space and displaying them using augmented reality technology" refers to a technology that overlays selected objects as digital data onto the user's actual space and provides a visual simulation.

[0357] "Means of integrating and presenting evaluation information" refers to technologies and processes for aggregating reviews and evaluations of selected items and presenting them to users in an easy-to-understand manner.

[0358] A "generative AI model" refers to an artificial intelligence algorithm used to analyze a user's emotional state and lifestyle patterns and provide individually optimized recommendations.

[0359] A "prompt" refers to the text input to give specific instructions to a generative AI model.

[0360] The following describes the specific operations and techniques used in embodiments for carrying out the present invention.

[0361] This system consists of a user terminal, a server that processes data, and an emotion engine that analyzes emotions.

[0362] The user first takes a picture of the room using their device. This picture is sent to the server via the device. The user also consents to the collection of data on their facial expressions and voice tone, allowing the emotion engine to analyze it.

[0363] The server analyzes the received image data using image recognition software to extract spatial information such as the physical layout and color scheme of the room. The server also collects data from social media and other sources to identify the user's lifestyle patterns and preferences.

[0364] The emotion engine uses multimodal data (facial expressions, voice, text) obtained from the user to analyze the user's current emotional state in detail. This analysis also takes into account the user's usual emotional history.

[0365] Based on all this information, the server utilizes a generated AI model to select the most suitable interior for the user. For example, if it determines that the user is seeking relaxation, it will suggest furniture in calming colors.

[0366] The device uses augmented reality (AR) technology to virtually place furniture information received from the server into the user's room. This allows the user to intuitively see how the furniture would actually look in their room.

[0367] Furthermore, the server aggregates sales information and customer reviews related to the selected furniture and presents them to the user via the terminal. This allows users to make informed decisions when purchasing interior items that resonate with their emotions.

[0368] For example, a prompt such as "Suggest furniture that will help the user relax" is input into the AI ​​model in an understandable format. This allows for personalized recommendations.

[0369] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0370] Step 1:

[0371] The user takes a picture of the room using their device. The user's action activates the device's camera application, capturing an image of the room. This image data is saved to the device's memory and prepared for transmission.

[0372] Step 2:

[0373] The user uploads image data from their device to the server. The device sends the collected image data to the server via a secure communication protocol. The input is the captured image data, and the output is the data stored on the server.

[0374] Step 3:

[0375] The server receives and analyzes image data. Using an image recognition algorithm, it extracts spatial information such as physical structure and color tone. The input to this process is the received image data, and the output is structural and color tone information. Specifically, it identifies the location of walls and furniture through pixel analysis.

[0376] Step 4:

[0377] The user consents to the collection of emotional data through their device. With permission, the device transmits the user's facial expressions and voice data to the emotion engine. This input is real-time, multimodal data, and the output is data ready for emotion analysis.

[0378] Step 5:

[0379] The emotion engine analyzes the received data. Through facial expression analysis, voice tone analysis, and text analysis, it identifies the user's emotional state. This input is multimodal data, and the output is the user's emotional state and history. Specific operations include voice frequency analysis and facial expression change detection.

[0380] Step 6:

[0381] The server uses a generated AI model to recommend interior design suitable for the user. The input is analyzed emotional state and spatial information, and the output is a list of furniture suitable for the user. Specifically, the AI ​​combines past preference data and emotional data to select the optimal items.

[0382] Step 7:

[0383] The server sends recommended furniture information to the device. The device then uses AR technology to virtually place the received information in the room, providing the user with visual feedback. The input is furniture information, and the output is a visualization using augmented reality technology.

[0384] Step 8:

[0385] The server aggregates and displays sales information and customer reviews related to furniture. The input is a database of information on each piece of furniture, and the output is integrated evaluation information. Specifically, it uses recursive data queries to highlight highly-rated items.

[0386] Step 9:

[0387] The user selects and purchases the most suitable furniture. The user makes their purchase decision considering AR visualizations and evaluation information. The input is the presented interior information, and the output is a notification that the purchase process is complete. Specific actions include one-click purchase options and payment processing.

[0388] (Application Example 2)

[0389] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0390] Modern consumers tend to seek products quickly and accurately that suit their emotional state and individual circumstances when making purchasing decisions. However, many conventional systems have struggled to accurately grasp a user's emotional state and provide products optimized for it. Therefore, the challenge lies in realizing personalized product recommendations that respond to users' emotions and environmental information.

[0391] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0392] In this invention, the server includes means for analyzing the user's emotional state, means for analyzing the user's physical environment from images, and means for recommending items based on the user's emotional state and environmental information. This makes it possible to provide personalized items based on the user's real-time emotions and environment.

[0393] "Methods for analyzing a user's emotional state" refers to technologies that identify a user's current emotional state based on data obtained from the user's facial expressions, voice, text, etc.

[0394] "Methods for analyzing a user's physical environment from images" refers to technologies that process images of a space provided by the user and extract physical elements such as the room's structure, design, and colors.

[0395] "Means for recommending products based on the user's emotional state and environmental information" refers to a technology that uses analyzed emotional state and environmental data to select and present products suitable for the user.

[0396] "Means of virtually placing the item using augmented reality technology" refers to a method of using augmented reality technology to virtually visualize an item in the user's actual environment.

[0397] "Means for integrating and presenting product evaluation information" refers to a technology that collects third-party evaluations and related information about recommended products and presents them in an easy-to-understand manner for users.

[0398] To implement this invention, the user must first use a device such as a smartphone or head-mounted display to take pictures of their room and their own face. The device has the function of uploading this data to a server.

[0399] The server processes data using an emotion analysis engine to analyze the user's emotional state from facial expressions and voice tone. Open-source software OpenCV is used for facial recognition, and a proprietary voice analysis tool is applied for voice analysis. Furthermore, the server receives images of the room and analyzes the physical structure and color tones of the environment using OpenCV and other tools.

[0400] Next, the server uses a generative AI model to select the most suitable items based on the acquired emotional and environmental data. During this process, it generates specific prompt statements, and the AI ​​model recommends items based on these prompt statements. An example of such a prompt statement might be, "A prompt to suggest the most suitable furniture and interior design for a user experiencing stress."

[0401] The selected items are sent to the device and virtually placed in the room using augmented reality technology. This process utilizes the device's AR capabilities, allowing the user to visually see how the items blend in with the real-world space.

[0402] Furthermore, the server aggregates third-party evaluation information and sales information on items and provides it to the terminal. This enables users to make more informed decisions. Thus, the present invention provides a system that enables personalized item selection based on the user's emotions and environment.

[0403] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0404] Step 1:

[0405] Users use smartphones or head-mounted displays to capture images of their room and their own faces. This provides image data that captures the user's environment and emotions. This image data serves as foundational data for subsequent analysis.

[0406] Step 2:

[0407] The device uploads the acquired image data to the server. The input consists of images of the room and the user's face, and the output is the server receiving the data. The uploaded images are transferred to the server for analysis.

[0408] Step 3:

[0409] The server provides uploaded facial images to an emotion analysis engine for facial recognition and voice analysis. Using facial images and voice data as input, it obtains emotional state data as output. This emotional data is used to select items suitable for the user using a generative AI model.

[0410] Step 4:

[0411] The server analyzes images of a room to identify the structure and color tone of the physical environment. The input is image data of the room, and the output is spatial information and color tone data of the room. Image processing using OpenCV is used to extract physical features as digital data.

[0412] Step 5:

[0413] The server provides prompt text to a generating AI model based on emotional and environmental data, recommending the most suitable items for the user. The input includes emotional state and spatial information data, and the output is item recommendation results. It generates specific prompts such as "a prompt to suggest the most suitable furniture and interior for a user experiencing stress," and the AI ​​model selects the appropriate items.

[0414] Step 6:

[0415] The selected item information is sent to the terminal and virtually placed in the room using augmented reality (AR) functionality. The input is the item recommendation result, and the output is a visual simulation using AR. Through the terminal, the user can see in real time how the items will blend in with the room.

[0416] Step 7:

[0417] The server collects third-party reviews and sales information for recommended items and transmits it to the terminal. The input is item data, and the output is integrated reviews and sales information. This allows users to obtain data to support their decision-making regarding item selection.

[0418] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0419] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0420] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0421] [Third Embodiment]

[0422] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0423] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0424] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0425] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0426] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0427] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0428] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0429] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0430] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0431] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0432] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0433] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0434] This invention provides a method for realizing a system that enables users to select furniture efficiently and effectively. The system functions through the interaction of terminals, servers, and users.

[0435] First, the device provides the user with the ability to take photos of their room and upload them to the application. At this time, the user can grant the system access to their social media accounts. The device then sends this data to the server.

[0436] The server analyzes the received image data and uses an application engine to recognize the physical layout, color scheme, and existing interior style of the room. It also analyzes data from social media to identify the user's preferences and lifestyle. This process collects user attribute information.

[0437] Next, the server uses a generative AI model to search a global furniture database based on the collected attribute and room information, and selects suitable item candidates. This AI model has the ability to recommend highly relevant furniture based on past data and current user attributes.

[0438] Subsequently, the candidate furniture generated by the server is sent to the terminal. The terminal uses AR technology to provide a function that virtually places the suggested furniture in the user's room. Through the terminal, the user can visually check the virtually placed furniture and evaluate its harmony and fit within the space.

[0439] Furthermore, the server collects and integrates sales information, user reviews, and pricing information for each piece of furniture. This information is presented visually to the user via their terminal to assist in making decisions when choosing furniture.

[0440] As a concrete example, suppose a user is looking for modern-style furniture to decorate their new home. The user takes photos of the room with their device and allows the use of social media data. Based on this information, the server recommends several sofas and tables and provides a visualization of how they would look in the room using augmented reality (AR). The user then selects and purchases the best furniture, taking into account the latest sales information. This allows the user to achieve their ideal interior design while saving time and effort.

[0441] The following describes the processing flow.

[0442] Step 1:

[0443] The user takes photos of the room using their device. The photos are uploaded to the server through the application. The user can also grant the application access to their social media accounts.

[0444] Step 2:

[0445] The server processes the received photos of the room using image analysis technology to identify the room's size, color scheme, and existing interior. This extracts the room's physical characteristics as digital data.

[0446] Step 3:

[0447] The server analyzes data collected from social media to identify user preferences and lifestyles. This analysis provides attribute information such as the user's preferred style, colors, and themes.

[0448] Step 4:

[0449] The server utilizes a generative AI model to select the most suitable item candidates from a global furniture database based on collected room characteristics and user attribute information. The AI ​​model extracts furniture that meets the selection criteria.

[0450] Step 5:

[0451] The server sends information about the selected furniture candidates to the terminal. The terminal uses this information and augmented reality (AR) functionality to virtually place the furniture in the room. This allows the user to visually see how the furniture will be arranged.

[0452] Step 6:

[0453] Users operate a device to view virtually placed furniture from various angles and judge its harmony with the overall room and its suitability for the space.

[0454] Step 7:

[0455] The server collects and integrates sales information and customer reviews for each piece of furniture from online data. It then provides users with information that summarizes the advantages and disadvantages of each option.

[0456] Step 8:

[0457] Users select the most suitable furniture based on the information provided through their device and proceed with the purchase. This allows users to make decisions quickly.

[0458] (Example 1)

[0459] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0460] Traditionally, selecting furniture and interior items has presented challenges in making optimal choices that suit individual user preferences and spaces. Furthermore, a lack of visual support for making appropriate selections without physically inspecting the items meant that the selection process was time-consuming and laborious. Additionally, centrally collecting and presenting relevant pricing and user review information was difficult.

[0461] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0462] In this invention, the server includes means for collecting and analyzing user image information, means for recommending optimal items using a generating AI model based on the analyzed spatial structure and color tone, means for virtually arranging and visualizing the recommended items in space, and means for collecting, integrating, and displaying sales information and evaluation information related to the items. This enables the selection of items optimized for the user, allows for visual confirmation of their arrangement, and enables efficient decision-making by comprehensively viewing related information.

[0463] "Means for collecting and analyzing user image information" refers to a system that acquires image data provided by users and processes that data to identify the layout, color tone, and style of a space.

[0464] "A method for recommending optimal items using a generative AI model" refers to a system that utilizes AI technology to present highly relevant furniture and interior items based on collected user attribute information and spatial characteristics.

[0465] "A means of virtually placing and visualizing recommended items in a space" refers to a system that uses AR technology or similar methods to simulate and display selected furniture and interior items in the user's room in real time, providing an image of how they would actually look in place.

[0466] "Means for collecting, integrating, and displaying sales information and evaluation information related to goods" refers to a system that collects and organizes market price information, user reviews, and sales information related to presented furniture and interior items, and presents them comprehensively to users.

[0467] This invention provides a system for users to efficiently select furniture and interior items. This system operates through the interaction of a terminal, a server, and the user.

[0468] The device primarily provides the functionality for users to take photos of their rooms and upload them to the application. Users can also grant the system permission to access their social media account data. The captured images are sent from the device to the server, and the social media data is similarly transferred to the server.

[0469] The server uses an image analysis engine to analyze received image data and recognize the physical layout, color scheme, and interior style of the room. For social media data, attribute information is collected by identifying user preferences and lifestyles using natural language processing techniques. In this process, for example, the TensorFlow library is used for image analysis.

[0470] Next, the server uses a generative AI model to select highly relevant furniture from a global database based on the collected attribute and room information. The generative AI model uses machine learning algorithms to recommend furniture based on past data and current user attributes. At this point, the server sends the generative AI a prompt message stating, "Please suggest the most suitable furniture based on user attribute data."

[0471] Subsequently, the furniture selected by the server is sent to the terminal, which uses augmented reality (AR) technology to virtually place the furniture in the user's room. This process allows the user to visually confirm how the furniture will look. Furthermore, the server collects sales information, user reviews, and pricing information and presents it to the user as integrated information through the terminal. This enables the user to make efficient decisions.

[0472] For example, if a user wants to furnish their new home with modern-style furniture, they first take a photo of the room with their device and allow the use of their social media data. The server analyzes this information, recommends suitable sofas and tables, and provides a visualization of how they would look in the room using augmented reality (AR) technology. As a result, the user can effectively and efficiently select the interior items they desire.

[0473] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0474] Step 1:

[0475] The user takes a picture of their room using their device and uploads it to the application. The user can also grant permission for the application to use data from their social media account. The input consists of the captured image of the room and social media authentication information, while the output is a data file that can be sent. Specifically, the user activates the device's camera function and presses the capture button to acquire the image.

[0476] Step 2:

[0477] The device sends captured image data and data from social media to the server. The input is the data file created earlier, and the output is the server's reception status. The device sends the data using the secure HTTPS protocol and displays a progress bar to notify the user when the transmission is complete.

[0478] Step 3:

[0479] The server analyzes the received image data and retrieves the results. The input is image data received from the terminal, which is processed by the analysis engine, and the output is information about the room layout and color tone. Specifically, the server uses an image analysis engine (e.g., TensorFlow) to identify the color tone and layout of the image and stores it in JSON format.

[0480] Step 4:

[0481] The server analyzes SNS data to identify user preferences and lifestyles. The input is SNS profile information, and natural language processing techniques are used for analysis. The output is user attribute information. Specifically, the server uses a natural language processing library to analyze text data, performing keyword extraction and topic modeling.

[0482] Step 5:

[0483] The server uses a generative AI model to recommend furniture. The input consists of analyzed user attribute information and room information. The AI ​​model searches for the most suitable furniture from the big data, and the output is a list of recommended furniture. Specifically, the server sends a prompt to the AI ​​model, "Please suggest the most suitable furniture based on the user attribute data," and receives a response from the model.

[0484] Step 6:

[0485] A list of furniture selected from the server is sent to the terminal. The terminal then uses AR technology to virtually place these furniture items in the user's room. The input is furniture data from the server, and the output is an AR display visually presented to the user. Specifically, the furniture is displayed as virtual objects on the terminal's screen and can be viewed in real time by overlaying it with the camera image.

[0486] Step 7:

[0487] The server collects and integrates sales information, reviews, and pricing information related to the proposed furniture. The input is the furniture's ID information, and the output is presented to the user in a centralized information pack. Specifically, the server uses external APIs to retrieve the necessary data, organizes it, and sends it to the terminal.

[0488] (Application Example 1)

[0489] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0490] In today's retail environment, consumers are required to choose the optimal product from a vast selection, which is often a time-consuming and laborious task. Furthermore, even if consumers visually examine products in a physical store, it is difficult to confirm their suitability in their actual living space beforehand. This leads to problems such as increased dissatisfaction and the risk of returns after purchase.

[0491] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0492] In this invention, the server includes means for utilizing visual equipment to extract user requests, means for recommending optimal products based on the user's preferences and location structure information, and means for displaying and arranging products three-dimensionally using virtual reality technology. This allows consumers to check the suitability of products in their actual living spaces in advance and make purchase decisions efficiently.

[0493] "User requirements" refer to the specific needs and criteria that users have when choosing a product.

[0494] "Visual devices" refer to augmented reality devices and wearable devices that users use to virtually inspect products.

[0495] "Location structure information" refers to data about the physical layout and design of a user's residence or commercial space.

[0496] The "optimal product" is an item that has been determined to be the most suitable based on the user's requirements and preferences.

[0497] "Virtual reality technology" refers to the technology that uses computer technology to construct a virtual world, allowing users to visually experience objects within that world.

[0498] The embodiments for carrying out the invention will now be described. To realize this invention, the system is configured as follows.

[0499] The terminal will primarily utilize smart glasses or headset devices as visual devices worn by the user. These visual devices will capture images of in-store merchandise and transmit that data to a server. Specifically, smart glasses (e.g., Microsoft HoloLens, Google Glass) will be used.

[0500] The server receives data transmitted from visual devices and analyzes the structural information of the location using an image analysis library (e.g., OpenCV). It also analyzes past purchase history and information collected from social media to identify user attributes and extract user preferences and requests.

[0501] Furthermore, the server uses a generative AI model to recommend the most suitable products from the product database based on the analyzed data. In this process, prompts are used to instruct the AI ​​model. For example, a prompt might say, "Suggest the most modern style furniture for this living room. Please select based on this image and the user's preferences."

[0502] The terminal uses information received from the server and virtual reality technology to provide a visualization of how the product would look if it were actually placed in the user's living space. This allows the user to see in real time how the product would look in their own space.

[0503] As a concrete example, when a user visiting a furniture store uses smart glasses to view a sofa, they can see a virtual view of the sofa placed in their own living room. In this way, users can confirm the suitability of their product selection before purchasing it.

[0504] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0505] Step 1:

[0506] The user wears a visual device and focuses their gaze on an item of interest in the store. The visual device (smart glasses) captures the image and sends it to the server along with the user's gaze data. At this point, the input is the image data of the product and the user's gaze data, and the output is the transmission of data to the server.

[0507] Step 2:

[0508] The server analyzes the received video data using an image analysis library (e.g., OpenCV) to identify structural information of the location. This analysis extracts positional information and color tones related to the placement of products. The input is the captured video data, and the output is the spatial information obtained through the analysis.

[0509] Step 3:

[0510] The server analyzes the user's past purchase history and data from social media to identify user preferences and requests. This process uses data mining techniques to extract user attributes. The input is data about user attributes, and the output is the extracted preferences and purchasing trends.

[0511] Step 4:

[0512] The server utilizes a generative AI model to search and recommend the most suitable products from its product database based on analyzed spatial information and user preferences. Advanced recommendations are provided by giving instructions to the AI ​​using prompts. Input consists of spatial information and user preferences, while output is a list of recommended products.

[0513] Step 5:

[0514] The terminal uses recommended product information received from the server to visualize the product placement using virtual reality technology. It presents a virtual product placement to the user, making it appear as if the products are actually present in that space. The input is recommended product information, and the output is a virtual view with the products placed.

[0515] Step 6:

[0516] Users review the visualized product placement and evaluate how well the product fits into their living space. Based on this evaluation, users determine their purchase intent and make a selection. The input is visual information of the virtual product placement, and the output is the user's evaluation and decision.

[0517] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0518] This invention provides a method for realizing a system that recognizes a user's emotional state using an emotion engine and personalizes furniture selection and presentation based on that information. This system functions through interaction between a terminal, a server, and an emotion engine.

[0519] Users can first take photos of their room using their device and upload that data to the server via the application. Furthermore, users grant access to a feature that allows the emotion engine to analyze facial expressions and voice data.

[0520] The server analyzes the physical structure and color scheme from received images of the room, and analyzes social media data to identify the user's preferences and lifestyle patterns. In addition, the emotion engine analyzes multimodal data to capture emotions, such as the user's facial expressions, voice tone, and entered text. This analysis allows the user to obtain their current emotional state and past emotional history.

[0521] The server takes data obtained from the emotion engine into account and uses a generative AI model to recommend furniture that suits the user's preferences and emotional state. For example, if the analysis determines that the user is in an emotional state of wanting to relax, it can suggest sofas in calming colors and furniture with relaxing designs.

[0522] Next, the furniture candidates selected by the server are sent to the terminal. The terminal virtually places these furniture pieces using AR technology, providing the user with a visual simulation. Through this visualization, the user can see how the selected furniture harmonizes with the overall space.

[0523] Furthermore, the server aggregates sales information and customer reviews for each piece of furniture. This information is presented to the user via their device and used as a criterion for selection. Users can then select and purchase the most suitable furniture, taking into account their own emotional state and the actual condition of their room.

[0524] For example, if a user is feeling stressed and wants to add a calming element to their living space, the emotion engine will pick up on this information and recommend relaxing interior design. In this way, users can quickly create a comfortable living space that corresponds to their emotions.

[0525] The following describes the processing flow.

[0526] Step 1:

[0527] The user takes photos of the room using their device and uploads the data to the server through the application. The user also grants access to the emotion engine and configures settings to allow input of facial expressions and voice.

[0528] Step 2:

[0529] The server analyzes the received images of the room to identify its size, color scheme, furniture arrangement, and other characteristics. This image analysis generates digital data representing the physical features of the room.

[0530] Step 3:

[0531] The server analyzes the user's social media data to obtain attribute information about the user's preferences. This includes the user's preferred design style and color palette.

[0532] Step 4:

[0533] The emotion engine analyzes the user's facial expressions and voice to understand their emotional state in real time. Furthermore, by considering their past emotional history, it provides data to make more appropriate recommendations.

[0534] Step 5:

[0535] Based on data from the emotion engine, the server utilizes a generative AI model to recommend the most suitable furniture for the user's preferences and emotional state. For example, if the emotion of wanting to relax is recognized, furniture with a calming design will be selected.

[0536] Step 6:

[0537] The furniture candidates selected by the server are sent to the terminal. The terminal uses AR technology to virtually place the furniture in the user's room, providing a realistic visual simulation.

[0538] Step 7:

[0539] Users operate a device to view virtually placed furniture. They evaluate the harmony of the furniture with the space and its placement while completing a visual challenge.

[0540] Step 8:

[0541] The server collects and organizes sales information and reviews for each furniture item. This information is then presented to the user via their terminal for final decision-making.

[0542] Step 9:

[0543] Users select the most suitable furniture based on the information presented through their device and complete the purchase process. As a result, users can choose the optimal interior design that suits their emotions.

[0544] (Example 2)

[0545] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0546] To improve the comfort of modern living spaces, personalized interior design suggestions tailored to the user's emotional state are crucial. However, conventional systems have struggled to make recommendations that take emotional states into account, failing to adequately meet individual user needs. Therefore, there is a need for a system that accurately recognizes the user's emotional state and recommends the most suitable items based on that understanding.

[0547] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0548] In this invention, the server includes means for collecting data to recognize the user's emotional state, means for recommending the most suitable items based on the emotional state and spatial information, and means for arranging the items in a virtual space and displaying them using augmented reality technology. This makes it possible to suggest personalized interiors that correspond to the user's emotional state.

[0549] "Users" refer to individuals who use the system to receive interior design suggestions based on their emotional state.

[0550] "Emotional state" refers to the psychological condition obtained by analyzing multimodal data such as the user's facial expressions, voice tone, and entered text.

[0551] "Means of data collection" refers to the technical processes and devices used to obtain information such as images, audio, and text from users and transmit it to a server.

[0552] "Spatial information" refers to information that includes physical characteristics such as structure and color scheme related to the user's living space.

[0553] "Means of recommending optimal items" refers to algorithms and technologies that propose interiors deemed optimal for the user based on analyzed emotional states and spatial information.

[0554] "Means of placing objects in a virtual space and displaying them using augmented reality technology" refers to a technology that overlays selected objects as digital data onto the user's actual space and provides a visual simulation.

[0555] "Means of integrating and presenting evaluation information" refers to technologies and processes for aggregating reviews and evaluations of selected items and presenting them to users in an easy-to-understand manner.

[0556] A "generative AI model" refers to an artificial intelligence algorithm used to analyze a user's emotional state and lifestyle patterns and provide individually optimized recommendations.

[0557] A "prompt" refers to the text input to give specific instructions to a generative AI model.

[0558] The following describes the specific operations and techniques used in embodiments for carrying out the present invention.

[0559] This system consists of a user terminal, a server that processes data, and an emotion engine that analyzes emotions.

[0560] The user first takes a picture of the room using their device. This picture is sent to the server via the device. The user also consents to the collection of data on their facial expressions and voice tone, allowing the emotion engine to analyze it.

[0561] The server analyzes the received image data using image recognition software to extract spatial information such as the physical layout and color scheme of the room. The server also collects data from social media and other sources to identify the user's lifestyle patterns and preferences.

[0562] The emotion engine uses multimodal data (facial expressions, voice, text) obtained from the user to analyze the user's current emotional state in detail. This analysis also takes into account the user's usual emotional history.

[0563] Based on all this information, the server utilizes a generated AI model to select the most suitable interior for the user. For example, if it determines that the user is seeking relaxation, it will suggest furniture in calming colors.

[0564] The device uses augmented reality (AR) technology to virtually place furniture information received from the server into the user's room. This allows the user to intuitively see how the furniture would actually look in their room.

[0565] Furthermore, the server aggregates sales information and customer reviews related to the selected furniture and presents them to the user via the terminal. This allows users to make informed decisions when purchasing interior items that resonate with their emotions.

[0566] For example, a prompt such as "Suggest furniture that will help the user relax" is input into the AI ​​model in an understandable format. This allows for personalized recommendations.

[0567] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0568] Step 1:

[0569] The user takes a picture of the room using their device. The user's action activates the device's camera application, capturing an image of the room. This image data is saved to the device's memory and prepared for transmission.

[0570] Step 2:

[0571] The user uploads image data from their device to the server. The device sends the collected image data to the server via a secure communication protocol. The input is the captured image data, and the output is the data stored on the server.

[0572] Step 3:

[0573] The server receives and analyzes image data. Using an image recognition algorithm, it extracts spatial information such as physical structure and color tone. The input to this process is the received image data, and the output is structural and color tone information. Specifically, it identifies the location of walls and furniture through pixel analysis.

[0574] Step 4:

[0575] The user consents to the collection of emotional data through their device. With permission, the device transmits the user's facial expressions and voice data to the emotion engine. This input is real-time, multimodal data, and the output is data ready for emotion analysis.

[0576] Step 5:

[0577] The emotion engine analyzes the received data. Through facial expression analysis, voice tone analysis, and text analysis, it identifies the user's emotional state. This input is multimodal data, and the output is the user's emotional state and history. Specific operations include voice frequency analysis and facial expression change detection.

[0578] Step 6:

[0579] The server uses a generated AI model to recommend interior design suitable for the user. The input is analyzed emotional state and spatial information, and the output is a list of furniture suitable for the user. Specifically, the AI ​​combines past preference data and emotional data to select the optimal items.

[0580] Step 7:

[0581] The server sends recommended furniture information to the device. The device then uses AR technology to virtually place the received information in the room, providing the user with visual feedback. The input is furniture information, and the output is a visualization using augmented reality technology.

[0582] Step 8:

[0583] The server aggregates and displays sales information and customer reviews related to furniture. The input is a database of information on each piece of furniture, and the output is integrated evaluation information. Specifically, it uses recursive data queries to highlight highly-rated items.

[0584] Step 9:

[0585] The user selects and purchases the most suitable furniture. The user makes their purchase decision considering AR visualizations and evaluation information. The input is the presented interior information, and the output is a notification that the purchase process is complete. Specific actions include one-click purchase options and payment processing.

[0586] (Application Example 2)

[0587] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] Modern consumers tend to seek products quickly and accurately that suit their emotional state and individual circumstances when making purchasing decisions. However, many conventional systems have struggled to accurately grasp a user's emotional state and provide products optimized for it. Therefore, the challenge lies in realizing personalized product recommendations that respond to users' emotions and environmental information.

[0589] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0590] In this invention, the server includes means for analyzing the user's emotional state, means for analyzing the user's physical environment from images, and means for recommending items based on the user's emotional state and environmental information. This makes it possible to provide personalized items based on the user's real-time emotions and environment.

[0591] "Methods for analyzing a user's emotional state" refers to technologies that identify a user's current emotional state based on data obtained from the user's facial expressions, voice, text, etc.

[0592] "Methods for analyzing a user's physical environment from images" refers to technologies that process images of a space provided by the user and extract physical elements such as the room's structure, design, and colors.

[0593] "Means for recommending products based on the user's emotional state and environmental information" refers to a technology that uses analyzed emotional state and environmental data to select and present products suitable for the user.

[0594] "Means of virtually placing the item using augmented reality technology" refers to a method of using augmented reality technology to virtually visualize an item in the user's actual environment.

[0595] "Means for integrating and presenting product evaluation information" refers to a technology that collects third-party evaluations and related information about recommended products and presents them in an easy-to-understand manner for users.

[0596] To implement this invention, the user must first use a device such as a smartphone or head-mounted display to take pictures of their room and their own face. The device has the function of uploading this data to a server.

[0597] The server processes data using an emotion analysis engine to analyze the user's emotional state from facial expressions and voice tone. Open-source software OpenCV is used for facial recognition, and a proprietary voice analysis tool is applied for voice analysis. Furthermore, the server receives images of the room and analyzes the physical structure and color tones of the environment using OpenCV and other tools.

[0598] Next, the server uses a generative AI model to select the most suitable items based on the acquired emotional and environmental data. During this process, it generates specific prompt statements, and the AI ​​model recommends items based on these prompt statements. An example of such a prompt statement might be, "A prompt to suggest the most suitable furniture and interior design for a user experiencing stress."

[0599] The selected items are sent to the device and virtually placed in the room using augmented reality technology. This process utilizes the device's AR capabilities, allowing the user to visually see how the items blend in with the real-world space.

[0600] Furthermore, the server aggregates third-party evaluation information and sales information on items and provides it to the terminal. This enables users to make more informed decisions. Thus, the present invention provides a system that enables personalized item selection based on the user's emotions and environment.

[0601] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0602] Step 1:

[0603] Users use smartphones or head-mounted displays to capture images of their room and their own faces. This provides image data that captures the user's environment and emotions. This image data serves as foundational data for subsequent analysis.

[0604] Step 2:

[0605] The device uploads the acquired image data to the server. The input consists of images of the room and the user's face, and the output is the server receiving the data. The uploaded images are transferred to the server for analysis.

[0606] Step 3:

[0607] The server provides uploaded facial images to an emotion analysis engine for facial recognition and voice analysis. Using facial images and voice data as input, it obtains emotional state data as output. This emotional data is used to select items suitable for the user using a generative AI model.

[0608] Step 4:

[0609] The server analyzes images of a room to identify the structure and color tone of the physical environment. The input is image data of the room, and the output is spatial information and color tone data of the room. Image processing using OpenCV is used to extract physical features as digital data.

[0610] Step 5:

[0611] The server provides prompt text to a generating AI model based on emotional and environmental data, recommending the most suitable items for the user. The input includes emotional state and spatial information data, and the output is item recommendation results. It generates specific prompts such as "a prompt to suggest the most suitable furniture and interior for a user experiencing stress," and the AI ​​model selects the appropriate items.

[0612] Step 6:

[0613] The selected item information is sent to the terminal and virtually placed in the room using augmented reality (AR) functionality. The input is the item recommendation result, and the output is a visual simulation using AR. Through the terminal, the user can see in real time how the items will blend in with the room.

[0614] Step 7:

[0615] The server collects third-party reviews and sales information for recommended items and transmits it to the terminal. The input is item data, and the output is integrated reviews and sales information. This allows users to obtain data to support their decision-making regarding item selection.

[0616] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0617] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0618] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0619] [Fourth Embodiment]

[0620] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0621] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0622] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0623] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0624] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0625] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0626] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0627] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0628] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0629] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0630] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0631] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0632] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0633] This invention provides a method for realizing a system that enables users to select furniture efficiently and effectively. The system functions through the interaction of terminals, servers, and users.

[0634] First, the device provides the user with the ability to take photos of their room and upload them to the application. At this time, the user can grant the system access to their social media accounts. The device then sends this data to the server.

[0635] The server analyzes the received image data and uses an application engine to recognize the physical layout, color scheme, and existing interior style of the room. It also analyzes data from social media to identify the user's preferences and lifestyle. This process collects user attribute information.

[0636] Next, the server uses a generative AI model to search a global furniture database based on the collected attribute and room information, and selects suitable item candidates. This AI model has the ability to recommend highly relevant furniture based on past data and current user attributes.

[0637] Subsequently, the candidate furniture generated by the server is sent to the terminal. The terminal uses AR technology to provide a function that virtually places the suggested furniture in the user's room. Through the terminal, the user can visually check the virtually placed furniture and evaluate its harmony and fit within the space.

[0638] Furthermore, the server collects and integrates sales information, user reviews, and pricing information for each piece of furniture. This information is presented visually to the user via their terminal to assist in making decisions when choosing furniture.

[0639] As a concrete example, suppose a user is looking for modern-style furniture to decorate their new home. The user takes photos of the room with their device and allows the use of social media data. Based on this information, the server recommends several sofas and tables and provides a visualization of how they would look in the room using augmented reality (AR). The user then selects and purchases the best furniture, taking into account the latest sales information. This allows the user to achieve their ideal interior design while saving time and effort.

[0640] The following describes the processing flow.

[0641] Step 1:

[0642] The user takes photos of the room using their device. The photos are uploaded to the server through the application. The user can also grant the application access to their social media accounts.

[0643] Step 2:

[0644] The server processes the received photos of the room using image analysis technology to identify the room's size, color scheme, and existing interior. This extracts the room's physical characteristics as digital data.

[0645] Step 3:

[0646] The server analyzes data collected from social media to identify user preferences and lifestyles. This analysis provides attribute information such as the user's preferred style, colors, and themes.

[0647] Step 4:

[0648] The server utilizes a generative AI model to select the most suitable item candidates from a global furniture database based on collected room characteristics and user attribute information. The AI ​​model extracts furniture that meets the selection criteria.

[0649] Step 5:

[0650] The server sends information about the selected furniture candidates to the terminal. The terminal uses this information and augmented reality (AR) functionality to virtually place the furniture in the room. This allows the user to visually see how the furniture will be arranged.

[0651] Step 6:

[0652] Users operate a device to view virtually placed furniture from various angles and judge its harmony with the overall room and its suitability for the space.

[0653] Step 7:

[0654] The server collects and integrates sales information and customer reviews for each piece of furniture from online data. It then provides users with information that summarizes the advantages and disadvantages of each option.

[0655] Step 8:

[0656] Users select the most suitable furniture based on the information provided through their device and proceed with the purchase. This allows users to make decisions quickly.

[0657] (Example 1)

[0658] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0659] Traditionally, selecting furniture and interior items has presented challenges in making optimal choices that suit individual user preferences and spaces. Furthermore, a lack of visual support for making appropriate selections without physically inspecting the items meant that the selection process was time-consuming and laborious. Additionally, centrally collecting and presenting relevant pricing and user review information was difficult.

[0660] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0661] In this invention, the server includes means for collecting and analyzing user image information, means for recommending optimal items using a generating AI model based on the analyzed spatial structure and color tone, means for virtually arranging and visualizing the recommended items in space, and means for collecting, integrating, and displaying sales information and evaluation information related to the items. This enables the selection of items optimized for the user, allows for visual confirmation of their arrangement, and enables efficient decision-making by comprehensively viewing related information.

[0662] "Means for collecting and analyzing user image information" refers to a system that acquires image data provided by users and processes that data to identify the layout, color tone, and style of a space.

[0663] "A method for recommending optimal items using a generative AI model" refers to a system that utilizes AI technology to present highly relevant furniture and interior items based on collected user attribute information and spatial characteristics.

[0664] "A means of virtually placing and visualizing recommended items in a space" refers to a system that uses AR technology or similar methods to simulate and display selected furniture and interior items in the user's room in real time, providing an image of how they would actually look in place.

[0665] "Means for collecting, integrating, and displaying sales information and evaluation information related to goods" refers to a system that collects and organizes market price information, user reviews, and sales information related to presented furniture and interior items, and presents them comprehensively to users.

[0666] This invention provides a system for users to efficiently select furniture and interior items. This system operates through the interaction of a terminal, a server, and the user.

[0667] The device primarily provides the functionality for users to take photos of their rooms and upload them to the application. Users can also grant the system permission to access their social media account data. The captured images are sent from the device to the server, and the social media data is similarly transferred to the server.

[0668] The server uses an image analysis engine to analyze received image data and recognize the physical layout, color scheme, and interior style of the room. For social media data, attribute information is collected by identifying user preferences and lifestyles using natural language processing techniques. In this process, for example, the TensorFlow library is used for image analysis.

[0669] Next, the server uses a generative AI model to select highly relevant furniture from a global database based on the collected attribute and room information. The generative AI model uses machine learning algorithms to recommend furniture based on past data and current user attributes. At this point, the server sends the generative AI a prompt message stating, "Please suggest the most suitable furniture based on user attribute data."

[0670] Subsequently, the furniture selected by the server is sent to the terminal, which uses augmented reality (AR) technology to virtually place the furniture in the user's room. This process allows the user to visually confirm how the furniture will look. Furthermore, the server collects sales information, user reviews, and pricing information and presents it to the user as integrated information through the terminal. This enables the user to make efficient decisions.

[0671] For example, if a user wants to furnish their new home with modern-style furniture, they first take a photo of the room with their device and allow the use of their social media data. The server analyzes this information, recommends suitable sofas and tables, and provides a visualization of how they would look in the room using augmented reality (AR) technology. As a result, the user can effectively and efficiently select the interior items they desire.

[0672] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0673] Step 1:

[0674] The user takes a picture of their room using their device and uploads it to the application. The user can also grant permission for the application to use data from their social media account. The input consists of the captured image of the room and social media authentication information, while the output is a data file that can be sent. Specifically, the user activates the device's camera function and presses the capture button to acquire the image.

[0675] Step 2:

[0676] The device sends captured image data and data from social media to the server. The input is the data file created earlier, and the output is the server's reception status. The device sends the data using the secure HTTPS protocol and displays a progress bar to notify the user when the transmission is complete.

[0677] Step 3:

[0678] The server analyzes the received image data and retrieves the results. The input is image data received from the terminal, which is processed by the analysis engine, and the output is information about the room layout and color tone. Specifically, the server uses an image analysis engine (e.g., TensorFlow) to identify the color tone and layout of the image and stores it in JSON format.

[0679] Step 4:

[0680] The server analyzes SNS data to identify user preferences and lifestyles. The input is SNS profile information, and natural language processing techniques are used for analysis. The output is user attribute information. Specifically, the server uses a natural language processing library to analyze text data, performing keyword extraction and topic modeling.

[0681] Step 5:

[0682] The server uses a generative AI model to recommend furniture. The input consists of analyzed user attribute information and room information. The AI ​​model searches for the most suitable furniture from the big data, and the output is a list of recommended furniture. Specifically, the server sends a prompt to the AI ​​model, "Please suggest the most suitable furniture based on the user attribute data," and receives a response from the model.

[0683] Step 6:

[0684] A list of furniture selected from the server is sent to the terminal. The terminal then uses AR technology to virtually place these furniture items in the user's room. The input is furniture data from the server, and the output is an AR display visually presented to the user. Specifically, the furniture is displayed as virtual objects on the terminal's screen and can be viewed in real time by overlaying it with the camera image.

[0685] Step 7:

[0686] The server collects and integrates sales information, reviews, and pricing information related to the proposed furniture. The input is the furniture's ID information, and the output is presented to the user in a centralized information pack. Specifically, the server uses external APIs to retrieve the necessary data, organizes it, and sends it to the terminal.

[0687] (Application Example 1)

[0688] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0689] In today's retail environment, consumers are required to choose the optimal product from a vast selection, which is often a time-consuming and laborious task. Furthermore, even if consumers visually examine products in a physical store, it is difficult to confirm their suitability in their actual living space beforehand. This leads to problems such as increased dissatisfaction and the risk of returns after purchase.

[0690] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0691] In this invention, the server includes means for utilizing visual equipment to extract user requests, means for recommending optimal products based on the user's preferences and location structure information, and means for displaying and arranging products three-dimensionally using virtual reality technology. This allows consumers to check the suitability of products in their actual living spaces in advance and make purchase decisions efficiently.

[0692] "User requirements" refer to the specific needs and criteria that users have when choosing a product.

[0693] "Visual devices" refer to augmented reality devices and wearable devices that users use to virtually inspect products.

[0694] "Location structure information" refers to data about the physical layout and design of a user's residence or commercial space.

[0695] The "optimal product" is an item that has been determined to be the most suitable based on the user's requirements and preferences.

[0696] "Virtual reality technology" refers to the technology that uses computer technology to construct a virtual world, allowing users to visually experience objects within that world.

[0697] The embodiments for carrying out the invention will now be described. To realize this invention, the system is configured as follows.

[0698] The terminal will primarily utilize smart glasses or headset devices as visual devices worn by the user. These visual devices will capture images of in-store merchandise and transmit that data to a server. Specifically, smart glasses (e.g., Microsoft HoloLens, Google Glass) will be used.

[0699] The server receives data transmitted from visual devices and analyzes the structural information of the location using an image analysis library (e.g., OpenCV). It also analyzes past purchase history and information collected from social media to identify user attributes and extract user preferences and requests.

[0700] Furthermore, the server uses a generative AI model to recommend the most suitable products from the product database based on the analyzed data. In this process, prompts are used to instruct the AI ​​model. For example, a prompt might say, "Suggest the most modern style furniture for this living room. Please select based on this image and the user's preferences."

[0701] The terminal uses information received from the server and virtual reality technology to provide a visualization of how the product would look if it were actually placed in the user's living space. This allows the user to see in real time how the product would look in their own space.

[0702] As a concrete example, when a user visiting a furniture store uses smart glasses to view a sofa, they can see a virtual view of the sofa placed in their own living room. In this way, users can confirm the suitability of their product selection before purchasing it.

[0703] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0704] Step 1:

[0705] The user wears a visual device and focuses their gaze on an item of interest in the store. The visual device (smart glasses) captures the image and sends it to the server along with the user's gaze data. At this point, the input is the image data of the product and the user's gaze data, and the output is the transmission of data to the server.

[0706] Step 2:

[0707] The server analyzes the received video data using an image analysis library (e.g., OpenCV) to identify structural information of the location. This analysis extracts positional information and color tones related to the placement of products. The input is the captured video data, and the output is the spatial information obtained through the analysis.

[0708] Step 3:

[0709] The server analyzes the user's past purchase history and data from social media to identify user preferences and requests. This process uses data mining techniques to extract user attributes. The input is data about user attributes, and the output is the extracted preferences and purchasing trends.

[0710] Step 4:

[0711] The server utilizes a generative AI model to search and recommend the most suitable products from its product database based on analyzed spatial information and user preferences. Advanced recommendations are provided by giving instructions to the AI ​​using prompts. Input consists of spatial information and user preferences, while output is a list of recommended products.

[0712] Step 5:

[0713] The terminal uses recommended product information received from the server to visualize the product placement using virtual reality technology. It presents a virtual product placement to the user, making it appear as if the products are actually present in that space. The input is recommended product information, and the output is a virtual view with the products placed.

[0714] Step 6:

[0715] Users review the visualized product placement and evaluate how well the product fits into their living space. Based on this evaluation, users determine their purchase intent and make a selection. The input is visual information of the virtual product placement, and the output is the user's evaluation and decision.

[0716] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0717] This invention provides a method for realizing a system that recognizes a user's emotional state using an emotion engine and personalizes furniture selection and presentation based on that information. This system functions through interaction between a terminal, a server, and an emotion engine.

[0718] Users can first take photos of their room using their device and upload that data to the server via the application. Furthermore, users grant access to a feature that allows the emotion engine to analyze facial expressions and voice data.

[0719] The server analyzes the physical structure and color scheme from received images of the room, and analyzes social media data to identify the user's preferences and lifestyle patterns. In addition, the emotion engine analyzes multimodal data to capture emotions, such as the user's facial expressions, voice tone, and entered text. This analysis allows the user to obtain their current emotional state and past emotional history.

[0720] The server takes data obtained from the emotion engine into account and uses a generative AI model to recommend furniture that suits the user's preferences and emotional state. For example, if the analysis determines that the user is in an emotional state of wanting to relax, it can suggest sofas in calming colors and furniture with relaxing designs.

[0721] Next, the furniture candidates selected by the server are sent to the terminal. The terminal virtually places these furniture pieces using AR technology, providing the user with a visual simulation. Through this visualization, the user can see how the selected furniture harmonizes with the overall space.

[0722] Furthermore, the server aggregates sales information and customer reviews for each piece of furniture. This information is presented to the user via their device and used as a criterion for selection. Users can then select and purchase the most suitable furniture, taking into account their own emotional state and the actual condition of their room.

[0723] For example, if a user is feeling stressed and wants to add a calming element to their living space, the emotion engine will pick up on this information and recommend relaxing interior design. In this way, users can quickly create a comfortable living space that corresponds to their emotions.

[0724] The following describes the processing flow.

[0725] Step 1:

[0726] The user takes photos of the room using their device and uploads the data to the server through the application. The user also grants access to the emotion engine and configures settings to allow input of facial expressions and voice.

[0727] Step 2:

[0728] The server analyzes the received images of the room to identify its size, color scheme, furniture arrangement, and other characteristics. This image analysis generates digital data representing the physical features of the room.

[0729] Step 3:

[0730] The server analyzes the user's social media data to obtain attribute information about the user's preferences. This includes the user's preferred design style and color palette.

[0731] Step 4:

[0732] The emotion engine analyzes the user's facial expressions and voice to understand their emotional state in real time. Furthermore, by considering their past emotional history, it provides data to make more appropriate recommendations.

[0733] Step 5:

[0734] Based on data from the emotion engine, the server utilizes a generative AI model to recommend the most suitable furniture for the user's preferences and emotional state. For example, if the emotion of wanting to relax is recognized, furniture with a calming design will be selected.

[0735] Step 6:

[0736] The furniture candidates selected by the server are sent to the terminal. The terminal uses AR technology to virtually place the furniture in the user's room, providing a realistic visual simulation.

[0737] Step 7:

[0738] Users operate a device to view virtually placed furniture. They evaluate the harmony of the furniture with the space and its placement while completing a visual challenge.

[0739] Step 8:

[0740] The server collects and organizes sales information and reviews for each furniture item. This information is then presented to the user via their terminal for final decision-making.

[0741] Step 9:

[0742] Users select the most suitable furniture based on the information presented through their device and complete the purchase process. As a result, users can choose the optimal interior design that suits their emotions.

[0743] (Example 2)

[0744] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0745] To improve the comfort of modern living spaces, personalized interior design suggestions tailored to the user's emotional state are crucial. However, conventional systems have struggled to make recommendations that take emotional states into account, failing to adequately meet individual user needs. Therefore, there is a need for a system that accurately recognizes the user's emotional state and recommends the most suitable items based on that understanding.

[0746] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0747] In this invention, the server includes means for collecting data to recognize the user's emotional state, means for recommending the most suitable items based on the emotional state and spatial information, and means for arranging the items in a virtual space and displaying them using augmented reality technology. This makes it possible to suggest personalized interiors that correspond to the user's emotional state.

[0748] "Users" refer to individuals who use the system to receive interior design suggestions based on their emotional state.

[0749] "Emotional state" refers to the psychological condition obtained by analyzing multimodal data such as the user's facial expressions, voice tone, and entered text.

[0750] "Means of data collection" refers to the technical processes and devices used to obtain information such as images, audio, and text from users and transmit it to a server.

[0751] "Spatial information" refers to information that includes physical characteristics such as structure and color scheme related to the user's living space.

[0752] "Means of recommending optimal items" refers to algorithms and technologies that propose interiors deemed optimal for the user based on analyzed emotional states and spatial information.

[0753] "Means of placing objects in a virtual space and displaying them using augmented reality technology" refers to a technology that overlays selected objects as digital data onto the user's actual space and provides a visual simulation.

[0754] "Means of integrating and presenting evaluation information" refers to technologies and processes for aggregating reviews and evaluations of selected items and presenting them to users in an easy-to-understand manner.

[0755] A "generative AI model" refers to an artificial intelligence algorithm used to analyze a user's emotional state and lifestyle patterns and provide individually optimized recommendations.

[0756] A "prompt" refers to the text input to give specific instructions to a generative AI model.

[0757] The following describes the specific operations and techniques used in embodiments for carrying out the present invention.

[0758] This system consists of a user terminal, a server that processes data, and an emotion engine that analyzes emotions.

[0759] The user first takes a picture of the room using their device. This picture is sent to the server via the device. The user also consents to the collection of data on their facial expressions and voice tone, allowing the emotion engine to analyze it.

[0760] The server analyzes the received image data using image recognition software to extract spatial information such as the physical layout and color scheme of the room. The server also collects data from social media and other sources to identify the user's lifestyle patterns and preferences.

[0761] The emotion engine uses multimodal data (facial expressions, voice, text) obtained from the user to analyze the user's current emotional state in detail. This analysis also takes into account the user's usual emotional history.

[0762] Based on all this information, the server utilizes a generated AI model to select the most suitable interior for the user. For example, if it determines that the user is seeking relaxation, it will suggest furniture in calming colors.

[0763] The device uses augmented reality (AR) technology to virtually place furniture information received from the server into the user's room. This allows the user to intuitively see how the furniture would actually look in their room.

[0764] Furthermore, the server aggregates sales information and customer reviews related to the selected furniture and presents them to the user via the terminal. This allows users to make informed decisions when purchasing interior items that resonate with their emotions.

[0765] For example, a prompt such as "Suggest furniture that will help the user relax" is input into the AI ​​model in an understandable format. This allows for personalized recommendations.

[0766] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0767] Step 1:

[0768] The user takes a picture of the room using their device. The user's action activates the device's camera application, capturing an image of the room. This image data is saved to the device's memory and prepared for transmission.

[0769] Step 2:

[0770] The user uploads image data from their device to the server. The device sends the collected image data to the server via a secure communication protocol. The input is the captured image data, and the output is the data stored on the server.

[0771] Step 3:

[0772] The server receives and analyzes image data. Using an image recognition algorithm, it extracts spatial information such as physical structure and color tone. The input to this process is the received image data, and the output is structural and color tone information. Specifically, it identifies the location of walls and furniture through pixel analysis.

[0773] Step 4:

[0774] The user consents to the collection of emotional data through their device. With permission, the device transmits the user's facial expressions and voice data to the emotion engine. This input is real-time, multimodal data, and the output is data ready for emotion analysis.

[0775] Step 5:

[0776] The emotion engine analyzes the received data. Through facial expression analysis, voice tone analysis, and text analysis, it identifies the user's emotional state. This input is multimodal data, and the output is the user's emotional state and history. Specific operations include voice frequency analysis and facial expression change detection.

[0777] Step 6:

[0778] The server uses a generated AI model to recommend interior design suitable for the user. The input is analyzed emotional state and spatial information, and the output is a list of furniture suitable for the user. Specifically, the AI ​​combines past preference data and emotional data to select the optimal items.

[0779] Step 7:

[0780] The server sends recommended furniture information to the device. The device then uses AR technology to virtually place the received information in the room, providing the user with visual feedback. The input is furniture information, and the output is a visualization using augmented reality technology.

[0781] Step 8:

[0782] The server aggregates and displays sales information and customer reviews related to furniture. The input is a database of information on each piece of furniture, and the output is integrated evaluation information. Specifically, it uses recursive data queries to highlight highly-rated items.

[0783] Step 9:

[0784] The user selects and purchases the most suitable furniture. The user makes their purchase decision considering AR visualizations and evaluation information. The input is the presented interior information, and the output is a notification that the purchase process is complete. Specific actions include one-click purchase options and payment processing.

[0785] (Application Example 2)

[0786] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0787] Modern consumers tend to seek products quickly and accurately that suit their emotional state and individual circumstances when making purchasing decisions. However, many conventional systems have struggled to accurately grasp a user's emotional state and provide products optimized for it. Therefore, the challenge lies in realizing personalized product recommendations that respond to users' emotions and environmental information.

[0788] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0789] In this invention, the server includes means for analyzing the user's emotional state, means for analyzing the user's physical environment from images, and means for recommending items based on the user's emotional state and environmental information. This makes it possible to provide personalized items based on the user's real-time emotions and environment.

[0790] "Methods for analyzing a user's emotional state" refers to technologies that identify a user's current emotional state based on data obtained from the user's facial expressions, voice, text, etc.

[0791] "Methods for analyzing a user's physical environment from images" refers to technologies that process images of a space provided by the user and extract physical elements such as the room's structure, design, and colors.

[0792] "Means for recommending products based on the user's emotional state and environmental information" refers to a technology that uses analyzed emotional state and environmental data to select and present products suitable for the user.

[0793] "Means of virtually placing the item using augmented reality technology" refers to a method of using augmented reality technology to virtually visualize an item in the user's actual environment.

[0794] "Means for integrating and presenting product evaluation information" refers to a technology that collects third-party evaluations and related information about recommended products and presents them in an easy-to-understand manner for users.

[0795] To implement this invention, the user must first use a device such as a smartphone or head-mounted display to take pictures of their room and their own face. The device has the function of uploading this data to a server.

[0796] The server processes data using an emotion analysis engine to analyze the user's emotional state from facial expressions and voice tone. Open-source software OpenCV is used for facial recognition, and a proprietary voice analysis tool is applied for voice analysis. Furthermore, the server receives images of the room and analyzes the physical structure and color tones of the environment using OpenCV and other tools.

[0797] Next, the server uses a generative AI model to select the most suitable items based on the acquired emotional and environmental data. During this process, it generates specific prompt statements, and the AI ​​model recommends items based on these prompt statements. An example of such a prompt statement might be, "A prompt to suggest the most suitable furniture and interior design for a user experiencing stress."

[0798] The selected items are sent to the device and virtually placed in the room using augmented reality technology. This process utilizes the device's AR capabilities, allowing the user to visually see how the items blend in with the real-world space.

[0799] Furthermore, the server aggregates third-party evaluation information and sales information on items and provides it to the terminal. This enables users to make more informed decisions. Thus, the present invention provides a system that enables personalized item selection based on the user's emotions and environment.

[0800] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0801] Step 1:

[0802] Users use smartphones or head-mounted displays to capture images of their room and their own faces. This provides image data that captures the user's environment and emotions. This image data serves as foundational data for subsequent analysis.

[0803] Step 2:

[0804] The device uploads the acquired image data to the server. The input consists of images of the room and the user's face, and the output is the server receiving the data. The uploaded images are transferred to the server for analysis.

[0805] Step 3:

[0806] The server provides uploaded facial images to an emotion analysis engine for facial recognition and voice analysis. Using facial images and voice data as input, it obtains emotional state data as output. This emotional data is used to select items suitable for the user using a generative AI model.

[0807] Step 4:

[0808] The server analyzes images of a room to identify the structure and color tone of the physical environment. The input is image data of the room, and the output is spatial information and color tone data of the room. Image processing using OpenCV is used to extract physical features as digital data.

[0809] Step 5:

[0810] The server provides prompt text to a generating AI model based on emotional and environmental data, recommending the most suitable items for the user. The input includes emotional state and spatial information data, and the output is item recommendation results. It generates specific prompts such as "a prompt to suggest the most suitable furniture and interior for a user experiencing stress," and the AI ​​model selects the appropriate items.

[0811] Step 6:

[0812] The selected item information is sent to the terminal and virtually placed in the room using augmented reality (AR) functionality. The input is the item recommendation result, and the output is a visual simulation using AR. Through the terminal, the user can see in real time how the items will blend in with the room.

[0813] Step 7:

[0814] The server collects third-party reviews and sales information for recommended items and transmits it to the terminal. The input is item data, and the output is integrated reviews and sales information. This allows users to obtain data to support their decision-making regarding item selection.

[0815] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0816] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0817] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0818] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0819] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0820] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0821] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0822] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0823] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0824] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0825] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0826] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0827] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0828] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0829] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0830] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0831] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0832] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0833] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0834] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0835] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0836] The following is further disclosed regarding the embodiments described above.

[0837] (Claim 1)

[0838] Means for collecting user information,

[0839] A means of recommending the most suitable items based on the user's preferences and spatial information,

[0840] A means for virtually arranging and displaying items,

[0841] A means of integrating and presenting evaluation information for goods,

[0842] A system that includes this.

[0843] (Claim 2)

[0844] The system according to claim 1, which analyzes collected image data to identify spatial structure and color tone.

[0845] (Claim 3)

[0846] The system according to claim 1, which analyzes user attribute data and generates personalized recommended items.

[0847] "Example 1"

[0848] (Claim 1)

[0849] A means of collecting and analyzing user image information,

[0850] A means of recommending the most suitable item using a generative AI model based on the analyzed spatial structure and color tone,

[0851] A means of virtually placing and visualizing the recommended items in space,

[0852] A means of collecting, integrating, and displaying sales information and evaluation information related to goods,

[0853] A system that includes this.

[0854] (Claim 2)

[0855] The system according to claim 1, which processes collected image information to identify spatial configuration and interior style.

[0856] (Claim 3)

[0857] The system according to claim 1, which analyzes user preference data and generates personalized product candidates based on that data.

[0858] "Application Example 1"

[0859] (Claim 1)

[0860] A means of utilizing visual devices to extract user requirements,

[0861] A means of recommending the most suitable product based on user preferences and location structure information,

[0862] A means of displaying and arranging products in 3D using virtual reality technology,

[0863] A means of providing integrated market information and pricing information for a product,

[0864] A system that includes this.

[0865] (Claim 2)

[0866] The system according to claim 1, which analyzes collected visual data to identify the layout and color scheme of a place.

[0867] (Claim 3)

[0868] The system according to claim 1, which analyzes individual user information and generates personalized recommended products.

[0869] "Example 2 of combining an emotion engine"

[0870] (Claim 1)

[0871] A means of collecting data to recognize the emotional state of users,

[0872] A means of recommending the most suitable items based on emotional state and spatial information,

[0873] A means of placing objects in a virtual space and displaying them using augmented reality technology,

[0874] A means of integrating and presenting evaluation information for items that corresponds to the emotional state of the user,

[0875] A system that includes this.

[0876] (Claim 2)

[0877] The system according to claim 1, which analyzes collected multimodal data to identify the emotional state of the user.

[0878] (Claim 3)

[0879] The system according to claim 1, which uses a generative AI model to generate personalized recommended items based on the user's emotional state and lifestyle patterns.

[0880] "Application example 2 when combining with an emotional engine"

[0881] (Claim 1)

[0882] A means of analyzing the emotional state of users,

[0883] A means of analyzing the user's physical environment from images,

[0884] A means of recommending products based on the user's emotional state and environmental information,

[0885] A means of virtually placing the item using augmented reality technology,

[0886] A means of integrating and presenting evaluation information for goods,

[0887] A system that includes this.

[0888] (Claim 2)

[0889] The system according to claim 1, which analyzes and identifies spatial information from image data.

[0890] (Claim 3)

[0891] The system according to claim 1, which generates personalized recommended items based on emotion analysis. [Explanation of Symbols]

[0892] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for collecting user information, A means of recommending the most suitable items based on the user's preferences and spatial information, A means for virtually arranging and displaying items, A means of integrating and presenting evaluation information for goods, A system that includes this.

2. The system according to claim 1, which analyzes collected image data to identify spatial structure and color tone.

3. The system according to claim 1, which analyzes user attribute data and generates personalized recommended items.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A