system

A system using image analysis and AI optimizes home environments by identifying dead spaces and suggesting personalized product choices, enhancing space utilization and convenience.

JP2026074928APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Individuals face challenges in optimally organizing and utilizing their home environments due to the difficulty in identifying dead spaces and making efficient choices among furniture and storage items, exacerbated by the need for space optimization in telecommuting and remote lifestyles.

Method used

A system that uses image analysis and AI to identify environmental information, generate storage and display suggestions, select products from market information, and provide purchase links, thereby optimizing home space utilization and addressing dead spaces.

Benefits of technology

Enables users to efficiently utilize their home spaces by identifying dead spaces and providing personalized product suggestions, improving the quality of life and convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074928000001_ABST
    Figure 2026074928000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for receiving image data acquired by a camera and processing the image data to identify environmental information, A generating device that generates proposals regarding storage and display based on the aforementioned environmental information, A selection device that, by referring to market information, selects products related to the aforementioned proposal and provides purchase links for said products, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern busy lifestyles, many individuals face the problem that they cannot optimally organize and arrange their home environments. This problem has become more prominent in combination with the social background where the efficient use of home space is required due to the spread of telecommuting and remote lifestyles. Also, there is a problem that although there are a variety of furniture and storage items in the market, it is not easy to make an optimal choice among them. Furthermore, since it is also difficult to find ways to utilize dead spaces that users themselves do not notice, there is a current situation where efficient space utilization is difficult to achieve.

Means for Solving the Problems

[0005] This invention provides a means for identifying environmental information based on image data acquired using a camera. Furthermore, based on the identified environmental information, it is possible to make optimal suggestions using a storage and display generation device. The invention also includes a selection device that references market information to select products related to the suggestions and provides purchase links for those products. In this way, users can easily and efficiently utilize their home space. Additionally, it is possible to detect dead spaces from the generated suggestions and propose storage solutions based on that. This allows users to effectively utilize spaces they may not have noticed themselves, thereby improving their quality of life.

[0006] A "photography device" is a device used to acquire visual information about the environment, and includes devices such as cameras and smartphones.

[0007] "Image data" refers to data that records visual information acquired in digital format by a camera or other imaging device.

[0008] "Environmental information" refers to detailed information about the room layout and furniture arrangement extracted from image data.

[0009] A "generation device" is a system that has the function of generating suggestions regarding storage and display based on identified environmental information.

[0010] "Market information" refers to a database containing the latest product information on furniture and storage items, which is referenced to provide users with the best possible recommendations.

[0011] A "selection device" is a system that selects products related to a proposal from market information and provides purchase links for those products.

[0012] "Dead space" refers to unused space in terms of furniture arrangement or room design, representing potential areas that can be utilized efficiently.

[0013] "Storage methods" refer to methods and items for organizing and keeping belongings, and are techniques for effectively using room space. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention provides a storage and display suggestion system using image analysis technology and AI to enable users to easily optimize their rooms at home. This system is generally accessible and easy to use via smartphones and tablets, and is implemented in the following form.

[0036] First, the user takes a photo of a specific room in their home using a mobile device such as a smartphone. It is recommended that the photo be taken in a way that shows the overall layout of the room. The captured image data is sent to the server via the application.

[0037] The server executes a deep learning-based image recognition algorithm to analyze the received image data. This identifies environmental information such as the room layout, furniture types and placement, and space usage. In particular, the image analysis can detect dead spaces that are not being used efficiently.

[0038] Next, a generator built into the server generates suggestions for improving storage and display based on environmental information. These suggestions include new storage methods utilizing dead space, furniture rearrangement, and harmonious display methods. Furthermore, the suggested ideas are optimized for the detected dead space.

[0039] Furthermore, the server accesses a market information database and selects appropriate furniture and storage items related to the suggestion. This selection is based on the user's needs and the condition of the room, ensuring the most suitable choice. The selected items are associated with online store purchase links, allowing the user to easily proceed with the purchase process.

[0040] The terminal displays suggestions and product selection information received from the server to the user. The user interface consists of easy-to-understand graphics and text, allowing users to easily understand the suggestions and use purchase links as needed.

[0041] For example, if a user takes a photo of their living room, the server will identify the placement of the sofa and TV stand and suggest a rack that can be installed in the corner space. A link to an online store will also be provided, allowing the user to purchase the recommended rack and secure new storage space. In this way, the present invention helps to easily and efficiently improve the user's living environment.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user launches a dedicated app on their smartphone and takes a photo that captures the entire room. The app then provides shooting guidelines to help the user take the best possible image.

[0045] Step 2:

[0046] The device compresses the captured image data within the app and sends it to the server via the network. During this process, an appropriate data transfer method is selected, taking into account the user's network bandwidth.

[0047] Step 3:

[0048] The server analyzes the received image data. Using a deep learning model, it identifies the room layout, furniture types, and their placement, and stores this information in a database as environmental data.

[0049] Step 4:

[0050] Based on the analyzed environmental information, the server generates optimal storage and display suggestions using a generation device. These suggestions include specific ideas for utilizing identified dead spaces.

[0051] Step 5:

[0052] The server selects furniture and storage items related to the generated proposal by referring to a market information database. Purchase links to online stores are added to the selected products.

[0053] Step 6:

[0054] The server formats the proposed content and selected product information and sends it to the user's terminal. This data is structured in a format suitable for the user interface.

[0055] Step 7:

[0056] The device visually presents the received data to the user. The app screen clearly displays suggested storage ideas and purchase links, supporting the user's decision-making.

[0057] Step 8:

[0058] Users review the displayed suggestions and purchase information and click on product links that interest them. This directs them to the online store, where they can proceed with the purchase.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] Conventional storage and display suggestion systems made it difficult for users to obtain concrete guidance on how to efficiently utilize their home space. Furthermore, the selection and purchase of items using market information was complex and time-consuming for users. Additionally, the inability to incorporate feedback on the suggested content made it difficult to improve the system's accuracy.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes a device that receives image information acquired by a shooting function and analyzes the image information to identify spatial information; a generation device that creates suggestions regarding storage and display based on the spatial information; and a selection device that refers to market information and provides the selection of items related to the suggestions and links to purchase those items. As a result, users can receive specific suggestions for effectively utilizing their home space and make quick and appropriate item selections and purchases based on market information. It is also possible to improve the accuracy of the system by utilizing feedback from users.

[0064] The "shooting function" is a feature that allows users to acquire visual information of their environment using their own devices.

[0065] "Image information" refers to visual data acquired through the camera's capture function, and forms the basis for analysis.

[0066] "Spatial information" refers to information about the physical structure and arrangement of the environment obtained from analyzed image information.

[0067] A "generation device" is a device that has the function of creating proposals for efficient storage and display based on spatial information.

[0068] "Market information" refers to data related to the provision of goods and services, and is used to select items related to the proposed content.

[0069] A "selection device" is a device that, based on market information, selects items that match the proposed content and provides the user with a purchase link.

[0070] A "display device" is a device used to visually present selected items to the user, encouraging them to confirm the details of the proposal.

[0071] "Unused space" refers to space that is not currently being effectively utilized, and is a target for proposals regarding new storage methods and furniture arrangements.

[0072] "Feedback information" refers to the opinions and evaluations that users provide regarding the proposed content, and is used as feedback for system improvement.

[0073] This invention utilizes a shooting function, server, terminal, and generating AI model to enable users to efficiently optimize their home space. Specific embodiments of this system are described below.

[0074] Users use devices such as smartphones or tablets to take pictures of specific rooms in their homes. The captured image information is automatically sent to a server via the application. Through this process, data is collected on the server in real time.

[0075] The server utilizes deep learning technology to analyze the received image information. In particular, it uses Convolutional Neural Networks (CNNs) and other techniques to perform object recognition and spatial analysis within the image, thereby identifying spatial information. This allows for the understanding of details such as dead space and furniture placement.

[0076] Next, a generation device integrated into the server uses a generation AI model based on the analyzed spatial information to create storage and display proposals. This model is pre-trained and provides efficient suggestions for diverse spatial layouts. The generated proposals include ideas for effectively utilizing underutilized space.

[0077] Furthermore, the server references market information and selects the most suitable items related to the proposed improvements. The selection device lists the suitable items and associates each with a purchase link to an online store. This allows users to easily purchase the products.

[0078] The terminal graphically displays the received proposals and item information. The user interface is designed to allow users to visually review the proposals, making them easy to understand and act upon.

[0079] For example, when a user takes a photo of their living room, the server recognizes the arrangement of sofas and tables and identifies unused space. The generator then suggests the best storage solution for this space, finds suitable furniture online, and provides purchase links. An example of a prompt message is, "Analyze the photo of my living room and suggest an optimization for storage space."

[0080] This invention enables users to utilize their daily living spaces more efficiently, significantly improving convenience.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] Users take photos of their rooms at home using their smartphones or tablets. It is recommended that the photos be composed in a way that shows the overall layout of the room. These images are automatically sent to the server via the application. The input is the image data taken by the user, and the output is the transfer of the image data to the server.

[0084] Step 2:

[0085] The server uses deep learning technology to analyze the received image data. Specifically, it uses a Convolutional Neural Network (CNN) to identify objects within the image and determine spatial information such as floor plan, furniture arrangement, color, and shape. The input is image data sent by the user, and the output is detailed spatial information obtained through image analysis.

[0086] Step 3:

[0087] The server uses a generative AI model based on acquired spatial information to create storage and display suggestions. This generative AI model is trained to generate efficient storage methods and furniture arrangement patterns. The input is analyzed spatial information, and the output is specific improvement suggestions. This process includes ideas for utilizing underutilized spaces.

[0088] Step 4:

[0089] The server automatically selects suitable items by referencing market information related to the generated proposal. This selection considers factors such as item size, design, and price. The selected items are also accompanied by online store purchase links. The input is the generated proposal and market information, and the output is a list of items aligned with the proposal and their purchase links.

[0090] Step 5:

[0091] The terminal displays the suggested content and item information received from the server via a user interface. This interface is designed to be visually intuitive, allowing users to verify suggested storage methods and furniture arrangements using 3D models. Input consists of suggestions and item information from the server, while output is a visual display and information provided to the user.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] In modern brick-and-mortar stores, product display and space utilization are largely based on experience, making automation and the use of information technology difficult. In particular, store staff need considerable time and effort to visually analyze space and determine the optimal display. Furthermore, while new display strategies are constantly needed to enhance customer purchasing intent, obtaining objective and real-time optimization suggestions remains challenging.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes means for receiving image information acquired by a camera and processing the image information to identify the spatial situation; generation means for generating suggestions regarding storage and display based on the spatial situation; and selection means for selecting products related to the suggestions by referring to market information and providing purchase links for those products. This makes it possible to present efficient use of space and optimization of product display in physical stores in real time through a wearable information display device.

[0097] "Photography equipment" refers to a device used to acquire image information of a specific subject.

[0098] "Image information" refers to data acquired by photographic equipment, which visually represents the spatial situation and the arrangement of objects.

[0099] "Spatial conditions" refer to information that describes the physical environment and the arrangement of objects in a specific location.

[0100] "Generating means" refers to a function for creating proposals regarding storage and display based on a specified spatial situation.

[0101] "Market information" refers to data related to the market, such as product types, prices, and distribution information.

[0102] "Selection method" refers to a function that selects the most suitable product based on market information and presents relevant information to the user.

[0103] A "wearable information display device" is a device that is worn by a user to visually present information.

[0104] "Wasted space" refers to space that is not physically used or is not being utilized efficiently.

[0105] "User feedback" refers to information that includes opinions and evaluations from users regarding a proposal.

[0106] A "learning tool" is a function that uses feedback information to improve the system's performance and the accuracy of its suggestions.

[0107] To implement this invention, the following system configuration can be used. The server receives image information transmitted from the imaging device and runs a deep learning model using image recognition technology to analyze it. This analysis identifies spatial conditions and wasted space. For example, the open-source deep learning library TENSORFLOW® can be used.

[0108] Once the analysis is complete, the server generates storage and display suggestions based on the results using a generation mechanism. These suggestions include efficient product display and ways to utilize wasted space. The generated suggestions are displayed in real time on a wearable information display device, such as smart glasses.

[0109] Next, the server references market information, selects products suitable for the user using selection methods, and provides relevant information and purchase links. This process involves gathering information from online databases.

[0110] Users can view information presented through smart glasses and intuitively adjust displays in the actual store space. When a specific suggestion is selected, user feedback is accumulated in the system, and suggestions are improved through learning mechanisms.

[0111] As a concrete example, when a staff member at a sporting goods store uses smart glasses to take pictures of the store, a server analyzes the wasted space on the shelves. Subsequently, suggestions for displaying new basketballs in the empty spaces are displayed, along with a purchase link for those products. This is expected to improve the efficiency of product placement and enhance customer satisfaction.

[0112] An example of a prompt using a generative AI model would be: "Analyze the space in the input image and suggest an efficient way to display products. For example, show what kind of products should be displayed in the empty space in the corner."

[0113] Such a system can enable efficient use of space and improved customer experience, especially in physical stores.

[0114] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0115] Step 1:

[0116] The user wears smart glasses and walks around the store, capturing image data of designated areas. This image data is then used as input for the system.

[0117] Step 2:

[0118] The terminal sends the acquired image information to the server. The input data is sent in image formats such as JPEG and PNG. The server receives this image information and uses a deep learning model to analyze the spatial situation. In this analysis, an object recognition algorithm identifies wasted space and existing display arrangements. As a result of the analysis, information on wasted space is output.

[0119] Step 3:

[0120] The server uses an AI model generated based on the analysis results to produce suggestions for storage and display that utilize wasted space. The input data includes spatial information from the analysis results and information on existing merchandise. The generated suggestions become the server's output, presenting optimal product placement and display methods.

[0121] Step 4:

[0122] The server uses the generated suggestions as a reference, queries relevant market information, and selects products using selection criteria. The input data is the generated suggestions, and the output is information on products that match the suggestions and purchase links. This information is obtained from an online database.

[0123] Step 5:

[0124] The terminal displays the suggested content and selected product information received from the server on the user's wearable information display device. The input is display data generated by the server, and the output is display information visually presented on smart glasses. This allows the user to confirm how to use the space while looking at the actual wasted space.

[0125] Step 6:

[0126] Users can implement the proposed installation method and test the results. They can also provide feedback on the usefulness of the proposal and the results of adopting it. This feedback is recorded as input to the system and used by the server's learning mechanism to improve the accuracy of future proposals.

[0127] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0128] This invention combines a system that analyzes image data related to the user's room layout and proposes optimal storage and display solutions with an emotion engine that recognizes the user's emotions and reflects them in the suggestions. As a result, users can receive suggestions optimized for their own emotional state, enabling them to create a more satisfying living space.

[0129] First, the user takes a picture of a specific room in their home using their smartphone camera. This image data is then sent to a server via a dedicated app.

[0130] The server executes an image recognition algorithm to analyze the received image data, identifying environmental information such as room layout, furniture arrangement, and dead space. This analysis builds the foundational data for proposals that enable efficient use of space.

[0131] Next, the server activates an emotion engine to recognize the user's emotions from their facial expressions and tone of voice during the photo shoot. This emotion data is used to generate suggestions that reflect the user's current mood and preferences. For example, if the user prefers a relaxed atmosphere, the server might suggest soft-colored curtains or the placement of houseplants.

[0132] Furthermore, the server references market information and selects products based on sentiment recognition results. This selection includes furniture and interior goods that reflect the user's emotions and preferences. Purchase links for the selected products are provided so that users can access them immediately if they are interested.

[0133] The device displays these suggestions to the user in an easy-to-understand interface. Users can browse the displayed suggestions and products and make selections that match their emotional state. They can also provide feedback on the suggestions, and this feedback is used to improve the accuracy of future suggestions through the emotion engine's learning function.

[0134] As a concrete example, if a user takes a photo of their living room, the server analyzes the image to identify the placement of sofas and tables, and uses an emotion engine to generate suggestions for creating a relaxing space. Furthermore, purchase links for products related to the suggested interior are provided, allowing the user to access the purchase page with a single click. In this way, the present invention, which combines an emotion engine, supports the creation of a more personalized space by providing customized suggestions that take the user's emotions into consideration.

[0135] The following describes the processing flow.

[0136] Step 1:

[0137] The user takes photos of their room at home using a dedicated app. During this process, the app activates the camera to detect the user's facial expressions and records their emotions.

[0138] Step 2:

[0139] The device sends captured image data and user facial expression data to the server. This data is treated as basic information for analysis.

[0140] Step 3:

[0141] The server analyzes the received image data and uses an AI model to identify the room layout and furniture arrangement. At the same time, it detects dead space and registers it in a database.

[0142] Step 4:

[0143] The server activates an emotion engine to recognize emotions from the user's facial expression data. This emotion data is used to adjust the suggestions.

[0144] Step 5:

[0145] The server generates optimal storage and display suggestions based on environmental information and emotional data. The generation device then performs this process, including color coordination and item selection tailored to the emotional state.

[0146] Step 6:

[0147] Based on the proposal, the server selects relevant products from the market information database. Purchase links are then set up for the selected products to allow for easy access.

[0148] Step 7:

[0149] The server sends the suggested content and product information to the terminal. The data is formatted in a way that is easy for the user to understand intuitively.

[0150] Step 8:

[0151] The device displays suggestions through its user interface. Users can view the presented information and click on product links that interest them.

[0152] Step 9:

[0153] Users access external online stores through selected links and complete purchases. They can also provide feedback on the suggestions they receive within the app.

[0154] (Example 2)

[0155] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0156] In modern living spaces, users often face the challenge of efficiently utilizing limited space while creating an environment that is optimal for their individual emotions and preferences. Furthermore, selecting the right products from the many options available on the market is not easy. Therefore, there is a need for a system that automatically provides suggestions that match the user's emotional state and preferences, thereby improving the efficiency of storage and interior design selection.

[0157] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0158] In this invention, the server includes means for receiving visual data acquired by an imaging device and analyzing the visual data to identify structural information; a generation device for creating storage and display suggestions based on the structural information; and an emotion analysis device for analyzing emotional information from audio and visual information and adjusting suggestions based on the user's psychological state. This makes it possible to efficiently utilize limited space and provide personalized suggestions that match the user's emotions and preferences.

[0159] An "imaging device" is a hardware device used to acquire visual data and is used to capture images of the user's environment.

[0160] "Visual data" refers to image information acquired by an imaging device, and serves as the basis for analyzing the structure and layout of a room.

[0161] "Structural information" refers to information about the room layout and furniture arrangement, identified through the analysis of visual data.

[0162] A "generation device" is a functional unit within a system that generates storage and display suggestions for the user based on structural information.

[0163] "Audio and visual information" refers to data including the user's voice and facial expressions, and is used as input data when analyzing emotional information.

[0164] "Emotional information" refers to data about the user's psychological state and emotions, extracted from audio and visual information.

[0165] An "emotion analysis device" is a component within a system that analyzes audio and visual information, generates emotional information, and adjusts suggestions to match the user's psychological state.

[0166] "Market information" refers to information about goods in online or offline trading markets and is used to select items related to the proposal.

[0167] A "display unit" is a functional unit that visually displays the generated proposals and provides an interface for receiving feedback from users.

[0168] "Unused areas" refer to spaces that are not normally utilized, as detected from visual data, and are used to propose storage methods.

[0169] A "purchase channel" is a link or means for purchasing selected items, provided in a format that is easily accessible to the user.

[0170] This invention is a system for providing optimal storage and display suggestions in a user's living space, offering personalized suggestions that take into account the user's emotional state.

[0171] Hardware and Software Overview

[0172] Users acquire visual data of their living space using imaging devices such as smartphones or tablets. This image data is transmitted to a server via a dedicated application.

[0173] The server runs in a cloud computing environment and analyzes visual data using image processing libraries (e.g., OpenCV). It identifies structural information such as room layout and furniture arrangement.

[0174] Furthermore, the server uses speech recognition and facial recognition algorithms (e.g., TensorFlow) to extract emotional information from the user's voice and visual information.

[0175] The generation AI model automatically generates storage and interior design suggestions tailored to the user, based on structural and emotional information.

[0176] The device visually presents these suggestions to the user and provides an intuitive interface.

[0177] Specific examples and prompt statements

[0178] For example, if a user takes a photo of their living room, the server analyzes the image to evaluate the efficiency of the furniture arrangement and suggests ways to utilize dead space. If the system determines that the user is in a relaxed emotional state, it will suggest soft lighting and plant placement. Purchase links for selected items are also provided, allowing the user to easily access the purchase page.

[0179] An example of a prompt message is: "Based on the user's living room image and emotional data, create three suggestions for creating a relaxing space and provide links to related products."

[0180] This invention allows users to make more effective use of their home space and create a comfortable environment optimized for their emotional state.

[0181] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0182] Step 1:

[0183] The user takes a picture of the room with their smartphone camera. The input data is high-resolution visual data, which enables accurate analysis of the environment. The captured image data is sent to a server via a dedicated app. The output is the transfer of the image data to the server.

[0184] Step 2:

[0185] The server stores the received visual data in a cloud environment. Next, it analyzes the data using an image processing library (e.g., OpenCV). The input is image data sent by the user. The server applies an object detection algorithm to identify the room layout and furniture arrangement. In this process, it extracts dead space and important structural information. The output is the structural information of the user's room.

[0186] Step 3:

[0187] The server extracts emotional information using speech and facial recognition algorithms. Input includes audio data and visual facial data. A machine learning model using TensorFlow is applied to analyze emotions from facial expressions and voice tone. The output is emotional data indicating the user's psychological state.

[0188] Step 4:

[0189] The server utilizes a generative AI model to combine structural and emotional information to generate storage and interior design suggestions. The input consists of analyzed structural and emotional information, which is used to automatically generate suggestions. The AI ​​model uses natural language processing to generate interior designs tailored to the user. The output is the proposed interior and storage strategy.

[0190] Step 5:

[0191] The server refers to a product information database and selects products related to the generated suggestions. The input is the suggestion content. Based on market information, it identifies purchase links for related products and selects products that match the user's emotional state. The output is the product links combined with the suggestions.

[0192] Step 6:

[0193] The device displays suggestions and selected products to the user through an intuitive interface. Input consists of the suggested items and product links. The visual UI allows users to easily browse suggestions and access the purchase page for selected products with a click. Output consists of the suggestions and product purchase links that the user sees.

[0194] Step 7:

[0195] Users provide feedback on the suggestions. This feedback is sent to the server via a dedicated application. The input is the user's feedback data. The output is the transmission of feedback information, which is used to train the sentiment engine and improve the quality of future suggestions.

[0196] (Application Example 2)

[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0198] In recent years, there has been a growing demand for technologies to improve the customer experience in physical stores. However, traditional methods fail to adequately provide personalized product recommendations that reflect customer emotions and preferences. Furthermore, conventional technologies struggle to respond immediately to changing customer needs.

[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0200] In this invention, the server includes means for receiving image data acquired by a camera and processing the image data to identify environmental information; a generation device for generating suggestions regarding storage and display based on the environmental information and emotion recognition; a selection device for selecting products related to the suggestions and providing means for purchasing those products by referring to market information; and a presentation device for providing product suggestions to customers using an external display device. This makes it possible to provide product suggestions optimized for customer emotions in physical stores and improve the customer experience.

[0201] A "photography device" is a device used to acquire image data and to photograph surrounding objects and space.

[0202] "Environmental information" refers to information about the room's structure, furniture arrangement, and spatial characteristics, identified based on image data acquired by the camera.

[0203] "Emotion recognition" is the process of determining a user's emotional state from their facial expressions and tone of voice, and acquiring that information.

[0204] A "generation device" is a system that creates storage and display suggestions based on received environmental information and the results of emotion recognition.

[0205] "Market information" refers to information related to consumer purchasing activities, including databases of products in circulation and sales information.

[0206] A "selection device" is a means of selecting products related to a proposal by referring to market information and providing users with opportunities to purchase them.

[0207] An "external display device" is a display device used to provide users with proposed information and information on selected products.

[0208] A specific system for carrying out this invention is configured as follows.

[0209] The server receives image data acquired from imaging devices such as smart glasses and head-mounted displays. The received image data is analyzed using image recognition algorithms such as TensorFlow and OpenCV to identify environmental information such as the structure of a room or store and the arrangement of items. Based on these analysis results, Amazon Rekognition or Microsoft® Azure® Face API are used as emotion engines to identify emotions from the user's facial expressions and tone of voice, and generate suggestions tailored to the current situation.

[0210] The store's customer service staff, who are the users, receive suggestion information from the server through external display devices such as smart glasses or HMDs. The server then references market information to select products optimized for the user's emotional state, and presents the results via the display device. The market information used includes the latest product databases and sales data.

[0211] This system works as follows: For example, if a staff member serving a customer in a store uses smart glasses to capture the customer's restless expression or tired voice, the system uses this information to suggest products that enhance relaxation, such as an aroma diffuser or a relaxation chair. This allows customers to receive recommendations for products that are best suited to their situation.

[0212] An example of a prompt sentence to input into a generative AI model is: "A store clerk wearing smart glasses observes customers in the store. If the customer appears relaxed, what products would the clerk suggest?"

[0213] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0214] Step 1:

[0215] The server receives image data captured through cameras on smart glasses or head-mounted displays. This image data is provided in real time from devices worn by store staff, who are the users of the system. The input images are treated as material necessary for analyzing environmental information.

[0216] Step 2:

[0217] The server analyzes the received image data using TensorFlow and OpenCV. The image recognition algorithm identifies environmental information such as the structure of the room or store and the arrangement of objects, which is then obtained as output. This output data serves as the basis for emotion recognition in the next step.

[0218] Step 3:

[0219] Based on identified environmental information, the server uses Amazon Rekognition and Microsoft Azure Face APIs to identify the user's emotions. Inputs include the user's facial expressions and tone of voice, while output provides the user's current emotional state. This emotion data is a crucial element in suggestion generation.

[0220] Step 4:

[0221] The server uses a generator to create optimal storage and display suggestions for the user based on acquired emotional data and environmental information. The inputs used are the analysis results and the results of emotional recognition, and the output is specific suggestions for the user. These suggestions are then used to reference market information.

[0222] Step 5:

[0223] The server references market information and selects products relevant to the requested suggestions. It uses the latest product database and sales information as input, and outputs a list of products optimized for the user's emotional state. The selected products are then processed by a selection device that provides purchasing options.

[0224] Step 6:

[0225] Product suggestions provided by a server are displayed on the terminal, such as smart glasses or a head-mounted display. Users can make real-time product suggestions to customers through the display device. This output information becomes a direct purchase suggestion to the customer, making in-store customer service activities more effective.

[0226] Through these steps, the system enables the real-time product recommendations in physical stores that take customer emotions into consideration, thereby improving the customer experience.

[0227] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0228] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0229] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0230] [Second Embodiment]

[0231] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0232] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0233] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0234] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0235] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0236] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0237] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0238] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0239] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0240] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0241] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0242] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0243] This invention provides a storage and display suggestion system using image analysis technology and AI to enable users to easily optimize their rooms at home. This system is generally accessible and easy to use via smartphones and tablets, and is implemented in the following form.

[0244] First, the user takes a photo of a specific room in their home using a mobile device such as a smartphone. It is recommended that the photo be taken in a way that shows the overall layout of the room. The captured image data is sent to the server via the application.

[0245] The server executes a deep learning-based image recognition algorithm to analyze the received image data. This identifies environmental information such as the room layout, furniture types and placement, and space usage. In particular, the image analysis can detect dead spaces that are not being used efficiently.

[0246] Next, a generator built into the server generates suggestions for improving storage and display based on environmental information. These suggestions include new storage methods utilizing dead space, furniture rearrangement, and harmonious display methods. Furthermore, the suggested ideas are optimized for the detected dead space.

[0247] Furthermore, the server accesses a market information database and selects appropriate furniture and storage items related to the suggestion. This selection is based on the user's needs and the condition of the room, ensuring the most suitable choice. The selected items are associated with online store purchase links, allowing the user to easily proceed with the purchase process.

[0248] The terminal displays suggestions and product selection information received from the server to the user. The user interface consists of easy-to-understand graphics and text, allowing users to easily understand the suggestions and use purchase links as needed.

[0249] For example, if a user takes a photo of their living room, the server will identify the placement of the sofa and TV stand and suggest a rack that can be installed in the corner space. A link to an online store will also be provided, allowing the user to purchase the recommended rack and secure new storage space. In this way, the present invention helps to easily and efficiently improve the user's living environment.

[0250] The following describes the processing flow.

[0251] Step 1:

[0252] The user launches a dedicated app on their smartphone and takes a photo that captures the entire room. The app then provides shooting guidelines to help the user take the best possible image.

[0253] Step 2:

[0254] The device compresses the captured image data within the app and sends it to the server via the network. During this process, an appropriate data transfer method is selected, taking into account the user's network bandwidth.

[0255] Step 3:

[0256] The server analyzes the received image data. Using a deep learning model, it identifies the room layout, furniture types, and their placement, and stores this information in a database as environmental data.

[0257] Step 4:

[0258] Based on the analyzed environmental information, the server generates optimal storage and display suggestions using a generation device. These suggestions include specific ideas for utilizing identified dead spaces.

[0259] Step 5:

[0260] The server selects furniture and storage items related to the generated proposal by referring to a market information database. Purchase links to online stores are added to the selected products.

[0261] Step 6:

[0262] The server formats the proposed content and selected product information and sends it to the user's terminal. This data is structured in a format suitable for the user interface.

[0263] Step 7:

[0264] The device visually presents the received data to the user. The app screen clearly displays suggested storage ideas and purchase links, supporting the user's decision-making.

[0265] Step 8:

[0266] Users review the displayed suggestions and purchase information and click on product links that interest them. This directs them to the online store, where they can proceed with the purchase.

[0267] (Example 1)

[0268] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0269] Conventional storage and display suggestion systems made it difficult for users to obtain concrete guidance on how to efficiently utilize their home space. Furthermore, the selection and purchase of items using market information was complex and time-consuming for users. Additionally, the inability to incorporate feedback on the suggested content made it difficult to improve the system's accuracy.

[0270] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0271] In this invention, the server includes a device that receives image information acquired by a shooting function and analyzes the image information to identify spatial information; a generation device that creates suggestions regarding storage and display based on the spatial information; and a selection device that refers to market information and provides the selection of items related to the suggestions and links to purchase those items. As a result, users can receive specific suggestions for effectively utilizing their home space and make quick and appropriate item selections and purchases based on market information. It is also possible to improve the accuracy of the system by utilizing feedback from users.

[0272] The "shooting function" is a feature that allows users to acquire visual information of their environment using their own devices.

[0273] "Image information" refers to visual data acquired through the camera's capture function, and forms the basis for analysis.

[0274] "Spatial information" refers to information about the physical structure and arrangement of the environment obtained from analyzed image information.

[0275] A "generation device" is a device that has the function of creating proposals for efficient storage and display based on spatial information.

[0276] "Market information" refers to data related to the provision of goods and services, and is used to select items related to the proposed content.

[0277] A "selection device" is a device that, based on market information, selects items that match the proposed content and provides the user with a purchase link.

[0278] A "display device" is a device used to visually present selected items to the user, encouraging them to confirm the details of the proposal.

[0279] "Unused space" refers to space that is not currently being effectively utilized, and is a target for proposals regarding new storage methods and furniture arrangements.

[0280] "Return information" refers to the opinions and evaluations made by users regarding the proposed content and is utilized as feedback for system improvement.

[0281] This invention uses a camera function, a server, a terminal, and a generative AI model for users to efficiently optimize the space in their homes. The following shows specific embodiments of this system.

[0282] Users use a terminal such as a smartphone or tablet to take pictures of specific rooms in their homes. The captured image information is automatically sent to the server through an application. Through this process, data is collected on the server in real time. <*

[0283] The server utilizes deep learning technology to analyze the received image information. In particular, it uses Convolutional Neural Network (CNN) and others to perform object recognition and spatial analysis within the image and identify spatial information. As a result, details such as dead space and furniture arrangement are grasped.

[0284] Next, the generation device incorporated in the server uses the generative AI model based on the analyzed spatial information to create proposals for storage and display. This model has been pre-trained and provides efficient proposals for various spatial layouts. The generated proposals include ideas for making effective use of unused space.

[0285] Furthermore, the server refers to market information and selects the optimal items related to the proposed improvement content. The selection device lists up the compatible items and associates each with a purchase link to an online store. Thereby, users can easily purchase the products.

[0286] [[ID=*25]]On the terminal, the received proposal content and item information are graphically displayed. The user interface is designed to enable visual confirmation of the proposals, and users can easily understand and operate the proposals.

[0287] As a specific example, when a user takes a photo of a living room, the server recognizes the arrangement of the sofa and table and identifies the unused space that is not being utilized. The generation device proposes the optimal storage method for this space, searches for suitable furniture online, and provides a purchase link. An example of the prompt sentence is, "Analyze the photo of the living room and propose an optimization plan for the storage space."

[0288] With this invention, users can utilize the space in their daily lives more efficiently, significantly improving convenience.

[0289] The flow of the specific process in Example 1 will be described using FIG. 11.

[0290] Step 1:

[0291] The user uses a smartphone or tablet to take a photo of a room in their home. It is recommended to take the photo in a composition that allows the layout of the entire room to be grasped. This image is automatically transmitted to the server through the application. The input is the image data taken by the user, and the output is the transfer of the image data to the server.

[0292] Step 2:

[0293] The server uses deep learning technology to analyze the received image data. Specifically, it uses a Convolutional Neural Network (CNN) to identify objects in the image and specify spatial information such as the floor plan, furniture arrangement, color, and shape. The input is the image data transmitted from the user, and the output is the detailed spatial information obtained through image analysis.

[0294] Step 3:

[0295] The server uses a generative AI model based on acquired spatial information to create storage and display suggestions. This generative AI model is trained to generate efficient storage methods and furniture arrangement patterns. The input is analyzed spatial information, and the output is specific improvement suggestions. This process includes ideas for utilizing underutilized spaces.

[0296] Step 4:

[0297] The server automatically selects suitable items by referencing market information related to the generated proposal. This selection considers factors such as item size, design, and price. The selected items are also accompanied by online store purchase links. The input is the generated proposal and market information, and the output is a list of items aligned with the proposal and their purchase links.

[0298] Step 5:

[0299] The terminal displays the suggested content and item information received from the server via a user interface. This interface is designed to be visually intuitive, allowing users to verify suggested storage methods and furniture arrangements using 3D models. Input consists of suggestions and item information from the server, while output is a visual display and information provided to the user.

[0300] (Application Example 1)

[0301] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0302] In modern brick-and-mortar stores, product display and space utilization are largely based on experience, making automation and the use of information technology difficult. In particular, store staff need considerable time and effort to visually analyze space and determine the optimal display. Furthermore, while new display strategies are constantly needed to enhance customer purchasing intent, obtaining objective and real-time optimization suggestions remains challenging.

[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means respectively.

[0304] In this invention, the server includes means for receiving image information acquired by imaging equipment, processing the image information to identify the spatial situation, generation means for generating proposals regarding storage and display based on the spatial situation, and selection means for selecting products related to the proposals by referring to market information and providing purchase links for the products. Thereby, through the wearable information display device, it becomes possible to present in real time the efficient utilization of space and the optimization of product display in physical stores.

[0305] The "imaging equipment" is a device for acquiring image information of a specific object.

[0306] The "image information" is data acquired by imaging equipment, which visually shows the spatial situation and the arrangement of objects.

[0307] The "spatial situation" is information indicating the physical environment and the arrangement state of objects at a specific location.

[0308] The "generation means" is a function for creating proposals regarding storage and display based on the identified spatial situation.

[0309] The "market information" refers to data related to the market, such as product types, prices, distribution information, etc.

[0310] The "selection means" is a function for selecting the optimal products based on market information and presenting relevant information to users.

[0311] The "wearable information display device" is a device for visually presenting information when worn by a user.

[0312] The "unused space" refers to space that is not physically utilized or not efficiently utilized.

[0313] "User feedback" refers to information that includes opinions and evaluations from users regarding a proposal.

[0314] A "learning tool" is a function that uses feedback information to improve the system's performance and the accuracy of its suggestions.

[0315] To implement this invention, the following system configuration can be used. The server receives image information transmitted from the imaging device and runs a deep learning model using image recognition technology to analyze it. This analysis identifies spatial conditions and wasted space. For example, TensorFlow, an open-source deep learning library, can be used.

[0316] Once the analysis is complete, the server generates storage and display suggestions based on the results using a generation mechanism. These suggestions include efficient product display and ways to utilize wasted space. The generated suggestions are displayed in real time on a wearable information display device, such as smart glasses.

[0317] Next, the server references market information, selects products suitable for the user using selection methods, and provides relevant information and purchase links. This process involves gathering information from online databases.

[0318] Users can view information presented through smart glasses and intuitively adjust displays in the actual store space. When a specific suggestion is selected, user feedback is accumulated in the system, and suggestions are improved through learning mechanisms.

[0319] As a concrete example, when a staff member at a sporting goods store uses smart glasses to take pictures of the store, a server analyzes the wasted space on the shelves. Subsequently, suggestions for displaying new basketballs in the empty spaces are displayed, along with a purchase link for those products. This is expected to improve the efficiency of product placement and enhance customer satisfaction.

[0320] An example of a prompt using a generative AI model would be: "Analyze the space in the input image and suggest an efficient way to display products. For example, show what kind of products should be displayed in the empty space in the corner."

[0321] Such a system can enable efficient use of space and improved customer experience, especially in physical stores.

[0322] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0323] Step 1:

[0324] The user wears smart glasses and walks around the store, capturing image data of designated areas. This image data is then used as input for the system.

[0325] Step 2:

[0326] The terminal sends the acquired image information to the server. The input data is sent in image formats such as JPEG and PNG. The server receives this image information and uses a deep learning model to analyze the spatial situation. In this analysis, an object recognition algorithm identifies wasted space and existing display arrangements. As a result of the analysis, information on wasted space is output.

[0327] Step 3:

[0328] The server uses an AI model generated based on the analysis results to produce suggestions for storage and display that utilize wasted space. The input data includes spatial information from the analysis results and information on existing merchandise. The generated suggestions become the server's output, presenting optimal product placement and display methods.

[0329] Step 4:

[0330] The server uses the generated suggestions as a reference, queries relevant market information, and selects products using selection criteria. The input data is the generated suggestions, and the output is information on products that match the suggestions and purchase links. This information is obtained from an online database.

[0331] Step 5:

[0332] The terminal displays the suggested content and selected product information received from the server on the user's wearable information display device. The input is display data generated by the server, and the output is display information visually presented on smart glasses. This allows the user to confirm how to use the space while looking at the actual wasted space.

[0333] Step 6:

[0334] Users can implement the proposed installation method and test the results. They can also provide feedback on the usefulness of the proposal and the results of adopting it. This feedback is recorded as input to the system and used by the server's learning mechanism to improve the accuracy of future proposals.

[0335] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0336] This invention combines a system that analyzes image data related to the user's room layout and proposes optimal storage and display solutions with an emotion engine that recognizes the user's emotions and reflects them in the suggestions. As a result, users can receive suggestions optimized for their own emotional state, enabling them to create a more satisfying living space.

[0337] First, the user takes a picture of a specific room in their home using their smartphone camera. This image data is then sent to a server via a dedicated app.

[0338] The server executes an image recognition algorithm to analyze the received image data, identifying environmental information such as room layout, furniture arrangement, and dead space. This analysis builds the foundational data for proposals that enable efficient use of space.

[0339] Next, the server activates an emotion engine to recognize the user's emotions from their facial expressions and tone of voice during the photo shoot. This emotion data is used to generate suggestions that reflect the user's current mood and preferences. For example, if the user prefers a relaxed atmosphere, the server might suggest soft-colored curtains or the placement of houseplants.

[0340] Furthermore, the server references market information and selects products based on sentiment recognition results. This selection includes furniture and interior goods that reflect the user's emotions and preferences. Purchase links for the selected products are provided so that users can access them immediately if they are interested.

[0341] The device displays these suggestions to the user in an easy-to-understand interface. Users can browse the displayed suggestions and products and make selections that match their emotional state. They can also provide feedback on the suggestions, and this feedback is used to improve the accuracy of future suggestions through the emotion engine's learning function.

[0342] As a concrete example, if a user takes a photo of their living room, the server analyzes the image to identify the placement of sofas and tables, and uses an emotion engine to generate suggestions for creating a relaxing space. Furthermore, purchase links for products related to the suggested interior are provided, allowing the user to access the purchase page with a single click. In this way, the present invention, which combines an emotion engine, supports the creation of a more personalized space by providing customized suggestions that take the user's emotions into consideration.

[0343] The following describes the processing flow.

[0344] Step 1:

[0345] The user takes photos of their room at home using a dedicated app. During this process, the app activates the camera to detect the user's facial expressions and records their emotions.

[0346] Step 2:

[0347] The device sends captured image data and user facial expression data to the server. This data is treated as basic information for analysis.

[0348] Step 3:

[0349] The server analyzes the received image data and uses an AI model to identify the room layout and furniture arrangement. At the same time, it detects dead space and registers it in a database.

[0350] Step 4:

[0351] The server activates an emotion engine to recognize emotions from the user's facial expression data. This emotion data is used to adjust the suggestions.

[0352] Step 5:

[0353] The server generates optimal storage and display suggestions based on environmental information and emotional data. The generation device then performs this process, including color coordination and item selection tailored to the emotional state.

[0354] Step 6:

[0355] Based on the proposal, the server selects relevant products from the market information database. Purchase links are then set up for the selected products to allow for easy access.

[0356] Step 7:

[0357] The server sends the suggested content and product information to the terminal. The data is formatted in a way that is easy for the user to understand intuitively.

[0358] Step 8:

[0359] The device displays suggestions through its user interface. Users can view the presented information and click on product links that interest them.

[0360] Step 9:

[0361] Users access external online stores through selected links and complete purchases. They can also provide feedback on the suggestions they receive within the app.

[0362] (Example 2)

[0363] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0364] In modern living spaces, users often face the challenge of efficiently utilizing limited space while creating an environment that is optimal for their individual emotions and preferences. Furthermore, selecting the right products from the many options available on the market is not easy. Therefore, there is a need for a system that automatically provides suggestions that match the user's emotional state and preferences, thereby improving the efficiency of storage and interior design selection.

[0365] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0366] In this invention, the server includes means for receiving visual data acquired by an imaging device and analyzing the visual data to identify structural information; a generation device for creating storage and display suggestions based on the structural information; and an emotion analysis device for analyzing emotional information from audio and visual information and adjusting suggestions based on the user's psychological state. This makes it possible to efficiently utilize limited space and provide personalized suggestions that match the user's emotions and preferences.

[0367] An "imaging device" is a hardware device used to acquire visual data and is used to capture images of the user's environment.

[0368] "Visual data" refers to image information acquired by an imaging device, and serves as the basis for analyzing the structure and layout of a room.

[0369] "Structural information" refers to information about the room layout and furniture arrangement, identified through the analysis of visual data.

[0370] A "generation device" is a functional unit within a system that generates storage and display suggestions for the user based on structural information.

[0371] "Audio and visual information" refers to data including the user's voice and facial expressions, and is used as input data when analyzing emotional information.

[0372] "Emotional information" refers to data about the user's psychological state and emotions, extracted from audio and visual information.

[0373] An "emotion analysis device" is a component within a system that analyzes audio and visual information, generates emotional information, and adjusts suggestions to match the user's psychological state.

[0374] "Market information" refers to information about goods in online or offline trading markets and is used to select items related to the proposal.

[0375] A "display unit" is a functional unit that visually displays the generated proposals and provides an interface for receiving feedback from users.

[0376] "Unused areas" refer to spaces that are not normally utilized, as detected from visual data, and are used to propose storage methods.

[0377] A "purchase channel" is a link or means for purchasing selected items, provided in a format that is easily accessible to the user.

[0378] This invention is a system for providing optimal storage and display suggestions in a user's living space, offering personalized suggestions that take into account the user's emotional state.

[0379] Hardware and Software Overview

[0380] Users acquire visual data of their living space using imaging devices such as smartphones or tablets. This image data is transmitted to a server via a dedicated application.

[0381] The server runs in a cloud computing environment and analyzes visual data using image processing libraries (e.g., OpenCV). It identifies structural information such as room layout and furniture arrangement.

[0382] Furthermore, the server uses speech recognition and facial recognition algorithms (e.g., TensorFlow) to extract emotional information from the user's voice and visual information.

[0383] The generation AI model automatically generates storage and interior design suggestions tailored to the user, based on structural and emotional information.

[0384] The device visually presents these suggestions to the user and provides an intuitive interface.

[0385] Specific examples and prompt statements

[0386] For example, if a user takes a photo of their living room, the server analyzes the image to evaluate the efficiency of the furniture arrangement and suggests ways to utilize dead space. If the system determines that the user is in a relaxed emotional state, it will suggest soft lighting and plant placement. Purchase links for selected items are also provided, allowing the user to easily access the purchase page.

[0387] An example of a prompt message is: "Based on the user's living room image and emotional data, create three suggestions for creating a relaxing space and provide links to related products."

[0388] This invention allows users to make more effective use of their home space and create a comfortable environment optimized for their emotional state.

[0389] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0390] Step 1:

[0391] The user takes a picture of the room with their smartphone camera. The input data is high-resolution visual data, which enables accurate analysis of the environment. The captured image data is sent to a server via a dedicated app. The output is the transfer of the image data to the server.

[0392] Step 2:

[0393] The server stores the received visual data in a cloud environment. Next, it analyzes the data using an image processing library (e.g., OpenCV). The input is image data sent by the user. The server applies an object detection algorithm to identify the room layout and furniture arrangement. In this process, it extracts dead space and important structural information. The output is the structural information of the user's room.

[0394] Step 3:

[0395] The server extracts emotional information using speech and facial recognition algorithms. Input includes audio data and visual facial data. A machine learning model using TensorFlow is applied to analyze emotions from facial expressions and voice tone. The output is emotional data indicating the user's psychological state.

[0396] Step 4:

[0397] The server utilizes a generative AI model to combine structural and emotional information to generate storage and interior design suggestions. The input consists of analyzed structural and emotional information, which is used to automatically generate suggestions. The AI ​​model uses natural language processing to generate interior designs tailored to the user. The output is the proposed interior and storage strategy.

[0398] Step 5:

[0399] The server refers to a product information database and selects products related to the generated suggestions. The input is the suggestion content. Based on market information, it identifies purchase links for related products and selects products that match the user's emotional state. The output is the product links combined with the suggestions.

[0400] Step 6:

[0401] The device displays suggestions and selected products to the user through an intuitive interface. Input consists of the suggested items and product links. The visual UI allows users to easily browse suggestions and access the purchase page for selected products with a click. Output consists of the suggestions and product purchase links that the user sees.

[0402] Step 7:

[0403] Users provide feedback on the suggestions. This feedback is sent to the server via a dedicated application. The input is the user's feedback data. The output is the transmission of feedback information, which is used to train the sentiment engine and improve the quality of future suggestions.

[0404] (Application Example 2)

[0405] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0406] In recent years, there has been a growing demand for technologies to improve the customer experience in physical stores. However, traditional methods fail to adequately provide personalized product recommendations that reflect customer emotions and preferences. Furthermore, conventional technologies struggle to respond immediately to changing customer needs.

[0407] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0408] In this invention, the server includes means for receiving image data acquired by a camera and processing the image data to identify environmental information; a generation device for generating suggestions regarding storage and display based on the environmental information and emotion recognition; a selection device for selecting products related to the suggestions and providing means for purchasing those products by referring to market information; and a presentation device for providing product suggestions to customers using an external display device. This makes it possible to provide product suggestions optimized for customer emotions in physical stores and improve the customer experience.

[0409] A "photography device" is a device used to acquire image data and to photograph surrounding objects and space.

[0410] "Environmental information" refers to information about the room's structure, furniture arrangement, and spatial characteristics, identified based on image data acquired by the camera.

[0411] "Emotion recognition" is the process of determining a user's emotional state from their facial expressions and tone of voice, and acquiring that information.

[0412] A "generation device" is a system that creates storage and display suggestions based on received environmental information and the results of emotion recognition.

[0413] "Market information" refers to information related to consumer purchasing activities, including databases of products in circulation and sales information.

[0414] A "selection device" is a means of selecting products related to a proposal by referring to market information and providing users with opportunities to purchase them.

[0415] An "external display device" is a display device used to provide users with proposed information and information on selected products.

[0416] A specific system for carrying out this invention is configured as follows.

[0417] The server receives image data acquired from imaging devices such as smart glasses and head-mounted displays. The received image data is analyzed using image recognition algorithms such as TensorFlow and OpenCV to identify environmental information such as the structure of a room or store and the arrangement of items. Based on these analysis results, Amazon Rekognition or Microsoft Azure Face API are used as emotion engines to identify emotions from the user's facial expressions and tone of voice, and generate suggestions tailored to the current situation.

[0418] The store's customer service staff, who are the users, receive suggestion information from the server through external display devices such as smart glasses or HMDs. The server then references market information to select products optimized for the user's emotional state, and presents the results via the display device. The market information used includes the latest product databases and sales data.

[0419] This system works as follows: For example, if a staff member serving a customer in a store uses smart glasses to capture the customer's restless expression or tired voice, the system uses this information to suggest products that enhance relaxation, such as an aroma diffuser or a relaxation chair. This allows customers to receive recommendations for products that are best suited to their situation.

[0420] An example of a prompt sentence to input into a generative AI model is: "A store clerk wearing smart glasses observes customers in the store. If the customer appears relaxed, what products would the clerk suggest?"

[0421] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0422] Step 1:

[0423] The server receives image data captured through cameras on smart glasses or head-mounted displays. This image data is provided in real time from devices worn by store staff, who are the users of the system. The input images are treated as material necessary for analyzing environmental information.

[0424] Step 2:

[0425] The server analyzes the received image data using TensorFlow and OpenCV. The image recognition algorithm identifies environmental information such as the structure of the room or store and the arrangement of objects, which is then obtained as output. This output data serves as the basis for emotion recognition in the next step.

[0426] Step 3:

[0427] Based on identified environmental information, the server uses Amazon Rekognition and Microsoft Azure Face APIs to identify the user's emotions. Inputs include the user's facial expressions and tone of voice, while output provides the user's current emotional state. This emotion data is a crucial element in suggestion generation.

[0428] Step 4:

[0429] The server uses a generator to create optimal storage and display suggestions for the user based on acquired emotional data and environmental information. The inputs used are the analysis results and the results of emotional recognition, and the output is specific suggestions for the user. These suggestions are then used to reference market information.

[0430] Step 5:

[0431] The server references market information and selects products relevant to the requested suggestions. It uses the latest product database and sales information as input, and outputs a list of products optimized for the user's emotional state. The selected products are then processed by a selection device that provides purchasing options.

[0432] Step 6:

[0433] Product suggestions provided by a server are displayed on the terminal, such as smart glasses or a head-mounted display. Users can make real-time product suggestions to customers through the display device. This output information becomes a direct purchase suggestion to the customer, making in-store customer service activities more effective.

[0434] Through these steps, the system enables the real-time product recommendations in physical stores that take customer emotions into consideration, thereby improving the customer experience.

[0435] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0436] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0437] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0438] [Third Embodiment]

[0439] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0440] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0441] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0442] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0443] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0445] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0446] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0447] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0448] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0449] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0450] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0451] This invention provides a storage and display suggestion system using image analysis technology and AI to enable users to easily optimize their rooms at home. This system is generally accessible and easy to use via smartphones and tablets, and is implemented in the following form.

[0452] First, the user takes a photo of a specific room in their home using a mobile device such as a smartphone. It is recommended that the photo be taken in a way that shows the overall layout of the room. The captured image data is sent to the server via the application.

[0453] The server executes a deep learning-based image recognition algorithm to analyze the received image data. This identifies environmental information such as the room layout, furniture types and placement, and space usage. In particular, the image analysis can detect dead spaces that are not being used efficiently.

[0454] Next, a generator built into the server generates suggestions for improving storage and display based on environmental information. These suggestions include new storage methods utilizing dead space, furniture rearrangement, and harmonious display methods. Furthermore, the suggested ideas are optimized for the detected dead space.

[0455] Furthermore, the server accesses a market information database and selects appropriate furniture and storage items related to the suggestion. This selection is based on the user's needs and the condition of the room, ensuring the most suitable choice. The selected items are associated with online store purchase links, allowing the user to easily proceed with the purchase process.

[0456] The terminal displays suggestions and product selection information received from the server to the user. The user interface consists of easy-to-understand graphics and text, allowing users to easily understand the suggestions and use purchase links as needed.

[0457] For example, if a user takes a photo of their living room, the server will identify the placement of the sofa and TV stand and suggest a rack that can be installed in the corner space. A link to an online store will also be provided, allowing the user to purchase the recommended rack and secure new storage space. In this way, the present invention helps to easily and efficiently improve the user's living environment.

[0458] The following describes the processing flow.

[0459] Step 1:

[0460] The user launches a dedicated app on their smartphone and takes a photo that captures the entire room. The app then provides shooting guidelines to help the user take the best possible image.

[0461] Step 2:

[0462] The device compresses the captured image data within the app and sends it to the server via the network. During this process, an appropriate data transfer method is selected, taking into account the user's network bandwidth.

[0463] Step 3:

[0464] The server analyzes the received image data. Using a deep learning model, it identifies the room layout, furniture types, and their placement, and stores this information in a database as environmental data.

[0465] Step 4:

[0466] Based on the analyzed environmental information, the server generates optimal storage and display suggestions using a generation device. These suggestions include specific ideas for utilizing identified dead spaces.

[0467] Step 5:

[0468] The server selects furniture and storage items related to the generated proposal by referring to a market information database. Purchase links to online stores are added to the selected products.

[0469] Step 6:

[0470] The server formats the proposed content and selected product information and sends it to the user's terminal. This data is structured in a format suitable for the user interface.

[0471] Step 7:

[0472] The device visually presents the received data to the user. The app screen clearly displays suggested storage ideas and purchase links, supporting the user's decision-making.

[0473] Step 8:

[0474] Users review the displayed suggestions and purchase information and click on product links that interest them. This directs them to the online store, where they can proceed with the purchase.

[0475] (Example 1)

[0476] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0477] Conventional storage and display suggestion systems made it difficult for users to obtain concrete guidance on how to efficiently utilize their home space. Furthermore, the selection and purchase of items using market information was complex and time-consuming for users. Additionally, the inability to incorporate feedback on the suggested content made it difficult to improve the system's accuracy.

[0478] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0479] In this invention, the server includes a device that receives image information acquired by a shooting function and analyzes the image information to identify spatial information; a generation device that creates suggestions regarding storage and display based on the spatial information; and a selection device that refers to market information and provides the selection of items related to the suggestions and links to purchase those items. As a result, users can receive specific suggestions for effectively utilizing their home space and make quick and appropriate item selections and purchases based on market information. It is also possible to improve the accuracy of the system by utilizing feedback from users.

[0480] The "shooting function" is a feature that allows users to acquire visual information of their environment using their own devices.

[0481] "Image information" refers to visual data acquired through the camera's capture function, and forms the basis for analysis.

[0482] "Spatial information" refers to information about the physical structure and arrangement of the environment obtained from analyzed image information.

[0483] A "generation device" is a device that has the function of creating proposals for efficient storage and display based on spatial information.

[0484] "Market information" refers to data related to the provision of goods and services, and is used to select items related to the proposed content.

[0485] A "selection device" is a device that, based on market information, selects items that match the proposed content and provides the user with a purchase link.

[0486] A "display device" is a device used to visually present selected items to the user, encouraging them to confirm the details of the proposal.

[0487] "Unused space" refers to space that is not currently being effectively utilized, and is a target for proposals regarding new storage methods and furniture arrangements.

[0488] "Feedback information" refers to the opinions and evaluations that users provide regarding the proposed content, and is used as feedback for system improvement.

[0489] This invention utilizes a shooting function, server, terminal, and generating AI model to enable users to efficiently optimize their home space. Specific embodiments of this system are described below.

[0490] Users use devices such as smartphones or tablets to take pictures of specific rooms in their homes. The captured image information is automatically sent to a server via the application. Through this process, data is collected on the server in real time.

[0491] The server utilizes deep learning technology to analyze the received image information. In particular, it uses Convolutional Neural Networks (CNNs) and other techniques to perform object recognition and spatial analysis within the image, thereby identifying spatial information. This allows for the understanding of details such as dead space and furniture placement.

[0492] Next, a generation device integrated into the server uses a generation AI model based on the analyzed spatial information to create storage and display proposals. This model is pre-trained and provides efficient suggestions for diverse spatial layouts. The generated proposals include ideas for effectively utilizing underutilized space.

[0493] Furthermore, the server references market information and selects the most suitable items related to the proposed improvements. The selection device lists the suitable items and associates each with a purchase link to an online store. This allows users to easily purchase the products.

[0494] The terminal graphically displays the received proposals and item information. The user interface is designed to allow users to visually review the proposals, making them easy to understand and act upon.

[0495] For example, when a user takes a photo of their living room, the server recognizes the arrangement of sofas and tables and identifies unused space. The generator then suggests the best storage solution for this space, finds suitable furniture online, and provides purchase links. An example of a prompt message is, "Analyze the photo of my living room and suggest an optimization for storage space."

[0496] This invention enables users to utilize their daily living spaces more efficiently, significantly improving convenience.

[0497] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0498] Step 1:

[0499] Users take photos of their rooms at home using their smartphones or tablets. It is recommended that the photos be composed in a way that shows the overall layout of the room. These images are automatically sent to the server via the application. The input is the image data taken by the user, and the output is the transfer of the image data to the server.

[0500] Step 2:

[0501] The server uses deep learning technology to analyze the received image data. Specifically, it uses a Convolutional Neural Network (CNN) to identify objects within the image and determine spatial information such as floor plan, furniture arrangement, color, and shape. The input is image data sent by the user, and the output is detailed spatial information obtained through image analysis.

[0502] Step 3:

[0503] The server uses a generative AI model based on acquired spatial information to create storage and display suggestions. This generative AI model is trained to generate efficient storage methods and furniture arrangement patterns. The input is analyzed spatial information, and the output is specific improvement suggestions. This process includes ideas for utilizing underutilized spaces.

[0504] Step 4:

[0505] The server automatically selects suitable items by referencing market information related to the generated proposal. This selection considers factors such as item size, design, and price. The selected items are also accompanied by online store purchase links. The input is the generated proposal and market information, and the output is a list of items aligned with the proposal and their purchase links.

[0506] Step 5:

[0507] The terminal displays the suggested content and item information received from the server via a user interface. This interface is designed to be visually intuitive, allowing users to verify suggested storage methods and furniture arrangements using 3D models. Input consists of suggestions and item information from the server, while output is a visual display and information provided to the user.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] In modern brick-and-mortar stores, product display and space utilization are largely based on experience, making automation and the use of information technology difficult. In particular, store staff need considerable time and effort to visually analyze space and determine the optimal display. Furthermore, while new display strategies are constantly needed to enhance customer purchasing intent, obtaining objective and real-time optimization suggestions remains challenging.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes means for receiving image information acquired by a camera and processing the image information to identify the spatial situation; generation means for generating suggestions regarding storage and display based on the spatial situation; and selection means for selecting products related to the suggestions by referring to market information and providing purchase links for those products. This makes it possible to present efficient use of space and optimization of product display in physical stores in real time through a wearable information display device.

[0513] "Photography equipment" refers to a device used to acquire image information of a specific subject.

[0514] "Image information" refers to data acquired by photographic equipment, which visually represents the spatial situation and the arrangement of objects.

[0515] "Spatial conditions" refer to information that describes the physical environment and the arrangement of objects in a specific location.

[0516] "Generating means" refers to a function for creating proposals regarding storage and display based on a specified spatial situation.

[0517] "Market information" refers to data related to the market, such as product types, prices, and distribution information.

[0518] "Selection method" refers to a function that selects the most suitable product based on market information and presents relevant information to the user.

[0519] A "wearable information display device" is a device that is worn by a user to visually present information.

[0520] "Wasted space" refers to space that is not physically used or is not being utilized efficiently.

[0521] "User feedback" refers to information that includes opinions and evaluations from users regarding a proposal.

[0522] A "learning tool" is a function that uses feedback information to improve the system's performance and the accuracy of its suggestions.

[0523] To implement this invention, the following system configuration can be used. The server receives image information transmitted from the imaging device and runs a deep learning model using image recognition technology to analyze it. This analysis identifies spatial conditions and wasted space. For example, TensorFlow, an open-source deep learning library, can be used.

[0524] Once the analysis is complete, the server generates storage and display suggestions based on the results using a generation mechanism. These suggestions include efficient product display and ways to utilize wasted space. The generated suggestions are displayed in real time on a wearable information display device, such as smart glasses.

[0525] Next, the server references market information, selects products suitable for the user using selection methods, and provides relevant information and purchase links. This process involves gathering information from online databases.

[0526] Users can view information presented through smart glasses and intuitively adjust displays in the actual store space. When a specific suggestion is selected, user feedback is accumulated in the system, and suggestions are improved through learning mechanisms.

[0527] As a concrete example, when a staff member at a sporting goods store uses smart glasses to take pictures of the store, a server analyzes the wasted space on the shelves. Subsequently, suggestions for displaying new basketballs in the empty spaces are displayed, along with a purchase link for those products. This is expected to improve the efficiency of product placement and enhance customer satisfaction.

[0528] An example of a prompt using a generative AI model would be: "Analyze the space in the input image and suggest an efficient way to display products. For example, show what kind of products should be displayed in the empty space in the corner."

[0529] Such a system can enable efficient use of space and improved customer experience, especially in physical stores.

[0530] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0531] Step 1:

[0532] The user wears smart glasses and walks around the store, capturing image data of designated areas. This image data is then used as input for the system.

[0533] Step 2:

[0534] The terminal sends the acquired image information to the server. The input data is sent in image formats such as JPEG and PNG. The server receives this image information and uses a deep learning model to analyze the spatial situation. In this analysis, an object recognition algorithm identifies wasted space and existing display arrangements. As a result of the analysis, information on wasted space is output.

[0535] Step 3:

[0536] The server uses an AI model generated based on the analysis results to produce suggestions for storage and display that utilize wasted space. The input data includes spatial information from the analysis results and information on existing merchandise. The generated suggestions become the server's output, presenting optimal product placement and display methods.

[0537] Step 4:

[0538] The server uses the generated suggestions as a reference, queries relevant market information, and selects products using selection criteria. The input data is the generated suggestions, and the output is information on products that match the suggestions and purchase links. This information is obtained from an online database.

[0539] Step 5:

[0540] The terminal displays the suggested content and selected product information received from the server on the user's wearable information display device. The input is display data generated by the server, and the output is display information visually presented on smart glasses. This allows the user to confirm how to use the space while looking at the actual wasted space.

[0541] Step 6:

[0542] Users can implement the proposed installation method and test the results. They can also provide feedback on the usefulness of the proposal and the results of adopting it. This feedback is recorded as input to the system and used by the server's learning mechanism to improve the accuracy of future proposals.

[0543] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0544] This invention combines a system that analyzes image data related to the user's room layout and proposes optimal storage and display solutions with an emotion engine that recognizes the user's emotions and reflects them in the suggestions. As a result, users can receive suggestions optimized for their own emotional state, enabling them to create a more satisfying living space.

[0545] First, the user takes a picture of a specific room in their home using their smartphone camera. This image data is then sent to a server via a dedicated app.

[0546] The server executes an image recognition algorithm to analyze the received image data, identifying environmental information such as room layout, furniture arrangement, and dead space. This analysis builds the foundational data for proposals that enable efficient use of space.

[0547] Next, the server activates an emotion engine to recognize the user's emotions from their facial expressions and tone of voice during the photo shoot. This emotion data is used to generate suggestions that reflect the user's current mood and preferences. For example, if the user prefers a relaxed atmosphere, the server might suggest soft-colored curtains or the placement of houseplants.

[0548] Furthermore, the server references market information and selects products based on sentiment recognition results. This selection includes furniture and interior goods that reflect the user's emotions and preferences. Purchase links for the selected products are provided so that users can access them immediately if they are interested.

[0549] The device displays these suggestions to the user in an easy-to-understand interface. Users can browse the displayed suggestions and products and make selections that match their emotional state. They can also provide feedback on the suggestions, and this feedback is used to improve the accuracy of future suggestions through the emotion engine's learning function.

[0550] As a concrete example, if a user takes a photo of their living room, the server analyzes the image to identify the placement of sofas and tables, and uses an emotion engine to generate suggestions for creating a relaxing space. Furthermore, purchase links for products related to the suggested interior are provided, allowing the user to access the purchase page with a single click. In this way, the present invention, which combines an emotion engine, supports the creation of a more personalized space by providing customized suggestions that take the user's emotions into consideration.

[0551] The following describes the processing flow.

[0552] Step 1:

[0553] The user takes photos of their room at home using a dedicated app. During this process, the app activates the camera to detect the user's facial expressions and records their emotions.

[0554] Step 2:

[0555] The device sends captured image data and user facial expression data to the server. This data is treated as basic information for analysis.

[0556] Step 3:

[0557] The server analyzes the received image data and uses an AI model to identify the room layout and furniture arrangement. At the same time, it detects dead space and registers it in a database.

[0558] Step 4:

[0559] The server activates an emotion engine to recognize emotions from the user's facial expression data. This emotion data is used to adjust the suggestions.

[0560] Step 5:

[0561] The server generates optimal storage and display suggestions based on environmental information and emotional data. The generation device then performs this process, including color coordination and item selection tailored to the emotional state.

[0562] Step 6:

[0563] Based on the proposal, the server selects relevant products from the market information database. Purchase links are then set up for the selected products to allow for easy access.

[0564] Step 7:

[0565] The server sends the suggested content and product information to the terminal. The data is formatted in a way that is easy for the user to understand intuitively.

[0566] Step 8:

[0567] The device displays suggestions through its user interface. Users can view the presented information and click on product links that interest them.

[0568] Step 9:

[0569] Users access external online stores through selected links and complete purchases. They can also provide feedback on the suggestions they receive within the app.

[0570] (Example 2)

[0571] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] In modern living spaces, users often face the challenge of efficiently utilizing limited space while creating an environment that is optimal for their individual emotions and preferences. Furthermore, selecting the right products from the many options available on the market is not easy. Therefore, there is a need for a system that automatically provides suggestions that match the user's emotional state and preferences, thereby improving the efficiency of storage and interior design selection.

[0573] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0574] In this invention, the server includes means for receiving visual data acquired by an imaging device and analyzing the visual data to identify structural information; a generation device for creating storage and display suggestions based on the structural information; and an emotion analysis device for analyzing emotional information from audio and visual information and adjusting suggestions based on the user's psychological state. This makes it possible to efficiently utilize limited space and provide personalized suggestions that match the user's emotions and preferences.

[0575] An "imaging device" is a hardware device used to acquire visual data and is used to capture images of the user's environment.

[0576] "Visual data" refers to image information acquired by an imaging device, and serves as the basis for analyzing the structure and layout of a room.

[0577] "Structural information" refers to information about the room layout and furniture arrangement, identified through the analysis of visual data.

[0578] A "generation device" is a functional unit within a system that generates storage and display suggestions for the user based on structural information.

[0579] "Audio and visual information" refers to data including the user's voice and facial expressions, and is used as input data when analyzing emotional information.

[0580] "Emotional information" refers to data about the user's psychological state and emotions, extracted from audio and visual information.

[0581] An "emotion analysis device" is a component within a system that analyzes audio and visual information, generates emotional information, and adjusts suggestions to match the user's psychological state.

[0582] "Market information" refers to information about goods in online or offline trading markets and is used to select items related to the proposal.

[0583] A "display unit" is a functional unit that visually displays the generated proposals and provides an interface for receiving feedback from users.

[0584] "Unused areas" refer to spaces that are not normally utilized, as detected from visual data, and are used to propose storage methods.

[0585] A "purchase channel" is a link or means for purchasing selected items, provided in a format that is easily accessible to the user.

[0586] This invention is a system for providing optimal storage and display suggestions in a user's living space, offering personalized suggestions that take into account the user's emotional state.

[0587] Hardware and Software Overview

[0588] Users acquire visual data of their living space using imaging devices such as smartphones or tablets. This image data is transmitted to a server via a dedicated application.

[0589] The server runs in a cloud computing environment and analyzes visual data using image processing libraries (e.g., OpenCV). It identifies structural information such as room layout and furniture arrangement.

[0590] Furthermore, the server uses speech recognition and facial recognition algorithms (e.g., TensorFlow) to extract emotional information from the user's voice and visual information.

[0591] The generation AI model automatically generates storage and interior design suggestions tailored to the user, based on structural and emotional information.

[0592] The device visually presents these suggestions to the user and provides an intuitive interface.

[0593] Specific examples and prompt statements

[0594] For example, if a user takes a photo of their living room, the server analyzes the image to evaluate the efficiency of the furniture arrangement and suggests ways to utilize dead space. If the system determines that the user is in a relaxed emotional state, it will suggest soft lighting and plant placement. Purchase links for selected items are also provided, allowing the user to easily access the purchase page.

[0595] An example of a prompt message is: "Based on the user's living room image and emotional data, create three suggestions for creating a relaxing space and provide links to related products."

[0596] This invention allows users to make more effective use of their home space and create a comfortable environment optimized for their emotional state.

[0597] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0598] Step 1:

[0599] The user takes a picture of the room with their smartphone camera. The input data is high-resolution visual data, which enables accurate analysis of the environment. The captured image data is sent to a server via a dedicated app. The output is the transfer of the image data to the server.

[0600] Step 2:

[0601] The server stores the received visual data in a cloud environment. Next, it analyzes the data using an image processing library (e.g., OpenCV). The input is image data sent by the user. The server applies an object detection algorithm to identify the room layout and furniture arrangement. In this process, it extracts dead space and important structural information. The output is the structural information of the user's room.

[0602] Step 3:

[0603] The server extracts emotional information using speech and facial recognition algorithms. Input includes audio data and visual facial data. A machine learning model using TensorFlow is applied to analyze emotions from facial expressions and voice tone. The output is emotional data indicating the user's psychological state.

[0604] Step 4:

[0605] The server utilizes a generative AI model to combine structural and emotional information to generate storage and interior design suggestions. The input consists of analyzed structural and emotional information, which is used to automatically generate suggestions. The AI ​​model uses natural language processing to generate interior designs tailored to the user. The output is the proposed interior and storage strategy.

[0606] Step 5:

[0607] The server refers to a product information database and selects products related to the generated suggestions. The input is the suggestion content. Based on market information, it identifies purchase links for related products and selects products that match the user's emotional state. The output is the product links combined with the suggestions.

[0608] Step 6:

[0609] The device displays suggestions and selected products to the user through an intuitive interface. Input consists of the suggested items and product links. The visual UI allows users to easily browse suggestions and access the purchase page for selected products with a click. Output consists of the suggestions and product purchase links that the user sees.

[0610] Step 7:

[0611] Users provide feedback on the suggestions. This feedback is sent to the server via a dedicated application. The input is the user's feedback data. The output is the transmission of feedback information, which is used to train the sentiment engine and improve the quality of future suggestions.

[0612] (Application Example 2)

[0613] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0614] In recent years, there has been a growing demand for technologies to improve the customer experience in physical stores. However, traditional methods fail to adequately provide personalized product recommendations that reflect customer emotions and preferences. Furthermore, conventional technologies struggle to respond immediately to changing customer needs.

[0615] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0616] In this invention, the server includes means for receiving image data acquired by a camera and processing the image data to identify environmental information; a generation device for generating suggestions regarding storage and display based on the environmental information and emotion recognition; a selection device for selecting products related to the suggestions and providing means for purchasing those products by referring to market information; and a presentation device for providing product suggestions to customers using an external display device. This makes it possible to provide product suggestions optimized for customer emotions in physical stores and improve the customer experience.

[0617] A "photography device" is a device used to acquire image data and to photograph surrounding objects and space.

[0618] "Environmental information" refers to information about the room's structure, furniture arrangement, and spatial characteristics, identified based on image data acquired by the camera.

[0619] "Emotion recognition" is the process of determining a user's emotional state from their facial expressions and tone of voice, and acquiring that information.

[0620] A "generation device" is a system that creates storage and display suggestions based on received environmental information and the results of emotion recognition.

[0621] "Market information" refers to information related to consumer purchasing activities, including databases of products in circulation and sales information.

[0622] A "selection device" is a means of selecting products related to a proposal by referring to market information and providing users with opportunities to purchase them.

[0623] An "external display device" is a display device used to provide users with proposed information and information on selected products.

[0624] A specific system for carrying out this invention is configured as follows.

[0625] The server receives image data acquired from imaging devices such as smart glasses and head-mounted displays. The received image data is analyzed using image recognition algorithms such as TensorFlow and OpenCV to identify environmental information such as the structure of a room or store and the arrangement of items. Based on these analysis results, Amazon Rekognition or Microsoft Azure Face API are used as emotion engines to identify emotions from the user's facial expressions and tone of voice, and generate suggestions tailored to the current situation.

[0626] The store's customer service staff, who are the users, receive suggestion information from the server through external display devices such as smart glasses or HMDs. The server then references market information to select products optimized for the user's emotional state, and presents the results via the display device. The market information used includes the latest product databases and sales data.

[0627] This system works as follows: For example, if a staff member serving a customer in a store uses smart glasses to capture the customer's restless expression or tired voice, the system uses this information to suggest products that enhance relaxation, such as an aroma diffuser or a relaxation chair. This allows customers to receive recommendations for products that are best suited to their situation.

[0628] An example of a prompt sentence to input into a generative AI model is: "A store clerk wearing smart glasses observes customers in the store. If the customer appears relaxed, what products would the clerk suggest?"

[0629] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0630] Step 1:

[0631] The server receives image data captured through cameras on smart glasses or head-mounted displays. This image data is provided in real time from devices worn by store staff, who are the users of the system. The input images are treated as material necessary for analyzing environmental information.

[0632] Step 2:

[0633] The server analyzes the received image data using TensorFlow and OpenCV. The image recognition algorithm identifies environmental information such as the structure of the room or store and the arrangement of objects, which is then obtained as output. This output data serves as the basis for emotion recognition in the next step.

[0634] Step 3:

[0635] Based on identified environmental information, the server uses Amazon Rekognition and Microsoft Azure Face APIs to identify the user's emotions. Inputs include the user's facial expressions and tone of voice, while output provides the user's current emotional state. This emotion data is a crucial element in suggestion generation.

[0636] Step 4:

[0637] The server uses a generator to create optimal storage and display suggestions for the user based on acquired emotional data and environmental information. The inputs used are the analysis results and the results of emotional recognition, and the output is specific suggestions for the user. These suggestions are then used to reference market information.

[0638] Step 5:

[0639] The server references market information and selects products relevant to the requested suggestions. It uses the latest product database and sales information as input, and outputs a list of products optimized for the user's emotional state. The selected products are then processed by a selection device that provides purchasing options.

[0640] Step 6:

[0641] Product suggestions provided by a server are displayed on the terminal, such as smart glasses or a head-mounted display. Users can make real-time product suggestions to customers through the display device. This output information becomes a direct purchase suggestion to the customer, making in-store customer service activities more effective.

[0642] Through these steps, the system enables the real-time product recommendations in physical stores that take customer emotions into consideration, thereby improving the customer experience.

[0643] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0644] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0645] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0646] [Fourth Embodiment]

[0647] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0648] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0649] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0650] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0651] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0652] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0653] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0654] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0655] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0656] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0657] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0658] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0659] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0660] This invention provides a storage and display suggestion system using image analysis technology and AI to enable users to easily optimize their rooms at home. This system is generally accessible and easy to use via smartphones and tablets, and is implemented in the following form.

[0661] First, the user takes a photo of a specific room in their home using a mobile device such as a smartphone. It is recommended that the photo be taken in a way that shows the overall layout of the room. The captured image data is sent to the server via the application.

[0662] The server executes a deep learning-based image recognition algorithm to analyze the received image data. This identifies environmental information such as the room layout, furniture types and placement, and space usage. In particular, the image analysis can detect dead spaces that are not being used efficiently.

[0663] Next, a generator built into the server generates suggestions for improving storage and display based on environmental information. These suggestions include new storage methods utilizing dead space, furniture rearrangement, and harmonious display methods. Furthermore, the suggested ideas are optimized for the detected dead space.

[0664] Furthermore, the server accesses a market information database and selects appropriate furniture and storage items related to the suggestion. This selection is based on the user's needs and the condition of the room, ensuring the most suitable choice. The selected items are associated with online store purchase links, allowing the user to easily proceed with the purchase process.

[0665] The terminal displays suggestions and product selection information received from the server to the user. The user interface consists of easy-to-understand graphics and text, allowing users to easily understand the suggestions and use purchase links as needed.

[0666] For example, if a user takes a photo of their living room, the server will identify the placement of the sofa and TV stand and suggest a rack that can be installed in the corner space. A link to an online store will also be provided, allowing the user to purchase the recommended rack and secure new storage space. In this way, the present invention helps to easily and efficiently improve the user's living environment.

[0667] The following describes the processing flow.

[0668] Step 1:

[0669] The user launches a dedicated app on their smartphone and takes a photo that captures the entire room. The app then provides shooting guidelines to help the user take the best possible image.

[0670] Step 2:

[0671] The device compresses the captured image data within the app and sends it to the server via the network. During this process, an appropriate data transfer method is selected, taking into account the user's network bandwidth.

[0672] Step 3:

[0673] The server analyzes the received image data. Using a deep learning model, it identifies the room layout, furniture types, and their placement, and stores this information in a database as environmental data.

[0674] Step 4:

[0675] Based on the analyzed environmental information, the server generates optimal storage and display suggestions using a generation device. These suggestions include specific ideas for utilizing identified dead spaces.

[0676] Step 5:

[0677] The server selects furniture and storage items related to the generated proposal by referring to a market information database. Purchase links to online stores are added to the selected products.

[0678] Step 6:

[0679] The server formats the proposed content and selected product information and sends it to the user's terminal. This data is structured in a format suitable for the user interface.

[0680] Step 7:

[0681] The device visually presents the received data to the user. The app screen clearly displays suggested storage ideas and purchase links, supporting the user's decision-making.

[0682] Step 8:

[0683] Users review the displayed suggestions and purchase information and click on product links that interest them. This directs them to the online store, where they can proceed with the purchase.

[0684] (Example 1)

[0685] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0686] Conventional storage and display suggestion systems made it difficult for users to obtain concrete guidance on how to efficiently utilize their home space. Furthermore, the selection and purchase of items using market information was complex and time-consuming for users. Additionally, the inability to incorporate feedback on the suggested content made it difficult to improve the system's accuracy.

[0687] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0688] In this invention, the server includes a device that receives image information acquired by a shooting function and analyzes the image information to identify spatial information; a generation device that creates suggestions regarding storage and display based on the spatial information; and a selection device that refers to market information and provides the selection of items related to the suggestions and links to purchase those items. As a result, users can receive specific suggestions for effectively utilizing their home space and make quick and appropriate item selections and purchases based on market information. It is also possible to improve the accuracy of the system by utilizing feedback from users.

[0689] The "shooting function" is a feature that allows users to acquire visual information of their environment using their own devices.

[0690] "Image information" refers to visual data acquired through the camera's capture function, and forms the basis for analysis.

[0691] "Spatial information" refers to information about the physical structure and arrangement of the environment obtained from analyzed image information.

[0692] A "generation device" is a device that has the function of creating proposals for efficient storage and display based on spatial information.

[0693] "Market information" refers to data related to the provision of goods and services, and is used to select items related to the proposed content.

[0694] A "selection device" is a device that, based on market information, selects items that match the proposed content and provides the user with a purchase link.

[0695] A "display device" is a device used to visually present selected items to the user, encouraging them to confirm the details of the proposal.

[0696] "Unused space" refers to space that is not currently being effectively utilized, and is a target for proposals regarding new storage methods and furniture arrangements.

[0697] "Feedback information" refers to the opinions and evaluations that users provide regarding the proposed content, and is used as feedback for system improvement.

[0698] This invention utilizes a shooting function, server, terminal, and generating AI model to enable users to efficiently optimize their home space. Specific embodiments of this system are described below.

[0699] Users use devices such as smartphones or tablets to take pictures of specific rooms in their homes. The captured image information is automatically sent to a server via the application. Through this process, data is collected on the server in real time.

[0700] The server utilizes deep learning technology to analyze the received image information. In particular, it uses Convolutional Neural Networks (CNNs) and other techniques to perform object recognition and spatial analysis within the image, thereby identifying spatial information. This allows for the understanding of details such as dead space and furniture placement.

[0701] Next, a generation device integrated into the server uses a generation AI model based on the analyzed spatial information to create storage and display proposals. This model is pre-trained and provides efficient suggestions for diverse spatial layouts. The generated proposals include ideas for effectively utilizing underutilized space.

[0702] Furthermore, the server references market information and selects the most suitable items related to the proposed improvements. The selection device lists the suitable items and associates each with a purchase link to an online store. This allows users to easily purchase the products.

[0703] The terminal graphically displays the received proposals and item information. The user interface is designed to allow users to visually review the proposals, making them easy to understand and act upon.

[0704] For example, when a user takes a photo of their living room, the server recognizes the arrangement of sofas and tables and identifies unused space. The generator then suggests the best storage solution for this space, finds suitable furniture online, and provides purchase links. An example of a prompt message is, "Analyze the photo of my living room and suggest an optimization for storage space."

[0705] This invention enables users to utilize their daily living spaces more efficiently, significantly improving convenience.

[0706] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0707] Step 1:

[0708] Users take photos of their rooms at home using their smartphones or tablets. It is recommended that the photos be composed in a way that shows the overall layout of the room. These images are automatically sent to the server via the application. The input is the image data taken by the user, and the output is the transfer of the image data to the server.

[0709] Step 2:

[0710] The server uses deep learning technology to analyze the received image data. Specifically, it uses a Convolutional Neural Network (CNN) to identify objects within the image and determine spatial information such as floor plan, furniture arrangement, color, and shape. The input is image data sent by the user, and the output is detailed spatial information obtained through image analysis.

[0711] Step 3:

[0712] The server uses a generative AI model based on acquired spatial information to create storage and display suggestions. This generative AI model is trained to generate efficient storage methods and furniture arrangement patterns. The input is analyzed spatial information, and the output is specific improvement suggestions. This process includes ideas for utilizing underutilized spaces.

[0713] Step 4:

[0714] The server automatically selects suitable items by referencing market information related to the generated proposal. This selection considers factors such as item size, design, and price. The selected items are also accompanied by online store purchase links. The input is the generated proposal and market information, and the output is a list of items aligned with the proposal and their purchase links.

[0715] Step 5:

[0716] The terminal displays the suggested content and item information received from the server via a user interface. This interface is designed to be visually intuitive, allowing users to verify suggested storage methods and furniture arrangements using 3D models. Input consists of suggestions and item information from the server, while output is a visual display and information provided to the user.

[0717] (Application Example 1)

[0718] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0719] In modern brick-and-mortar stores, product display and space utilization are largely based on experience, making automation and the use of information technology difficult. In particular, store staff need considerable time and effort to visually analyze space and determine the optimal display. Furthermore, while new display strategies are constantly needed to enhance customer purchasing intent, obtaining objective and real-time optimization suggestions remains challenging.

[0720] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0721] In this invention, the server includes means for receiving image information acquired by a camera and processing the image information to identify the spatial situation; generation means for generating suggestions regarding storage and display based on the spatial situation; and selection means for selecting products related to the suggestions by referring to market information and providing purchase links for those products. This makes it possible to present efficient use of space and optimization of product display in physical stores in real time through a wearable information display device.

[0722] "Photography equipment" refers to a device used to acquire image information of a specific subject.

[0723] "Image information" refers to data acquired by photographic equipment, which visually represents the spatial situation and the arrangement of objects.

[0724] "Spatial conditions" refer to information that describes the physical environment and the arrangement of objects in a specific location.

[0725] "Generating means" refers to a function for creating proposals regarding storage and display based on a specified spatial situation.

[0726] "Market information" refers to data related to the market, such as product types, prices, and distribution information.

[0727] "Selection method" refers to a function that selects the most suitable product based on market information and presents relevant information to the user.

[0728] A "wearable information display device" is a device that is worn by a user to visually present information.

[0729] "Wasted space" refers to space that is not physically used or is not being utilized efficiently.

[0730] "User feedback" refers to information that includes opinions and evaluations from users regarding a proposal.

[0731] A "learning tool" is a function that uses feedback information to improve the system's performance and the accuracy of its suggestions.

[0732] To implement this invention, the following system configuration can be used. The server receives image information transmitted from the imaging device and runs a deep learning model using image recognition technology to analyze it. This analysis identifies spatial conditions and wasted space. For example, TensorFlow, an open-source deep learning library, can be used.

[0733] Once the analysis is complete, the server generates storage and display suggestions based on the results using a generation mechanism. These suggestions include efficient product display and ways to utilize wasted space. The generated suggestions are displayed in real time on a wearable information display device, such as smart glasses.

[0734] Next, the server references market information, selects products suitable for the user using selection methods, and provides relevant information and purchase links. This process involves gathering information from online databases.

[0735] Users can view information presented through smart glasses and intuitively adjust displays in the actual store space. When a specific suggestion is selected, user feedback is accumulated in the system, and suggestions are improved through learning mechanisms.

[0736] As a concrete example, when a staff member at a sporting goods store uses smart glasses to take pictures of the store, a server analyzes the wasted space on the shelves. Subsequently, suggestions for displaying new basketballs in the empty spaces are displayed, along with a purchase link for those products. This is expected to improve the efficiency of product placement and enhance customer satisfaction.

[0737] An example of a prompt using a generative AI model would be: "Analyze the space in the input image and suggest an efficient way to display products. For example, show what kind of products should be displayed in the empty space in the corner."

[0738] Such a system can enable efficient use of space and improved customer experience, especially in physical stores.

[0739] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0740] Step 1:

[0741] The user wears smart glasses and walks around the store, capturing image data of designated areas. This image data is then used as input for the system.

[0742] Step 2:

[0743] The terminal sends the acquired image information to the server. The input data is sent in image formats such as JPEG and PNG. The server receives this image information and uses a deep learning model to analyze the spatial situation. In this analysis, an object recognition algorithm identifies wasted space and existing display arrangements. As a result of the analysis, information on wasted space is output.

[0744] Step 3:

[0745] The server uses an AI model generated based on the analysis results to produce suggestions for storage and display that utilize wasted space. The input data includes spatial information from the analysis results and information on existing merchandise. The generated suggestions become the server's output, presenting optimal product placement and display methods.

[0746] Step 4:

[0747] The server uses the generated suggestions as a reference, queries relevant market information, and selects products using selection criteria. The input data is the generated suggestions, and the output is information on products that match the suggestions and purchase links. This information is obtained from an online database.

[0748] Step 5:

[0749] The terminal displays the suggested content and selected product information received from the server on the user's wearable information display device. The input is display data generated by the server, and the output is display information visually presented on smart glasses. This allows the user to confirm how to use the space while looking at the actual wasted space.

[0750] Step 6:

[0751] Users can implement the proposed installation method and test the results. They can also provide feedback on the usefulness of the proposal and the results of adopting it. This feedback is recorded as input to the system and used by the server's learning mechanism to improve the accuracy of future proposals.

[0752] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0753] This invention combines a system that analyzes image data related to the user's room layout and proposes optimal storage and display solutions with an emotion engine that recognizes the user's emotions and reflects them in the suggestions. As a result, users can receive suggestions optimized for their own emotional state, enabling them to create a more satisfying living space.

[0754] First, the user takes a picture of a specific room in their home using their smartphone camera. This image data is then sent to a server via a dedicated app.

[0755] The server executes an image recognition algorithm to analyze the received image data, identifying environmental information such as room layout, furniture arrangement, and dead space. This analysis builds the foundational data for proposals that enable efficient use of space.

[0756] Next, the server activates an emotion engine to recognize the user's emotions from their facial expressions and tone of voice during the photo shoot. This emotion data is used to generate suggestions that reflect the user's current mood and preferences. For example, if the user prefers a relaxed atmosphere, the server might suggest soft-colored curtains or the placement of houseplants.

[0757] Furthermore, the server references market information and selects products based on sentiment recognition results. This selection includes furniture and interior goods that reflect the user's emotions and preferences. Purchase links for the selected products are provided so that users can access them immediately if they are interested.

[0758] The device displays these suggestions to the user in an easy-to-understand interface. Users can browse the displayed suggestions and products and make selections that match their emotional state. They can also provide feedback on the suggestions, and this feedback is used to improve the accuracy of future suggestions through the emotion engine's learning function.

[0759] As a concrete example, if a user takes a photo of their living room, the server analyzes the image to identify the placement of sofas and tables, and uses an emotion engine to generate suggestions for creating a relaxing space. Furthermore, purchase links for products related to the suggested interior are provided, allowing the user to access the purchase page with a single click. In this way, the present invention, which combines an emotion engine, supports the creation of a more personalized space by providing customized suggestions that take the user's emotions into consideration.

[0760] The following describes the processing flow.

[0761] Step 1:

[0762] The user takes photos of their room at home using a dedicated app. During this process, the app activates the camera to detect the user's facial expressions and records their emotions.

[0763] Step 2:

[0764] The device sends captured image data and user facial expression data to the server. This data is treated as basic information for analysis.

[0765] Step 3:

[0766] The server analyzes the received image data and uses an AI model to identify the room layout and furniture arrangement. At the same time, it detects dead space and registers it in a database.

[0767] Step 4:

[0768] The server activates an emotion engine to recognize emotions from the user's facial expression data. This emotion data is used to adjust the suggestions.

[0769] Step 5:

[0770] The server generates optimal storage and display suggestions based on environmental information and emotional data. The generation device then performs this process, including color coordination and item selection tailored to the emotional state.

[0771] Step 6:

[0772] Based on the proposal, the server selects relevant products from the market information database. Purchase links are then set up for the selected products to allow for easy access.

[0773] Step 7:

[0774] The server sends the suggested content and product information to the terminal. The data is formatted in a way that is easy for the user to understand intuitively.

[0775] Step 8:

[0776] The device displays suggestions through its user interface. Users can view the presented information and click on product links that interest them.

[0777] Step 9:

[0778] Users access external online stores through selected links and complete purchases. They can also provide feedback on the suggestions they receive within the app.

[0779] (Example 2)

[0780] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0781] In modern living spaces, users often face the challenge of efficiently utilizing limited space while creating an environment that is optimal for their individual emotions and preferences. Furthermore, selecting the right products from the many options available on the market is not easy. Therefore, there is a need for a system that automatically provides suggestions that match the user's emotional state and preferences, thereby improving the efficiency of storage and interior design selection.

[0782] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0783] In this invention, the server includes means for receiving visual data acquired by an imaging device and analyzing the visual data to identify structural information; a generation device for creating storage and display suggestions based on the structural information; and an emotion analysis device for analyzing emotional information from audio and visual information and adjusting suggestions based on the user's psychological state. This makes it possible to efficiently utilize limited space and provide personalized suggestions that match the user's emotions and preferences.

[0784] An "imaging device" is a hardware device used to acquire visual data and is used to capture images of the user's environment.

[0785] "Visual data" refers to image information acquired by an imaging device, and serves as the basis for analyzing the structure and layout of a room.

[0786] "Structural information" refers to information about the room layout and furniture arrangement, identified through the analysis of visual data.

[0787] A "generation device" is a functional unit within a system that generates storage and display suggestions for the user based on structural information.

[0788] "Audio and visual information" refers to data including the user's voice and facial expressions, and is used as input data when analyzing emotional information.

[0789] "Emotional information" refers to data about the user's psychological state and emotions, extracted from audio and visual information.

[0790] An "emotion analysis device" is a component within a system that analyzes audio and visual information, generates emotional information, and adjusts suggestions to match the user's psychological state.

[0791] "Market information" refers to information about goods in online or offline trading markets and is used to select items related to the proposal.

[0792] A "display unit" is a functional unit that visually displays the generated proposals and provides an interface for receiving feedback from users.

[0793] "Unused areas" refer to spaces that are not normally utilized, as detected from visual data, and are used to propose storage methods.

[0794] A "purchase channel" is a link or means for purchasing selected items, provided in a format that is easily accessible to the user.

[0795] This invention is a system for providing optimal storage and display suggestions in a user's living space, offering personalized suggestions that take into account the user's emotional state.

[0796] Hardware and Software Overview

[0797] Users acquire visual data of their living space using imaging devices such as smartphones or tablets. This image data is transmitted to a server via a dedicated application.

[0798] The server runs in a cloud computing environment and analyzes visual data using image processing libraries (e.g., OpenCV). It identifies structural information such as room layout and furniture arrangement.

[0799] Furthermore, the server uses speech recognition and facial recognition algorithms (e.g., TensorFlow) to extract emotional information from the user's voice and visual information.

[0800] The generation AI model automatically generates storage and interior design suggestions tailored to the user, based on structural and emotional information.

[0801] The device visually presents these suggestions to the user and provides an intuitive interface.

[0802] Specific examples and prompt statements

[0803] For example, if a user takes a photo of their living room, the server analyzes the image to evaluate the efficiency of the furniture arrangement and suggests ways to utilize dead space. If the system determines that the user is in a relaxed emotional state, it will suggest soft lighting and plant placement. Purchase links for selected items are also provided, allowing the user to easily access the purchase page.

[0804] An example of a prompt message is: "Based on the user's living room image and emotional data, create three suggestions for creating a relaxing space and provide links to related products."

[0805] This invention allows users to make more effective use of their home space and create a comfortable environment optimized for their emotional state.

[0806] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0807] Step 1:

[0808] The user takes a picture of the room with their smartphone camera. The input data is high-resolution visual data, which enables accurate analysis of the environment. The captured image data is sent to a server via a dedicated app. The output is the transfer of the image data to the server.

[0809] Step 2:

[0810] The server stores the received visual data in a cloud environment. Next, it analyzes the data using an image processing library (e.g., OpenCV). The input is image data sent by the user. The server applies an object detection algorithm to identify the room layout and furniture arrangement. In this process, it extracts dead space and important structural information. The output is the structural information of the user's room.

[0811] Step 3:

[0812] The server extracts emotional information using speech and facial recognition algorithms. Input includes audio data and visual facial data. A machine learning model using TensorFlow is applied to analyze emotions from facial expressions and voice tone. The output is emotional data indicating the user's psychological state.

[0813] Step 4:

[0814] The server utilizes a generative AI model to combine structural and emotional information to generate storage and interior design suggestions. The input consists of analyzed structural and emotional information, which is used to automatically generate suggestions. The AI ​​model uses natural language processing to generate interior designs tailored to the user. The output is the proposed interior and storage strategy.

[0815] Step 5:

[0816] The server refers to a product information database and selects products related to the generated suggestions. The input is the suggestion content. Based on market information, it identifies purchase links for related products and selects products that match the user's emotional state. The output is the product links combined with the suggestions.

[0817] Step 6:

[0818] The device displays suggestions and selected products to the user through an intuitive interface. Input consists of the suggested items and product links. The visual UI allows users to easily browse suggestions and access the purchase page for selected products with a click. Output consists of the suggestions and product purchase links that the user sees.

[0819] Step 7:

[0820] Users provide feedback on the suggestions. This feedback is sent to the server via a dedicated application. The input is the user's feedback data. The output is the transmission of feedback information, which is used to train the sentiment engine and improve the quality of future suggestions.

[0821] (Application Example 2)

[0822] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0823] In recent years, there has been a growing demand for technologies to improve the customer experience in physical stores. However, traditional methods fail to adequately provide personalized product recommendations that reflect customer emotions and preferences. Furthermore, conventional technologies struggle to respond immediately to changing customer needs.

[0824] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0825] In this invention, the server includes means for receiving image data acquired by a camera and processing the image data to identify environmental information; a generation device for generating suggestions regarding storage and display based on the environmental information and emotion recognition; a selection device for selecting products related to the suggestions and providing means for purchasing those products by referring to market information; and a presentation device for providing product suggestions to customers using an external display device. This makes it possible to provide product suggestions optimized for customer emotions in physical stores and improve the customer experience.

[0826] A "photography device" is a device used to acquire image data and to photograph surrounding objects and space.

[0827] "Environmental information" refers to information about the room's structure, furniture arrangement, and spatial characteristics, identified based on image data acquired by the camera.

[0828] "Emotion recognition" is the process of determining a user's emotional state from their facial expressions and tone of voice, and acquiring that information.

[0829] A "generation device" is a system that creates storage and display suggestions based on received environmental information and the results of emotion recognition.

[0830] "Market information" refers to information related to consumer purchasing activities, including databases of products in circulation and sales information.

[0831] A "selection device" is a means of selecting products related to a proposal by referring to market information and providing users with opportunities to purchase them.

[0832] An "external display device" is a display device used to provide users with proposed information and information on selected products.

[0833] A specific system for carrying out this invention is configured as follows.

[0834] The server receives image data acquired from imaging devices such as smart glasses and head-mounted displays. The received image data is analyzed using image recognition algorithms such as TensorFlow and OpenCV to identify environmental information such as the structure of a room or store and the arrangement of items. Based on these analysis results, Amazon Rekognition or Microsoft Azure Face API are used as emotion engines to identify emotions from the user's facial expressions and tone of voice, and generate suggestions tailored to the current situation.

[0835] The store's customer service staff, who are the users, receive suggestion information from the server through external display devices such as smart glasses or HMDs. The server then references market information to select products optimized for the user's emotional state, and presents the results via the display device. The market information used includes the latest product databases and sales data.

[0836] This system works as follows: For example, if a staff member serving a customer in a store uses smart glasses to capture the customer's restless expression or tired voice, the system uses this information to suggest products that enhance relaxation, such as an aroma diffuser or a relaxation chair. This allows customers to receive recommendations for products that are best suited to their situation.

[0837] An example of a prompt sentence to input into a generative AI model is: "A store clerk wearing smart glasses observes customers in the store. If the customer appears relaxed, what products would the clerk suggest?"

[0838] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0839] Step 1:

[0840] The server receives image data captured through cameras on smart glasses or head-mounted displays. This image data is provided in real time from devices worn by store staff, who are the users of the system. The input images are treated as material necessary for analyzing environmental information.

[0841] Step 2:

[0842] The server analyzes the received image data using TensorFlow and OpenCV. The image recognition algorithm identifies environmental information such as the structure of the room or store and the arrangement of objects, which is then obtained as output. This output data serves as the basis for emotion recognition in the next step.

[0843] Step 3:

[0844] Based on identified environmental information, the server uses Amazon Rekognition and Microsoft Azure Face APIs to identify the user's emotions. Inputs include the user's facial expressions and tone of voice, while output provides the user's current emotional state. This emotion data is a crucial element in suggestion generation.

[0845] Step 4:

[0846] The server uses a generator to create optimal storage and display suggestions for the user based on acquired emotional data and environmental information. The inputs used are the analysis results and the results of emotional recognition, and the output is specific suggestions for the user. These suggestions are then used to reference market information.

[0847] Step 5:

[0848] The server references market information and selects products relevant to the requested suggestions. It uses the latest product database and sales information as input, and outputs a list of products optimized for the user's emotional state. The selected products are then processed by a selection device that provides purchasing options.

[0849] Step 6:

[0850] Product suggestions provided by a server are displayed on the terminal, such as smart glasses or a head-mounted display. Users can make real-time product suggestions to customers through the display device. This output information becomes a direct purchase suggestion to the customer, making in-store customer service activities more effective.

[0851] Through these steps, the system enables the real-time product recommendations in physical stores that take customer emotions into consideration, thereby improving the customer experience.

[0852] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0853] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0854] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0855] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0856] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0857] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0858] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0859] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0860] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0861] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0862] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0863] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0864] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0865] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0866] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0867] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0868] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0869] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0870] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0871] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0872] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0873] The following is further disclosed regarding the embodiments described above.

[0874] (Claim 1)

[0875] A means for receiving image data acquired by a camera and processing the image data to identify environmental information,

[0876] A generating device that generates proposals regarding storage and display based on the aforementioned environmental information,

[0877] A selection device that, by referring to market information, selects products related to the aforementioned proposal and provides purchase links for said products,

[0878] A system that includes this.

[0879] (Claim 2)

[0880] The system according to claim 1, which detects dead space from the aforementioned image data and proposes a storage means that utilizes said dead space.

[0881] (Claim 3)

[0882] The system according to claim 1, further comprising a learning device that receives user feedback and uses it to improve the generated suggestions.

[0883] "Example 1"

[0884] (Claim 1)

[0885] A device that receives image information acquired by a shooting function, analyzes the image information, and identifies spatial information,

[0886] A generating device that creates proposals for storage and display based on the aforementioned spatial information,

[0887] A selection device that refers to market information and provides the selection of items related to the aforementioned proposal and links to purchase said items,

[0888] A display device for visually displaying the selected items,

[0889] A system that includes this.

[0890] (Claim 2)

[0891] The system according to claim 1, which detects unused space from the aforementioned image information and proposes a storage means for utilizing the unused space.

[0892] (Claim 3)

[0893] The system according to claim 1, further comprising a learning device that receives feedback from users and uses it to improve the generated suggestions.

[0894] "Application Example 1"

[0895] (Claim 1)

[0896] A means for receiving image information acquired by a camera and processing said image information to identify the spatial situation,

[0897] A generation means for generating proposals regarding storage and display based on the aforementioned spatial conditions,

[0898] A selection means that selects products related to the aforementioned proposal by referring to market information and provides a purchase link for said products,

[0899] A means of presenting analysis results and optimization proposals in real space through a wearable information display device,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, which detects wasted space from the aforementioned image information, proposes a storage means that utilizes the wasted space, and visually presents it in real space by using a wearable information display device.

[0903] (Claim 3)

[0904] The system according to claim 1, further comprising a learning means for receiving user feedback and using it to improve the generated suggestions.

[0905] "Example 2 of combining an emotion engine"

[0906] (Claim 1)

[0907] A means for receiving visual data acquired by an imaging device and analyzing the visual data to identify structural information,

[0908] A generating device that creates proposals regarding storage and display based on the aforementioned structural information,

[0909] An emotion analysis device that analyzes emotional information from audio and visual information and adjusts suggestions based on the user's psychological state,

[0910] A selection device that, by referring to market information, provides the selection of articles related to the proposal and the purchasing route for said articles,

[0911] A display device that presents proposals to the user through a visual operation screen and accepts the user's opinions on said proposals,

[0912] A learning device used to improve the generated suggestions,

[0913] A system that includes this.

[0914] (Claim 2)

[0915] The system according to claim 1, which detects unused areas from the aforementioned visual data and proposes a storage method that utilizes said unused areas.

[0916] (Claim 3)

[0917] The system according to claim 1, wherein the generating device adjusts the proposed content based on emotional information provided by the user.

[0918] "Application example 2 of combining emotional engines"

[0919] (Claim 1)

[0920] A means for receiving image data acquired by a camera and processing the image data to identify environmental information,

[0921] A generating device that generates suggestions regarding storage and display based on the aforementioned environmental information and emotion recognition,

[0922] A selection device that, by referring to market information, provides a means for selecting and purchasing products related to the aforementioned proposal,

[0923] A presentation device that provides product suggestions to customers using an external display device,

[0924] A system that includes this.

[0925] (Claim 2)

[0926] The system according to claim 1, which detects spatially inefficient regions from the aforementioned image data and proposes a means for utilizing said inefficient regions.

[0927] (Claim 3)

[0928] The system according to claim 1, further comprising an educational device for receiving user responses and using them to improve the generated suggestions. [Explanation of symbols]

[0929] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving image data acquired by a camera and processing the image data to identify environmental information, A generating device that generates proposals regarding storage and display based on the aforementioned environmental information, A selection device that, by referring to market information, selects products related to the aforementioned proposal and provides purchase links for said products, A system that includes this.

2. The system according to claim 1, which detects dead space from the aforementioned image data and proposes a storage means that utilizes said dead space.

3. The system according to claim 1, further comprising a learning device that receives user feedback and uses it to improve the generated suggestions.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A