System

A system using video analysis and machine learning identifies tidying points and provides guidance to reduce the burden and increase motivation for tidying, while facilitating housekeeping service use.

JP2026019768APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121516
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Individuals, especially those who dislike tidying up or children, face challenges in starting and completing room tidying tasks due to lack of guidance, motivation, and appropriate methods, while parents lack effective ways to encourage their children, and users of housekeeping services struggle to communicate room conditions for quotes.

Method used

A system that captures room video, analyzes it using computer vision and machine learning to identify tidying points, provides step-by-step instructions, praises users, suggests further tasks, and matches with housekeeping services based on room conditions.

Benefits of technology

Reduces the burden and increases motivation for tidying by breaking tasks into manageable steps, helps develop tidying habits, and facilitates efficient use of housekeeping services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019768000001_ABST
    Figure 2026019768000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system including means for capturing a moving image of a room, means for analyzing the captured moving image to recognize an object and grasp a state of the room, means for identifying a cleanup point that can be completed within 15 minutes based on an analysis result, means for presenting the identified cleanup point and a procedure thereof to a user, means for displaying a message that praises the user after completion of the cleanup, and means for proposing a further cleanup point to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] For people who dislike tidying up or children who want to make tidying a habit, tidying up an entire room is a big burden, and they often don't know where to start and end up not doing anything at all. Furthermore, if motivation for tidying is low or time is limited, even small tasks can feel like a hassle. Another issue is the lack of appropriate methods and advice for parents to help their children develop the habit of tidying up. Additionally, there is also the problem that even those who want to use housekeeping services have no way to properly communicate the state of the room in advance, making it difficult to get a quote. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for capturing video of a room, a means for analyzing the captured video to recognize objects and grasp the state of the room, a means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, a means for presenting the identified tidying up points and the steps to the user, a means for displaying a message praising the user after the tidying up is completed, and a means for suggesting further tidying up points to the user. This system reduces the user's sense of burden in tidying up and increases motivation by setting small goals that can be achieved in a short time. Parents can also learn effective ways to encourage their children to tidy up and have them follow the app's guidance. Furthermore, by using a function that matches with housekeeping services based on the state of the room, advance estimates can be easily obtained, facilitating the use of housekeeping services.

[0006] "Video in a room" is video data that a user captures of part or the entirety of their daily living space.

[0007] "Means of capturing images" refers to the technology and functions for recording images using smart devices or mobile terminals equipped with camera functions.

[0008] "Analysis and object recognition" is the process of using computer vision and machine learning algorithms to identify and classify objects, furniture, etc. in the footage.

[0009] "Means for understanding the condition of a room" refers to the technology and its functions for evaluating the placement of items in a room and the level of clutter from the results of video analysis, and checking the overall condition.

[0010] A "15-minute tidying point" is a specific area or collection of items that a user can effectively tidy up in 15 minutes or less.

[0011] The "means of identification" refers to the technology and its functions that analyze the state of the room and automatically detect the target points for tidying up.

[0012] The "tidying up procedure" refers to specific work instructions and methods that allow the user to efficiently tidy up the identified tidying up points.

[0013] The "presentation means" refers to a user interface or notification function for displaying information such as tidying up points and procedures to the user.

[0014] A "praising message" is a word or phrase of praise intended to praise the user's efforts and motivate them after they have successfully tidied up.

[0015] "Display means" refers to the technology and functions that allow users to see information on the screen of a smart device or mobile terminal.

[0016] The "means for suggesting the next tidying point" is a technology and function that detects new tidying targets and notifies the user of that information when the user wants to continue tidying up.

[0017] "Matching with housekeeping services" is the process of selecting an appropriate housekeeping service provider based on information about the room's condition and providing the service.

[0018] The "means for sending a quote request" refers to the technology and its function for sending details of the room condition and cleaning request to a housekeeping service provider and obtaining the price of the service. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] The present invention provides an application system that makes it easier for users to tidy up their rooms, and an embodiment of the system is specifically described below. In particular, by presenting tidying up points that can be completed in a short time and showing the steps in an easy-to-understand manner, the system allows users to tackle tidying up without feeling any resistance.

[0041] This system allows users to use devices such as smartphones or tablets to take videos of their rooms and then analyze those videos. Specifically, the user launches the app and sends the video of the room's condition to the server. The server then analyzes the received video data, recognizes objects in the room, and determines their placement and the level of clutter. This analysis uses computer vision technology and machine learning algorithms.

[0042] Based on the video analysis results, the server identifies tidying up tasks that the user can complete within 15 minutes and generates instructions for doing so. For example, if the user is trying to tidy up documents scattered on a desk, the server will classify the documents by category and present instructions for putting them away in the appropriate storage location. This information is sent to the user's device, and the user can view the specific tidying up steps on the app screen.

[0043] After the user has finished tidying up, they press the "Tidy up complete" button in the app, and the device sends a completion notification to the server. The server receives this notification and generates a message praising the user. For example, the message might read, "Great job! Your desk is so tidy now!" If the user wants to continue tidying up, the option "Do you want to continue?" is presented. If the user answers "Yes," the server identifies new tidying points and sends the instructions to the device again. Through this series of steps, the user can gradually tidy up the entire room.

[0044] In addition, the system also provides a matching function for housekeeping services. When a user selects the "Ask a Professional" button within the app, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can review this quote and decide whether to use the service.

[0045] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The server then generates instructions such as "Classify documents into three categories (work, school, and personal) and put work-related documents in a red folder," and sends these to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying, "Great job! Your desk looks so tidy now!" If the user wants to continue tidying, the server asks, "Do you want to continue?". If the user answers "yes," it presents instructions for tidying up points (for example, inside the desk drawers) that can be completed within the next 15 minutes.

[0046] In this way, the system of the present invention provides the user with appropriate means and advice to help them tidy up their room effectively and efficiently, thereby reducing their resistance to tidying up and helping them develop good tidying habits.

[0047] The processing flow will be explained below.

[0048] Step 1:

[0049] The user launches the app using a smartphone or tablet and takes a video of the room.

[0050] Step 2:

[0051] The device stores the video data captured within the app and sends it to a server via the Internet.

[0052] Step 3:

[0053] The server receives the video data and the analysis module performs object recognition and classification within the video, using computer vision and machine learning algorithms.

[0054] Step 4:

[0055] Based on the analysis results, the server evaluates the arrangement and clutter of items and furniture in the room to grasp the overall condition.

[0056] Step 5:

[0057] The server identifies areas and tasks that can be completed within 15 minutes, and selects areas and tasks that are easy for users to complete in a short time.

[0058] Step 6:

[0059] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[0060] Step 7:

[0061] The server transmits the identified cleaning points and procedures to the terminal.

[0062] Step 8:

[0063] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[0064] Step 9:

[0065] The user follows the instructions to tidy up, and when tidying up is complete, presses the "tidy up complete" button.

[0066] Step 10:

[0067] The terminal sends a "tidying up completed" notification to the server.

[0068] Step 11:

[0069] The server receives the completion notification and generates a message praising the user.

[0070] Step 12:

[0071] A server-generated compliment message is sent to the device.

[0072] Step 13:

[0073] The terminal displays a complimentary message on the user interface.

[0074] Step 14:

[0075] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[0076] Step 15:

[0077] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[0078] Step 16:

[0079] The server checks the room status again and identifies areas that can be cleaned up within the next 15 minutes.

[0080] Step 17:

[0081] The server generates a new cleanup procedure and sends it to the terminal.

[0082] Step 18:

[0083] The terminal displays the new tidying up points and procedures received from the server on the user interface, and the user can confirm them and choose whether to tidy up again.

[0084] Optional Process:

[0085] Step A1:

[0086] The user selects the "Ask a Pro" button within the app.

[0087] Step A2:

[0088] The terminal transmits the room status information to the server.

[0089] Step A3:

[0090] Based on the received room status information, the server sends a quote request to an affiliated housekeeping service provider.

[0091] Step A4:

[0092] The housekeeping service provider creates a quote and sends it to the server.

[0093] Step A5:

[0094] The server sends the quote information to the terminal.

[0095] Step A6:

[0096] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[0097] Example 1

[0098] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0099] For many people, tidying up their rooms is a tedious and time-consuming task. This can lead to resistance or a sense of burden, making it difficult to keep a room clean. Even when hiring a professional, finding the right service can be a complicated process, resulting in cost and effort. This invention aims to solve these problems by providing a means for users to efficiently tidy up their rooms in a short amount of time.

[0100] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0101] In this invention, the server includes a means for recording images of the room, a means for analyzing the recorded images to identify objects and grasp the state of the room, and a means for identifying tidying up points that can be completed in a short time based on the analysis results. This allows the user to easily know the specific steps for tidying up the room, reducing the hassle. Furthermore, by matching with a housekeeping service, it is possible to reduce the effort required to hire a professional.

[0102] The "means for recording images within the room" refers to a camera function or a video recording function for recording the state of the room.

[0103] "Means for analyzing recorded images to identify objects and understand the state of a room" refers to methods that use computer vision techniques and machine learning algorithms to detect and identify objects from image data and understand the state of a room.

[0104] "Means for identifying tidying up points that can be completed in a short time based on the analysis results" refers to a method for selecting specific locations or items that a user can complete tidying up in a short time based on the analyzed image data.

[0105] "Means for displaying the identified tidying up points and the procedures to the user" refers to the functionality of a display or application for visually or textually showing the tidying up points and the procedures to the user.

[0106] The "means for displaying an acknowledgement message to the user after tidying up" refers to a method for displaying a message praising the user after tidying up is completed.

[0107] "Means for suggesting additional tidying up points to the user" refers to a method for suggesting additional tidying up tasks to the user.

[0108] "Means for sending a quote request to be used for matching with a housekeeping service" refers to a method for sending a quote request to a housekeeping service provider based on the condition of the room and matching with a user.

[0109] This invention is a system that allows a user to efficiently and effectively tidy up a room. Specifically, the system allows a user to take a video of the room using a device such as a smartphone or tablet, and analyzes the video to present tidying tips and procedures that the user can implement in a short amount of time. An embodiment of this system is described in detail below.

[0110] First, the user launches the application on their smartphone or tablet and records a video of the room's condition. This video is then sent from the device to the server. The server then analyzes the video data using computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow). The goal of the analysis is to identify objects in the room and understand the state of the room.

[0111] As a concrete example, consider the case where a user wants to tidy up their living room. The user uses a smartphone to record video of the living room and sends the video from the device to a server. The server processes the received video data and identifies objects in the living room (e.g., magazines, remote controls, cushions, etc.). Based on the analysis results, the server generates instructions such as "classify the magazines into three categories (fashion, sports, and news) and put them away in a specific place on the bookshelf." This instruction is then sent to the device and presented to the user.

[0112] The user follows the presented steps to tidy up. After completing the tidy up, the user presses the "Tidy up complete" button in the app. A completion notification is sent from the device to the server, and the server generates a message of praise for the user and displays it on the device. For example, a message such as "Great job! Your living room is now so tidy!" may be displayed.

[0113] If the user wants to continue tidying up, the system asks the user, "Do you want to continue?" If the user answers "yes," the server identifies the next tidying point and sends the instructions to the device again. For example, the next step might be, "Clean up the items on the living room table."

[0114] The system also provides a matching function with housekeeping services. When a user selects the "Ask a Professional" button, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The quote from the provider is sent to the device via the server, and the user can review it and decide whether to use the service.

[0115] An example of a prompt sentence is, "I took a picture of the state of my desk with my smartphone and sent it to the server. Please generate easy steps to tidy up the documents on my desk, and if necessary, please also suggest the next step to tidy up."

[0116] This allows users to gradually tidy up their rooms and easily use housekeeping services. The system aims to reduce users' resistance to tidying up and help them develop the habit of tidying up.

[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0118] Step 1:

[0119] The user takes a video of the room.

[0120] A user uses a smartphone or tablet to record a video of the state of a room. For example, they can record a video so that magazines, remote controls, and other items in the living room are visible. The input is video data of the current state of the room, and the output is a video file saved in the device's storage.

[0121] Step 2:

[0122] The device sends the video to the server.

[0123] The captured video is uploaded from the device to the server. At this time, the device compresses the video data before transferring it to reduce communication delays. The input is the video file stored in the device's storage, and the output is the video data uploaded to the server.

[0124] Step 3:

[0125] The server analyzes the video and identifies the object.

[0126] The server uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow) to analyze the received video data. In this process, an object detection algorithm is used to identify objects in the video (e.g., magazines, remote controls, cushions, etc.). The input is the video data uploaded to the server, and the output is the object identification results (object type and location).

[0127] Step 4:

[0128] The server identifies the tidying up point based on the analysis results.

[0129] Based on the location information of the identified objects, the server identifies specific tidying up steps that the user can complete in a short time. For example, it generates a procedure such as "sort the magazines in the living room into three categories (fashion, sports, and news) and store them in a specific place on the bookshelf." The input is the object identification result, and the output is the identification of the tidying up steps (specific tidying up steps).

[0130] Step 5:

[0131] The server sends the generated cleanup procedure to the terminal.

[0132] The specific cleanup procedure created by the server is sent to the terminal. The input is the cleanup procedure generated by the server, and the output is the cleanup procedure sent to the terminal.

[0133] Step 6:

[0134] The user follows the cleaning procedure to perform the cleaning.

[0135] The user follows the instructions displayed on the device application to tidy up. For example, the user performs a specific action such as "sorting magazines into fashion, sports, and news, and storing them on the bookshelf." The input is the tidying up instructions displayed on the device, and the output is the physical state of the room after the tidying up is completed.

[0136] Step 7:

[0137] The user sends a cleanup completion notification to the server.

[0138] After the user has finished cleaning up, they press the "Clean Up Complete" button in the application. This causes a completion notification to be sent from the device to the server. The input is the user's operation, and the output is the completion notification sent to the server.

[0139] Step 8:

[0140] The server generates a completion message and sends it to the terminal.

[0141] The server receives the completion notification and generates a message praising the user. For example, it creates a message such as "Great job! Your living room is now so tidy!" and sends it to the terminal. The input is the completion notification, and the output is the completion message sent to the terminal.

[0142] Step 9:

[0143] The user chooses whether to clean up further.

[0144] The terminal presents the user with the option "Do you want to continue?" If the user answers "Yes," the procedure to suggest the next tidying up point begins. The input is the user's selection, and the output is the start of the suggestion of the next tidying up point.

[0145] Step 10:

[0146] The server identifies the next cleanup point.

[0147] If the user answers "yes," the server identifies the next step to clean up and generates instructions for it. For example, it creates a specific step such as "Clean up the items on the living room table" and sends it to the device. The input is the user's selection, and the output is the next step to clean up that was sent to the device.

[0148] Step 11:

[0149] If the user selects a housekeeping service, the server performs the matching.

[0150] When a user selects the "Ask a Professional" button in the application, the device sends the room's condition information to the server. Based on this information, the server sends a quote request to a housekeeping service provider and displays the quote from the provider to the user. The input is the room's condition information, and the output is the quote for the housekeeping service.

[0151] (Application example 1)

[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0153] The present invention relates to an application system that allows users to easily tidy up their rooms. Conventional tidying support systems have the problem that when a user starts tidying up, it is unclear which object to start with, making it difficult to tidy up effectively. A similar problem exists in tidying up physical stores, where there is a lack of specific guidance for store staff to work efficiently. The present invention aims to solve these problems.

[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0155] In this invention, the server includes means for taking video of the room, means for analyzing the video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is complete, means for suggesting further tidying up points to the user, means for taking video of the tidying up state of the physical store, and means for presenting the tidying up procedures to store staff. This enables users and store staff to specifically and efficiently tidy up and organize in a short amount of time.

[0156] "Means for capturing video of the inside of a room" refers to equipment or a method for capturing video of the inside of a room using a camera and acquiring the video data.

[0157] "Means for analyzing captured video to recognize objects and grasp the state of a room" refers to equipment and methods that use video analysis technology to identify items present in a room based on acquired video data and grasp the overall situation of the room.

[0158] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method or algorithm that selects and identifies tidying up tasks that a user can complete in a short amount of time based on the results of video analysis.

[0159] "Means for presenting the identified tidying up points and procedures to the user" refers to equipment or methods for displaying or notifying the user of the identified work content and specific procedures so that the user can tidy up efficiently.

[0160] The "means for displaying a message praising the user after tidying up is completed" is a means for presenting a message praising the user's efforts when the user has finished tidying up.

[0161] The "means for suggesting further tidying up points to the user" is a means for presenting new tidying up tasks to the user when the user wants to continue tidying up.

[0162] "Means for photographing the tidiness and neatness of a physical store" refers to equipment or a method for using a camera to photograph the state of items and shelves in a physical store and obtain the video data.

[0163] "Means for presenting tidying procedures to store staff" refers to devices or methods for displaying or informing store staff of specified work content and specific procedures so that they can efficiently tidy up items.

[0164] The present invention provides a system that allows users to efficiently tidy up and organize their rooms and physical stores. Specific embodiments of the system will be described below.

[0165] This system uses devices such as smartphones and tablets to record footage of the state of a room or physical store, and then analyzes the video data to present cleaning and tidying procedures. Users use their devices to record video of the room or store and send the video to a server. The server analyzes the video data and uses computer vision technology and machine learning algorithms to recognize objects. Libraries such as OpenCV and TensorFlow can be used for this analysis.

[0166] Based on the results of the video analysis, the server identifies tidying up tasks that the user can complete within 15 minutes and automatically generates detailed instructions. The generated instructions are sent to the user's device, showing them exactly how to proceed with the tidying up. For example, specific steps such as "sort documents by category and put work-related documents in a red folder" are presented.

[0167] The user checks these steps through the app screen and performs the tidying task. After completing the tidying, the user presses the "Tidying up complete" button on the device, and the device sends a notification to the server. The server receives this notification and generates and displays a message of praise to the user, such as "Great job! Your desk is so tidy now!" If the user wants to continue tidying, it presents the option "Do you want to continue?". If the user answers "Yes," the server identifies new tidying points, generates new steps, and sends them to the device.

[0168] In the case of physical stores, the photographing and analysis procedures using the device are similar. The server identifies the state of organization of the physical store and presents specific organizational procedures to store staff. For example, it presents procedures for sorting products that are randomly placed on shelves and tidying them up by category. This allows store staff to work efficiently.

[0169] The hardware used is a smartphone, tablet, and server, and the software uses OpenCV for video analysis and TensorFlow for machine learning algorithms.

[0170] As a concrete example, if magazines are placed in a disorganized manner in a storeroom, a store clerk can use a smartphone to take a video of the situation and analyze the video. Based on the analysis results, a procedure will be generated, such as "sort the magazines by title and organize them on specific shelves." This allows the store clerk to follow the instructions and efficiently organize the magazines.

[0171] An example of a prompt for a generative AI model might be, "Please take a survey for an application that takes photos of the state of shelves in a physical store and generates the types of items and the procedures for organizing them."

[0172] This system allows users and store staff to tidy up and organize in a specific and efficient manner in a short amount of time.

[0173] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0174] Step 1:

[0175] Users use devices such as smartphones or tablets to shoot video of a room or a physical store.

[0176] Input: Video data captured using the device's camera function.

[0177] Output: The captured video file.

[0178] This video file will be used for later analysis.

[0179] Step 2:

[0180] The device sends the captured video to the server.

[0181] Input: Video files stored on your device.

[0182] Output: Video data sent to the server over the network.

[0183] The server receives this data and prepares it for analysis.

[0184] Step 3:

[0185] The server analyzes the received video data, recognizes objects, and understands the condition of the room or physical store.

[0186] Input: Video data sent to the server.

[0187] Output: Object recognition results in the video and room / store status information.

[0188] This analysis uses libraries such as OpenCV and TensorFlow to identify the location and type of object.

[0189] Step 4:

[0190] Based on the results of the video analysis, the server identifies tidying up points that the user can complete within 15 minutes.

[0191] Input: Object recognition results and room / store state information.

[0192] Output: Identified cleanup points and specific cleanup procedures.

[0193] Specific instructions are automatically generated, detailing which items to organize and how.

[0194] Step 5:

[0195] The server transmits the identified tidying up points and their procedures to the terminal and presents them to the user.

[0196] Input: Server-generated cleanup instructions.

[0197] Output: Cleanup instructions displayed on the user's terminal.

[0198] By following this procedure, the user can efficiently tidy up.

[0199] Step 6:

[0200] After the user has finished tidying up, he / she presses the "tidying up complete" button on the terminal.

[0201] Input: Cleanup completion notification action from user.

[0202] Output: Tidy-up completion notification sent to the server.

[0203] This notification lets the server know the progress of the cleanup.

[0204] Step 7:

[0205] The server receives the notification that the tidying up is complete, generates a message praising the user, such as "good job," and sends it to the terminal.

[0206] Input: Notification from user that cleanup is complete.

[0207] Output: A compliment message displayed on the terminal.

[0208] This gives the user a sense of accomplishment.

[0209] Step 8:

[0210] The server asks the user whether he / she wants to continue cleaning up, and if he / she does, it identifies a new cleaning point, generates a procedure, and sends it to the terminal.

[0211] Input: User confirms their desire to continue tidying.

[0212] Output: New cleanup procedure.

[0213] This allows the user to move on to the next tidying task.

[0214] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0215] This invention provides a system that presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[0216] The system begins when a user takes a video of their room using a device such as a smartphone or tablet and sends the video to the server. The server analyzes the received video to recognize objects and determine the state of the room. This analysis is carried out using computer vision and machine learning algorithms. Based on the results of this analysis, the server identifies tidying up points that the user can complete within 15 minutes and generates a procedure for doing so.

[0217] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotional state. This emotion engine analyzes the user's facial expressions and tone of voice via a camera and microphone to grasp the user's current emotional state. This allows the system to take the user's emotional state into consideration when presenting tidying points and procedures.

[0218] When a user tidies up through the app and presses the "Tidy up complete" button after completing the task, the device sends a completion notification to the server. The server generates a message of praise according to the user's emotional state based on the results of emotion analysis by the emotion engine. For example, if the user is tired, a message such as "You did a great job! Let's take a break" will be displayed.

[0219] Furthermore, if the user chooses to continue tidying up, the system also takes their emotional state into account when suggesting the next step: if the user is in a positive emotional state, it may suggest a larger or more complex task, while if the user is in a negative emotional state, it may suggest an easier task.

[0220] The system also provides a matching function for housekeeping services. When the user selects the "Ask a Professional" button, the device sends information about the room's condition and the results of analysis by the emotion engine to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can then review the quote and decide whether to use the service.

[0221] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The emotion engine also analyzes the user's facial expressions and recognizes that the user is motivated. The server then generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder" and sends them to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying "Good job! Your desk looks so tidy now!" If the user wants to continue tidying, the system asks "Do you want to continue?". If the user answers "yes," the system suggests tidying up points that can be completed within the next 15 minutes (for example, tasks related to the desk drawers).

[0222] In this way, the system of the present invention provides tidying up advice that takes into account the user's emotional state, thereby reducing resistance to tidying up and helping to form tidying up habits. Furthermore, by facilitating the use of housekeeping services, the system provides more efficient tidying up support.

[0223] The processing flow will be explained below.

[0224] Step 1:

[0225] The user launches the app using a smartphone or tablet and takes a video of the room.

[0226] Step 2:

[0227] The device stores the video data captured within the app and sends it to a server via the Internet.

[0228] Step 3:

[0229] The server receives the video data and the analysis module recognizes and classifies objects in the video using computer vision and machine learning algorithms. The server then determines the location, type, and overall condition of the recognized objects.

[0230] Step 4:

[0231] Based on the analysis results, the server identifies tidying up points that the user can complete within 15 minutes, and also selects areas and tasks that the user can easily accomplish in a short amount of time.

[0232] Step 5:

[0233] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[0234] Step 6:

[0235] The emotion engine analyzes the user's facial expressions and tone of voice through the device's camera and microphone to recognize the user's emotional state.

[0236] Step 7:

[0237] The server then adjusts the cleaning procedure and approach based on the emotional state of the user based on the analysis results of the emotion engine. For example, it adds positive comments to users in a positive emotional state and encouraging comments to users in a negative emotional state.

[0238] Step 8:

[0239] The server transmits the adjusted cleaning points and procedures to the terminal based on the emotional state.

[0240] Step 9:

[0241] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[0242] Step 10:

[0243] The user follows the instructions to tidy up, and when finished, presses the "tidy up complete" button.

[0244] Step 11:

[0245] The terminal sends a "tidying up completed" notification to the server.

[0246] Step 12:

[0247] The server generates a praising message according to the user's emotional state based on the emotion analysis results of the emotion engine.

[0248] Step 13:

[0249] A server-generated compliment message is sent to the device.

[0250] Step 14:

[0251] The device will display a praising message on the user interface. For example, if the user is tired, it will say "Great job! Let's take a break."

[0252] Step 15:

[0253] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[0254] Step 16:

[0255] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[0256] Step 17:

[0257] The server checks the room status again and identifies new cleanup points that can be completed within 15 minutes.

[0258] Step 18:

[0259] The server uses an emotion engine to analyze the user's current emotional state.

[0260] Step 19:

[0261] The server generates a new cleaning procedure based on the emotional state and sends it to the terminal.

[0262] Step 20:

[0263] The terminal displays the new tidying up points and procedures on the user interface, and the user can confirm this and choose whether to tidy up again.

[0264] Optional Process:

[0265] Step A1:

[0266] The user selects the "Ask a Pro" button within the app.

[0267] Step A2:

[0268] The device sends the room status and the emotion engine's analysis results to the server.

[0269] Step A3:

[0270] Based on the information received by the server, a request for an estimate is sent to an affiliated housekeeping service provider.

[0271] Step A4:

[0272] The housekeeping service provider creates a quote and sends it to the server.

[0273] Step A5:

[0274] The server sends the quote information to the terminal.

[0275] Step A6:

[0276] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[0277] As a concrete example, if a user wants to tidy up the documents on their desk, the process would proceed as follows: The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and understands the state of the desk. The emotion engine recognizes that the user is motivated. The server generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder," and sends these instructions to the device. When the user completes the tidying up by following the instructions and presses the "Tidying up complete" button, the server generates a message saying "Great job! Your desk looks so tidy now!" and displays it on the device. If the user wants to continue tidying up, the message "Do you want to continue?" will be displayed, and if the user answers "yes," the next tidying up point (for example, instructions for tidying up the desk drawers) will be suggested. This process allows the user to gradually tidy up an entire room.

[0278] Example 2

[0279] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0280] The present invention aims to provide effective tidying support for people who have difficulty tidying up and for children who want to make tidying a habit. Another objective of the present invention is to realize a system that takes into account the user's emotional state to increase motivation to tidy up and encourage tidying up in a short amount of time.

[0281] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0282] In this invention, the server includes means for taking a video of the room, means for analyzing the taken video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for recognizing the user's emotional state, means for presenting the identified tidying up points and their procedures to the user, means for displaying a message of praise according to the user's emotional state after tidying up is completed, and means for suggesting further tidying up points to the user. This makes it possible to present tidying up points that can be completed in a short time while taking the user's emotional state into consideration, reducing resistance to tidying up and increasing motivation.

[0283] "Means for capturing video in a room" is a function that allows a user to capture video of the entire room using a device such as a smartphone or tablet.

[0284] "Means of analyzing the captured video to recognize objects and understand the state of the room" refers to a function that enables the server to use computer vision and machine learning algorithms to identify objects in the video and determine the current state of the room.

[0285] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" is a function that allows the server to select tidying up tasks that the user can complete within 15 minutes based on the analysis results.

[0286] The "means for recognizing the user's emotional state" is a function in which the server analyzes the user's facial expressions and voice data collected through the device's camera and microphone to determine the user's emotional state.

[0287] The "means for presenting the identified tidying up points and the procedures to the user" is a function for displaying the tidying up tasks selected by the server and the procedures for executing them on the user's terminal.

[0288] The "means for displaying a praising message according to the user's emotional state after tidying up is completed" is a function for displaying an appropriate praising message on the user's terminal after the tidying up task is completed, based on the emotional state analyzed by the server.

[0289] The "means for suggesting further tidying up points to the user" is a function for selecting and suggesting new tidying up tasks to a user who wants to continue tidying up.

[0290] "Means for sending quotation requests to be used for matching with housekeeping services" is a function for sending quotation requests to housekeeping service providers based on the video footage and the results of sentiment analysis.

[0291] This system presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[0292] Hardware and Software Configuration

[0293] This system consists of a device such as a smartphone or tablet, a processing server, and an emotion engine. The specific hardware and software used are as follows:

[0294] Devices: Smartphones and tablets equipped with cameras and microphones

[0295] Server: GPU-equipped server

[0296] software:

[0297] Video analysis and object recognition: TensorFlow, OpenCV

[0298] Emotion recognition engine: Microsoft Azure Emotion API or other emotion recognition API

[0299] How it works

[0300] 1. Recording and sending video in the room:

[0301] User: Use the device camera to record a video of the entire room. It is recommended to capture key areas such as the desk, floor, and shelves.

[0302] Device: Temporarily stores the captured video and then uploads it to the server.

[0303] 2. Video Analysis:

[0304] Server: Analyzes the received video using computer vision and machine learning algorithms (e.g., TensorFlow and OpenCV) to recognize objects in the room.

[0305] Example: "Recognize documents and pens on a desk, the position of a chair, and objects on the floor."

[0306] 3. Recognition of emotional states:

[0307] User: Speaks and shows facial expressions into the device's camera and microphone to capture emotional state.

[0308] Terminal: Sends camera images and audio to the server.

[0309] Server: The emotion engine analyzes facial expressions and tone of voice to determine the user's current emotional state. For example, if the user is smiling, it is determined that the user is "motivated."

[0310] 4. Generate cleanup points and procedures:

[0311] Server: Based on the analysis of the video and the user's emotional state, it generates tidying up points and specific steps that the user can complete within 15 minutes.

[0312] Example: Generate a procedure such as "Classify the documents on your desk into three categories (important, temporary storage, and disposal)" and send it to the terminal.

[0313] 5. Feedback Generation:

[0314] User: Follow the steps provided to tidy up. When finished, press the "Clean up" button in the app.

[0315] Terminal: Sends a completion notification to the server.

[0316] Server: Based on the results of emotion analysis by the emotion engine, a message of praise is generated and sent to the device. A message such as "You did a great job! Let's take a short break" is displayed.

[0317] 6. Suggestions for the following tidying points:

[0318] Server: If the user wants to continue cleaning, the server suggests new cleaning points. Taking into account the user's emotional state, the server presents slightly larger tasks if the emotional state is positive, and easier tasks if the emotional state is negative.

[0319] Example: "I suggest organizing your desk drawers as your next tidying point."

[0320] Matching housekeeping services

[0321] If the user feels that cleaning up by themselves is difficult, they can select the "Ask a professional" button to use a housekeeping service. In this case, the system operates as follows:

[0322] Terminal: Sends information about the room's state and the emotion engine's analysis results to the server.

[0323] Server: Sends a quote request to affiliated housekeeping service providers and sends the quote results to the terminal.

[0324] User: Can review the quote and decide to use the service.

[0325] Prompt Sentence Examples

[0326] An example of a prompt for a generative AI model is:

[0327] "After the user takes a video of their room and sends it to the server, use computer vision and machine learning to analyze the state of the room. Then, use an emotion engine to understand the user's emotional state, generate tidying points and steps that can be completed in 15 minutes, and return appropriate praise messages."

[0328] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0329] Step 1: Record and send your video

[0330] The user uses a device such as a smartphone or tablet to record a video of the entire room, making sure to capture key areas such as the desk, floor, and shelves.

[0331] Input: A video of the room taken by the user

[0332] Specific operation: The user lifts the device and moves the camera to capture the entire room while shooting video.

[0333] Output: Recorded video file

[0334] The device temporarily stores the captured video files and then uploads them to a server using Wi-Fi or mobile data.

[0335] Input: Video files stored on the device

[0336] Specific operation: The device saves the video file and sends it to the server via the network.

[0337] Output: Video file sent to the server

[0338] Step 2: Analyze the video

[0339] The server analyzes the received video using computer vision and machine learning algorithms (TensorFlow and OpenCV), recognizes objects in the room, and determines the current state of the room.

[0340] Input: Video file sent to the server

[0341] What it does: The server extracts frames from a video file, uses computer vision algorithms to detect objects, and uses machine learning models to predict the category of each object.

[0342] Output: Object recognition results in the room (object position, type)

[0343] Step 3: Recognizing your emotional state

[0344] The user's emotional state is captured by speaking into the device's camera and microphone and showing facial expressions.

[0345] Input: User's facial expression and voice data

[0346] Specific actions: The user smiles at the camera and makes a short comment out loud.

[0347] Output: Photographed facial expressions and recorded audio

[0348] The device transmits camera images and audio to the server.

[0349] Input: Photographed facial expressions and recorded audio data

[0350] Specific operation: The device captures these data and sends them to the server.

[0351] Output: Facial expression and voice data sent to the server

[0352] The server uses an emotion engine to analyze facial expressions and tone of voice to determine the user's current emotional state.

[0353] Input: Transmitted facial expression and voice data

[0354] Specific operation: The server extracts facial and vocal features using an emotion engine and classifies the emotional state based on these features.

[0355] Output: Emotional state (positive, negative, neutral, etc.)

[0356] Step 4: Generate cleanup points and procedures

[0357] Based on the results of object recognition and emotional state analysis in the video, the server generates tidying up points and specific steps that the user can complete within 15 minutes.

[0358] Input: Object recognition results in the room and emotional state

[0359] Specific operation: The server compares the analysis results and selects the optimal tidying task. For example, it generates a procedure such as "classify the documents on the desk into three categories (important, temporary storage, and disposal)."

[0360] Output: Cleaning points and procedures

[0361] Step 5: Clean up and notify completion

[0362] The user follows the displayed steps to tidy up, and when they are done, they press the "Tidy up complete" button displayed in the app on their device.

[0363] Input: Cleaning points and procedures

[0364] Specific actions: The user follows the instructions to tidy up and presses the "tidy up complete" button when finished.

[0365] Output: Tidy up completion notification

[0366] The terminal notifies the server that the "tidy up complete" button has been pressed.

[0367] Input: Cleanup completion notification

[0368] Specific operation: The terminal sends a completion notification to the server.

[0369] Output: Completion notification sent to the server

[0370] Step 6: Generate feedback

[0371] Based on the emotion analysis results from the emotion engine, the server generates a praising message that corresponds to the user's emotional state and sends it to the terminal.

[0372] Input: Tidy-up completion notification and emotional state

[0373] Specific behavior: The server generates an appropriate praise message based on the emotional state, for example, "Good job! Let's take a break."

[0374] Output: Complimentary message

[0375] Step 7: Suggest the next step

[0376] If the user wants to continue tidying up, the server will suggest new tidying up points taking into account the user's emotional state.

[0377] Input: Emotional state and intention to continue tidying up

[0378] Specific behavior: The server checks the user's emotional state and suggests a slightly larger task if it is positive, or an easier task if it is negative. For example, it suggests, "As the next step in tidying up, we suggest organizing your desk drawers."

[0379] Output: Suggested next steps and procedures

[0380] Matching housekeeping services

[0381] If the user finds it difficult to clean up by themselves, they can use a housekeeping service by selecting the "Ask a professional" button.

[0382] Input: Select the Ask a Pro button

[0383] Specific operation: When the "Ask a Pro" button is pressed, the device sends information about the room's status and the emotion engine's analysis results to the server.

[0384] Output: Room status information and sentiment analysis results sent to the server

[0385] The server sends a request for a quote to an affiliated housekeeping service provider and sends the quote result to the terminal.

[0386] Input: Room status information and sentiment analysis results

[0387] Specific operation: The server sends this information to the housekeeping service provider and receives the estimate.

[0388] Output: Estimate results sent to the user's device

[0389] The user checks the estimate and decides to use the service if necessary.

[0390] Input: Estimate result

[0391] Specific operation: The user checks the quote and decides whether to use the service.

[0392] Output: Decision to use the service or cancellation

[0393] (Application example 2)

[0394] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0395] Conventional tidying support systems lacked sufficient means to effectively tidy up while maintaining user motivation. Furthermore, in physical stores, it was difficult to efficiently organize inventory and optimize product placement. This increased resistance to tidying up, making it difficult to organize regularly.

[0396] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video of the room, means for analyzing the captured video to recognize objects and understand the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is completed, means for recognizing the user's emotional state and generating a message according to the emotional state, means for suggesting further tidying up points based on the emotional state, and means for applying the results to product placement and inventory management in a physical store. This enables effective tidying up while maintaining the user's motivation, and enables efficient management of inventory management and product placement in a physical store.

[0397] "Means for capturing video in a room" refers to a device or method for capturing video of the entire room or a specific area and acquiring the video data.

[0398] "Means for analyzing captured video to recognize objects and understand the state of the room" refers to technology that processes captured video data, identifies objects and their placement within the room, and understands the state of the room as a whole.

[0399] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method for using the results of video analysis to determine specific locations and items of tidying up work that can be completed in a short amount of time.

[0400] The "means for presenting the identified cleanup points and procedures to the user" refers to an interface or notification mechanism for providing the user with clear cleanup instructions and guidelines.

[0401] The "means for displaying a message praising the user after the completion of tidying up" refers to a system or function for displaying a message praising the efforts of a user who has completed tidying up.

[0402] "Means for recognizing the user's emotional state and generating a message according to that emotional state" refers to technology that analyzes the user's facial expressions and tone of voice to understand their emotional state and generate an appropriate message accordingly.

[0403] The "means for suggesting further tidying up points based on the emotional state" is a method for suggesting the next tidying up task to be done, taking into account the user's current emotional state.

[0404] "Means for application to product placement and inventory management in physical stores" refers to systems and methods for applying similar video analysis and tidying procedure presentation technologies to improve the efficiency of product placement and inventory management in physical stores.

[0405] The present invention is a system that realizes efficient management of inventory organization and product placement in physical stores. Its distinctive feature is that users can take videos of the store interior using smartphones or smart glasses, and the system analyzes the video data to suggest quick tidying tips and procedures. The specific system configuration and its implementation are described below.

[0406] System program generation

[0407] The system consists of the following main components:

[0408] 1. Camera device: A camera device for capturing video inside the store, such as a smartphone or smart glasses.

[0409] 2. Server: A central processing unit for video analysis and sentiment analysis.

[0410] 3. User interface: An application that presents cleaning points and procedures to the user.

[0411] Hardware and Software Description

[0412] Hardware: Smartphone, smart glasses (camera function)

[0413] These devices are used to capture images of the conditions inside the store.

[0414] software:

[0415] OpenCV: Used for video capture and processing.

[0416] Emotion Engine Module: Used to analyze the user's emotional state.

[0417] Object Detection Module: Used to perform object recognition and identify the location and type of product.

[0418] Server: As the central processing unit, it performs video analysis, object recognition, and emotion analysis.

[0419] Data processing and calculation flow

[0420] 1. The server receives videos taken by users using their smartphones or smart glasses.

[0421] 2. The received video data is analyzed using OpenCV, and objects within the store are recognized and their locations are determined.

[0422] 3. Use the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to determine their emotional state.

[0423] 4. Based on the results of object recognition and emotion analysis, the server generates tidying up points and steps that can be completed within 15 minutes and presents them to the user.

[0424] 5. After the user has completed tidying up, a message of praise will be displayed according to the user's emotional state, and if the user wishes to continue tidying up, new tidying points will be suggested.

[0425] Adding specific examples

[0426] As a concrete example, consider the following scenario.

[0427] Specific examples

[0428] Imagine a user tidying up a product shelf in a physical store. The user takes a video of the product shelf with their smartphone and sends the video data to a server. The server analyzes the video and recognizes the types of products on the shelf and their arrangement. The Emotion Engine analyzes the user's facial expressions and recognizes that the user is a little tired. The server then generates instructions such as "Organize the three columns on the right side of this shelf and rearrange the products by category," and displays these instructions on the smartphone. When the user follows these instructions to organize the shelves and presses the "Tidying up complete" button, the server displays a message of praise saying, "Good job! The shelves are now very clean!" If the user wants to continue tidying, the server displays the message "Do you want to continue?" and if the user answers "Yes," it suggests tidying up points that can be completed within the next 15 minutes.

[0429] Prompt Sentence Examples

[0430] "Based on your current emotional state, prioritize your inventory. Evaluate the condition of your shelves and tell us where to organize next."

[0431] As described above, the system of the present invention supports effective tidying up in physical stores while taking into account the emotional state of the user.

[0432] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0433] Step 1:

[0434] Users use smartphones or smart glasses to take videos of the inside of a store. This device is used to capture the inside of the store in detail. The input is video data that represents the current state of the store. The output is a video file that has been taken.

[0435] Step 2:

[0436] The device sends the captured video data to the server. Here, the device uses a stable communication environment to quickly upload the captured video to the server. The input is the video file and communication data. The output is the video data sent to the server.

[0437] Step 3:

[0438] The server analyzes the received video data. Using video analysis software such as OpenCV, it recognizes objects and identifies their layout in the store. This process uses computer vision technology. The input is the video data sent to the server. The output is the analysis results that identify the object's location and type.

[0439] Step 4:

[0440] The server uses the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to recognize their emotional state. This analysis uses facial recognition algorithms and voice analysis technology. The input is video data containing the user's facial expressions and vocal characteristics. The output is data indicating the user's current emotional state.

[0441] Step 5:

[0442] The server identifies tidying up points that can be completed within 15 minutes based on the results of object recognition and emotion analysis. Specifically, it selects areas that are easy to tackle and will have the greatest impact. The input is the results of object recognition and emotion analysis. The output is data on tidying up points and the steps involved.

[0443] Step 6:

[0444] The server sends the identified cleanup points and procedures to the user's terminal. The user then follows the procedures to carry out the cleanup work. The input is the data on the cleanup points and procedures. The output is the cleanup instructions displayed on the user's terminal.

[0445] Step 7:

[0446] When the user finishes cleaning up and presses the "Clean Up Complete" button, the terminal sends a completion notification to the server. The input is the user's completion operation. The output is the completion notification sent to the server.

[0447] Step 8:

[0448] The server again uses the Emotion Engine to analyze the user's emotional state. Based on the results, it generates a praising message to display after the user has tidied up. The input is the user's latest emotional data. The output is a message based on the user's emotional state.

[0449] Step 9:

[0450] The server suggests further tidying up points based on the user's emotional state. If the user is positive, it suggests bigger tasks, and if negative, it suggests easier tasks. The input is the user's emotional state and the current tidying up situation. The output is new tidying up points and procedure data.

[0451] Step 10:

[0452] The server generates information to support the optimization of product placement and inventory management in physical stores and provides it to managers and staff. This enables efficient in-store management. The input is data on the placement of objects in the store and the status of organization. The output is optimized placement plans and organization procedures.

[0453] The above are the processing steps of this system.

[0454] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0455] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0456] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0457] [Second embodiment]

[0458] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0459] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0460] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0461] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0462] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0463] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0464] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0465] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0466] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0467] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0468] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0469] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0470] The present invention provides an application system that makes it easier for users to tidy up their rooms, and an embodiment of the system is specifically described below. In particular, by presenting tidying up points that can be completed in a short time and showing the steps in an easy-to-understand manner, the system allows users to tackle tidying up without feeling any resistance.

[0471] This system allows users to use devices such as smartphones or tablets to take videos of their rooms and then analyze those videos. Specifically, the user launches the app and sends the video of the room's condition to the server. The server then analyzes the received video data, recognizes objects in the room, and determines their placement and the level of clutter. This analysis uses computer vision technology and machine learning algorithms.

[0472] Based on the video analysis results, the server identifies tidying up tasks that the user can complete within 15 minutes and generates instructions for doing so. For example, if the user is trying to tidy up documents scattered on a desk, the server will classify the documents by category and present instructions for putting them away in the appropriate storage location. This information is sent to the user's device, and the user can view the specific tidying up steps on the app screen.

[0473] After the user has finished tidying up, they press the "Tidy up complete" button in the app, and the device sends a completion notification to the server. The server receives this notification and generates a message praising the user. For example, the message might read, "Great job! Your desk is so tidy now!" If the user wants to continue tidying up, the option "Do you want to continue?" is presented. If the user answers "Yes," the server identifies new tidying points and sends the instructions to the device again. Through this series of steps, the user can gradually tidy up the entire room.

[0474] In addition, the system also provides a matching function for housekeeping services. When a user selects the "Ask a Professional" button within the app, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can review this quote and decide whether to use the service.

[0475] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The server then generates instructions such as "Classify documents into three categories (work, school, and personal) and put work-related documents in a red folder," and sends these to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying, "Great job! Your desk looks so tidy now!" If the user wants to continue tidying, the server asks, "Do you want to continue?". If the user answers "yes," it presents instructions for tidying up points (for example, inside the desk drawers) that can be completed within the next 15 minutes.

[0476] In this way, the system of the present invention provides the user with appropriate means and advice to help them tidy up their room effectively and efficiently, thereby reducing their resistance to tidying up and helping them develop good tidying habits.

[0477] The processing flow will be explained below.

[0478] Step 1:

[0479] The user launches the app using a smartphone or tablet and takes a video of the room.

[0480] Step 2:

[0481] The device stores the video data captured within the app and sends it to a server via the Internet.

[0482] Step 3:

[0483] The server receives the video data and the analysis module performs object recognition and classification within the video, using computer vision and machine learning algorithms.

[0484] Step 4:

[0485] Based on the analysis results, the server evaluates the arrangement and clutter of items and furniture in the room to grasp the overall condition.

[0486] Step 5:

[0487] The server identifies areas and tasks that can be completed within 15 minutes, and selects areas and tasks that are easy for users to complete in a short time.

[0488] Step 6:

[0489] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[0490] Step 7:

[0491] The server transmits the identified cleaning points and procedures to the terminal.

[0492] Step 8:

[0493] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[0494] Step 9:

[0495] The user follows the instructions to tidy up, and when tidying up is complete, presses the "tidy up complete" button.

[0496] Step 10:

[0497] The terminal sends a "tidying up completed" notification to the server.

[0498] Step 11:

[0499] The server receives the completion notification and generates a message praising the user.

[0500] Step 12:

[0501] A server-generated compliment message is sent to the device.

[0502] Step 13:

[0503] The terminal displays a complimentary message on the user interface.

[0504] Step 14:

[0505] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[0506] Step 15:

[0507] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[0508] Step 16:

[0509] The server checks the room status again and identifies areas that can be cleaned up within the next 15 minutes.

[0510] Step 17:

[0511] The server generates a new cleanup procedure and sends it to the terminal.

[0512] Step 18:

[0513] The terminal displays the new tidying up points and procedures received from the server on the user interface, and the user can confirm them and choose whether to tidy up again.

[0514] Optional Process:

[0515] Step A1:

[0516] The user selects the "Ask a Pro" button within the app.

[0517] Step A2:

[0518] The terminal transmits the room status information to the server.

[0519] Step A3:

[0520] Based on the received room status information, the server sends a quote request to an affiliated housekeeping service provider.

[0521] Step A4:

[0522] The housekeeping service provider creates a quote and sends it to the server.

[0523] Step A5:

[0524] The server sends the quote information to the terminal.

[0525] Step A6:

[0526] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[0527] Example 1

[0528] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0529] For many people, tidying up their rooms is a tedious and time-consuming task. This can lead to resistance or a sense of burden, making it difficult to keep a room clean. Even when hiring a professional, finding the right service can be a complicated process, resulting in cost and effort. This invention aims to solve these problems by providing a means for users to efficiently tidy up their rooms in a short amount of time.

[0530] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0531] In this invention, the server includes a means for recording images of the room, a means for analyzing the recorded images to identify objects and grasp the state of the room, and a means for identifying tidying up points that can be completed in a short time based on the analysis results. This allows the user to easily know the specific steps for tidying up the room, reducing the hassle. Furthermore, by matching with a housekeeping service, it is possible to reduce the effort required to hire a professional.

[0532] The "means for recording images within the room" refers to a camera function or a video recording function for recording the state of the room.

[0533] "Means for analyzing recorded images to identify objects and understand the state of a room" refers to methods that use computer vision techniques and machine learning algorithms to detect and identify objects from image data and understand the state of a room.

[0534] "Means for identifying tidying up points that can be completed in a short time based on the analysis results" refers to a method for selecting specific locations or items that a user can complete tidying up in a short time based on the analyzed image data.

[0535] "Means for displaying the identified tidying up points and the procedures to the user" refers to the functionality of a display or application for visually or textually showing the tidying up points and the procedures to the user.

[0536] The "means for displaying an acknowledgement message to the user after tidying up" refers to a method for displaying a message praising the user after tidying up is completed.

[0537] "Means for suggesting additional tidying up points to the user" refers to a method for suggesting additional tidying up tasks to the user.

[0538] "Means for sending a quote request to be used for matching with a housekeeping service" refers to a method for sending a quote request to a housekeeping service provider based on the condition of the room and matching with a user.

[0539] This invention is a system that allows a user to efficiently and effectively tidy up a room. Specifically, the system allows a user to take a video of the room using a device such as a smartphone or tablet, and analyzes the video to present tidying tips and procedures that the user can implement in a short amount of time. An embodiment of this system is described in detail below.

[0540] First, the user launches the application on their smartphone or tablet and records a video of the room's condition. This video is then sent from the device to the server. The server then analyzes the video data using computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow). The goal of the analysis is to identify objects in the room and understand the state of the room.

[0541] As a concrete example, consider the case where a user wants to tidy up their living room. The user uses a smartphone to record video of the living room and sends the video from the device to a server. The server processes the received video data and identifies objects in the living room (e.g., magazines, remote controls, cushions, etc.). Based on the analysis results, the server generates instructions such as "classify the magazines into three categories (fashion, sports, and news) and put them away in a specific place on the bookshelf." This instruction is then sent to the device and presented to the user.

[0542] The user follows the presented steps to tidy up. After completing the tidy up, the user presses the "Tidy up complete" button in the app. A completion notification is sent from the device to the server, and the server generates a message of praise for the user and displays it on the device. For example, a message such as "Great job! Your living room is now so tidy!" may be displayed.

[0543] If the user wants to continue tidying up, the system asks the user, "Do you want to continue?" If the user answers "yes," the server identifies the next tidying point and sends the instructions to the device again. For example, the next step might be, "Clean up the items on the living room table."

[0544] The system also provides a matching function with housekeeping services. When a user selects the "Ask a Professional" button, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The quote from the provider is sent to the device via the server, and the user can review it and decide whether to use the service.

[0545] An example of a prompt sentence is, "I took a picture of the state of my desk with my smartphone and sent it to the server. Please generate easy steps to tidy up the documents on my desk, and if necessary, please also suggest the next step to tidy up."

[0546] This allows users to gradually tidy up their rooms and easily use housekeeping services. The system aims to reduce users' resistance to tidying up and help them develop the habit of tidying up.

[0547] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0548] Step 1:

[0549] The user takes a video of the room.

[0550] A user uses a smartphone or tablet to record a video of the state of a room. For example, they can record a video so that magazines, remote controls, and other items in the living room are visible. The input is video data of the current state of the room, and the output is a video file saved in the device's storage.

[0551] Step 2:

[0552] The device sends the video to the server.

[0553] The captured video is uploaded from the device to the server. At this time, the device compresses the video data before transferring it to reduce communication delays. The input is the video file stored in the device's storage, and the output is the video data uploaded to the server.

[0554] Step 3:

[0555] The server analyzes the video and identifies the object.

[0556] The server uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow) to analyze the received video data. In this process, an object detection algorithm is used to identify objects in the video (e.g., magazines, remote controls, cushions, etc.). The input is the video data uploaded to the server, and the output is the object identification results (object type and location).

[0557] Step 4:

[0558] The server identifies the tidying up point based on the analysis results.

[0559] Based on the location information of the identified objects, the server identifies specific tidying up steps that the user can complete in a short time. For example, it generates a procedure such as "sort the magazines in the living room into three categories (fashion, sports, and news) and store them in a specific place on the bookshelf." The input is the object identification result, and the output is the identification of the tidying up steps (specific tidying up steps).

[0560] Step 5:

[0561] The server sends the generated cleanup procedure to the terminal.

[0562] The specific cleanup procedure created by the server is sent to the terminal. The input is the cleanup procedure generated by the server, and the output is the cleanup procedure sent to the terminal.

[0563] Step 6:

[0564] The user follows the cleaning procedure to perform the cleaning.

[0565] The user follows the instructions displayed on the device application to tidy up. For example, the user performs a specific action such as "sorting magazines into fashion, sports, and news, and storing them on the bookshelf." The input is the tidying up instructions displayed on the device, and the output is the physical state of the room after the tidying up is completed.

[0566] Step 7:

[0567] The user sends a cleanup completion notification to the server.

[0568] After the user has finished cleaning up, they press the "Clean Up Complete" button in the application. This causes a completion notification to be sent from the device to the server. The input is the user's operation, and the output is the completion notification sent to the server.

[0569] Step 8:

[0570] The server generates a completion message and sends it to the terminal.

[0571] The server receives the completion notification and generates a message praising the user. For example, it creates a message such as "Great job! Your living room is now so tidy!" and sends it to the terminal. The input is the completion notification, and the output is the completion message sent to the terminal.

[0572] Step 9:

[0573] The user chooses whether to clean up further.

[0574] The terminal presents the user with the option "Do you want to continue?" If the user answers "Yes," the procedure to suggest the next tidying up point begins. The input is the user's selection, and the output is the start of the suggestion of the next tidying up point.

[0575] Step 10:

[0576] The server identifies the next cleanup point.

[0577] If the user answers "yes," the server identifies the next step to clean up and generates instructions for it. For example, it creates a specific step such as "Clean up the items on the living room table" and sends it to the device. The input is the user's selection, and the output is the next step to clean up that was sent to the device.

[0578] Step 11:

[0579] If the user selects a housekeeping service, the server performs the matching.

[0580] When a user selects the "Ask a Professional" button in the application, the device sends the room's condition information to the server. Based on this information, the server sends a quote request to a housekeeping service provider and displays the quote from the provider to the user. The input is the room's condition information, and the output is the quote for the housekeeping service.

[0581] (Application example 1)

[0582] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0583] The present invention relates to an application system that allows users to easily tidy up their rooms. Conventional tidying support systems have the problem that when a user starts tidying up, it is unclear which object to start with, making it difficult to tidy up effectively. A similar problem exists in tidying up physical stores, where there is a lack of specific guidance for store staff to work efficiently. The present invention aims to solve these problems.

[0584] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0585] In this invention, the server includes means for taking video of the room, means for analyzing the video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is complete, means for suggesting further tidying up points to the user, means for taking video of the tidying up state of the physical store, and means for presenting the tidying up procedures to store staff. This enables users and store staff to specifically and efficiently tidy up and organize in a short amount of time.

[0586] "Means for capturing video of the inside of a room" refers to equipment or a method for capturing video of the inside of a room using a camera and acquiring the video data.

[0587] "Means for analyzing captured video to recognize objects and grasp the state of a room" refers to equipment and methods that use video analysis technology to identify items present in a room based on acquired video data and grasp the overall situation of the room.

[0588] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method or algorithm that selects and identifies tidying up tasks that a user can complete in a short amount of time based on the results of video analysis.

[0589] "Means for presenting the identified tidying up points and procedures to the user" refers to equipment or methods for displaying or notifying the user of the identified work content and specific procedures so that the user can tidy up efficiently.

[0590] The "means for displaying a message praising the user after tidying up is completed" is a means for presenting a message praising the user's efforts when the user has finished tidying up.

[0591] The "means for suggesting further tidying up points to the user" is a means for presenting new tidying up tasks to the user when the user wants to continue tidying up.

[0592] "Means for photographing the tidiness and neatness of a physical store" refers to equipment or a method for using a camera to photograph the state of items and shelves in a physical store and obtain the video data.

[0593] "Means for presenting tidying procedures to store staff" refers to devices or methods for displaying or informing store staff of specified work content and specific procedures so that they can efficiently tidy up items.

[0594] The present invention provides a system that allows users to efficiently tidy up and organize their rooms and physical stores. Specific embodiments of the system will be described below.

[0595] This system uses devices such as smartphones and tablets to record footage of the state of a room or physical store, and then analyzes the video data to present cleaning and tidying procedures. Users use their devices to record video of the room or store and send the video to a server. The server analyzes the video data and uses computer vision technology and machine learning algorithms to recognize objects. Libraries such as OpenCV and TensorFlow can be used for this analysis.

[0596] Based on the results of the video analysis, the server identifies tidying up tasks that the user can complete within 15 minutes and automatically generates detailed instructions. The generated instructions are sent to the user's device, showing them exactly how to proceed with the tidying up. For example, specific steps such as "sort documents by category and put work-related documents in a red folder" are presented.

[0597] The user checks these steps through the app screen and performs the tidying task. After completing the tidying, the user presses the "Tidying up complete" button on the device, and the device sends a notification to the server. The server receives this notification and generates and displays a message of praise to the user, such as "Great job! Your desk is so tidy now!" If the user wants to continue tidying, it presents the option "Do you want to continue?". If the user answers "Yes," the server identifies new tidying points, generates new steps, and sends them to the device.

[0598] In the case of physical stores, the photographing and analysis procedures using the device are similar. The server identifies the state of organization of the physical store and presents specific organizational procedures to store staff. For example, it presents procedures for sorting products that are randomly placed on shelves and tidying them up by category. This allows store staff to work efficiently.

[0599] The hardware used is a smartphone, tablet, and server, and the software uses OpenCV for video analysis and TensorFlow for machine learning algorithms.

[0600] As a concrete example, if magazines are placed in a disorganized manner in a storeroom, a store clerk can use a smartphone to take a video of the situation and analyze the video. Based on the analysis results, a procedure will be generated, such as "sort the magazines by title and organize them on specific shelves." This allows the store clerk to follow the instructions and efficiently organize the magazines.

[0601] An example of a prompt for a generative AI model might be, "Please take a survey for an application that takes photos of the state of shelves in a physical store and generates the types of items and the procedures for organizing them."

[0602] This system allows users and store staff to tidy up and organize in a specific and efficient manner in a short amount of time.

[0603] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0604] Step 1:

[0605] Users use devices such as smartphones or tablets to shoot video of a room or a physical store.

[0606] Input: Video data captured using the device's camera function.

[0607] Output: The captured video file.

[0608] This video file will be used for later analysis.

[0609] Step 2:

[0610] The device sends the captured video to the server.

[0611] Input: Video files stored on your device.

[0612] Output: Video data sent to the server over the network.

[0613] The server receives this data and prepares it for analysis.

[0614] Step 3:

[0615] The server analyzes the received video data, recognizes objects, and understands the condition of the room or physical store.

[0616] Input: Video data sent to the server.

[0617] Output: Object recognition results in the video and room / store status information.

[0618] This analysis uses libraries such as OpenCV and TensorFlow to identify the location and type of object.

[0619] Step 4:

[0620] Based on the results of the video analysis, the server identifies tidying up points that the user can complete within 15 minutes.

[0621] Input: Object recognition results and room / store state information.

[0622] Output: Identified cleanup points and specific cleanup procedures.

[0623] Specific instructions are automatically generated, detailing which items to organize and how.

[0624] Step 5:

[0625] The server transmits the identified tidying up points and their procedures to the terminal and presents them to the user.

[0626] Input: Server-generated cleanup instructions.

[0627] Output: Cleanup instructions displayed on the user's terminal.

[0628] By following this procedure, the user can efficiently tidy up.

[0629] Step 6:

[0630] After the user has finished tidying up, he / she presses the "tidying up complete" button on the terminal.

[0631] Input: Cleanup completion notification action from user.

[0632] Output: Tidy-up completion notification sent to the server.

[0633] This notification lets the server know the progress of the cleanup.

[0634] Step 7:

[0635] The server receives the notification that the tidying up is complete, generates a message praising the user, such as "good job," and sends it to the terminal.

[0636] Input: Notification from user that cleanup is complete.

[0637] Output: A compliment message displayed on the terminal.

[0638] This gives the user a sense of accomplishment.

[0639] Step 8:

[0640] The server asks the user whether he / she wants to continue cleaning up, and if he / she does, it identifies a new cleaning point, generates a procedure, and sends it to the terminal.

[0641] Input: User confirms their desire to continue tidying.

[0642] Output: New cleanup procedure.

[0643] This allows the user to move on to the next tidying task.

[0644] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0645] This invention provides a system that presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[0646] The system begins when a user takes a video of their room using a device such as a smartphone or tablet and sends the video to the server. The server analyzes the received video to recognize objects and determine the state of the room. This analysis is carried out using computer vision and machine learning algorithms. Based on the results of this analysis, the server identifies tidying up points that the user can complete within 15 minutes and generates a procedure for doing so.

[0647] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotional state. This emotion engine analyzes the user's facial expressions and tone of voice via a camera and microphone to grasp the user's current emotional state. This allows the system to take the user's emotional state into consideration when presenting tidying points and procedures.

[0648] When a user tidies up through the app and presses the "Tidy up complete" button after completing the task, the device sends a completion notification to the server. The server generates a message of praise according to the user's emotional state based on the results of emotion analysis by the emotion engine. For example, if the user is tired, a message such as "You did a great job! Let's take a break" will be displayed.

[0649] Furthermore, if the user chooses to continue tidying up, the system also takes their emotional state into account when suggesting the next step: if the user is in a positive emotional state, it may suggest a larger or more complex task, while if the user is in a negative emotional state, it may suggest an easier task.

[0650] The system also provides a matching function for housekeeping services. When the user selects the "Ask a Professional" button, the device sends information about the room's condition and the results of analysis by the emotion engine to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can then review the quote and decide whether to use the service.

[0651] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The emotion engine also analyzes the user's facial expressions and recognizes that the user is motivated. The server then generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder" and sends them to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying "Good job! Your desk looks so tidy now!" If the user wants to continue tidying, the system asks "Do you want to continue?". If the user answers "yes," the system suggests tidying up points that can be completed within the next 15 minutes (for example, tasks related to the desk drawers).

[0652] In this way, the system of the present invention provides tidying up advice that takes into account the user's emotional state, thereby reducing resistance to tidying up and helping to form tidying up habits. Furthermore, by facilitating the use of housekeeping services, it provides more efficient tidying up support.

[0653] The processing flow will be explained below.

[0654] Step 1:

[0655] The user launches the app using a smartphone or tablet and takes a video of the room.

[0656] Step 2:

[0657] The device stores the video data captured within the app and sends it to a server via the Internet.

[0658] Step 3:

[0659] The server receives the video data and the analysis module recognizes and classifies objects in the video using computer vision and machine learning algorithms. The server then determines the location, type, and overall condition of the recognized objects.

[0660] Step 4:

[0661] Based on the analysis results, the server identifies tidying up points that the user can complete within 15 minutes, and also selects areas and tasks that the user can easily accomplish in a short amount of time.

[0662] Step 5:

[0663] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[0664] Step 6:

[0665] The emotion engine analyzes the user's facial expressions and tone of voice through the device's camera and microphone to recognize the user's emotional state.

[0666] Step 7:

[0667] The server then adjusts the cleaning procedure and approach based on the emotional state of the user based on the analysis results of the emotion engine. For example, it adds positive comments to users in a positive emotional state and encouraging comments to users in a negative emotional state.

[0668] Step 8:

[0669] The server transmits the adjusted cleaning points and procedures to the terminal based on the emotional state.

[0670] Step 9:

[0671] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[0672] Step 10:

[0673] The user follows the instructions to tidy up, and when finished, presses the "tidy up complete" button.

[0674] Step 11:

[0675] The terminal sends a "tidying up completed" notification to the server.

[0676] Step 12:

[0677] The server generates a praising message according to the user's emotional state based on the emotion analysis results of the emotion engine.

[0678] Step 13:

[0679] A server-generated compliment message is sent to the device.

[0680] Step 14:

[0681] The device will display a praising message on the user interface. For example, if the user is tired, it will say "Great job! Let's take a break."

[0682] Step 15:

[0683] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[0684] Step 16:

[0685] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[0686] Step 17:

[0687] The server checks the room status again and identifies new cleanup points that can be completed within 15 minutes.

[0688] Step 18:

[0689] The server uses an emotion engine to analyze the user's current emotional state.

[0690] Step 19:

[0691] The server generates a new cleaning procedure based on the emotional state and sends it to the terminal.

[0692] Step 20:

[0693] The terminal displays the new tidying up points and procedures on the user interface, and the user can confirm this and choose whether to tidy up again.

[0694] Optional Process:

[0695] Step A1:

[0696] The user selects the "Ask a Pro" button within the app.

[0697] Step A2:

[0698] The device sends the room status and the emotion engine's analysis results to the server.

[0699] Step A3:

[0700] Based on the information received by the server, a request for an estimate is sent to an affiliated housekeeping service provider.

[0701] Step A4:

[0702] The housekeeping service provider creates a quote and sends it to the server.

[0703] Step A5:

[0704] The server sends the quote information to the terminal.

[0705] Step A6:

[0706] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[0707] As a concrete example, if a user wants to tidy up the documents on their desk, the process would proceed as follows: The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and understands the state of the desk. The emotion engine recognizes that the user is motivated. The server generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder," and sends these instructions to the device. When the user completes the tidying up by following the instructions and presses the "Tidying up complete" button, the server generates a message saying "Great job! Your desk looks so tidy now!" and displays it on the device. If the user wants to continue tidying up, the message "Do you want to continue?" will be displayed, and if the user answers "yes," the next tidying up point (for example, instructions for tidying up the desk drawers) will be suggested. This process allows the user to gradually tidy up an entire room.

[0708] Example 2

[0709] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0710] The present invention aims to provide effective tidying support for people who have difficulty tidying up and for children who want to make tidying a habit. Another objective of the present invention is to realize a system that takes into account the user's emotional state to increase motivation to tidy up and encourage tidying up in a short amount of time.

[0711] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0712] In this invention, the server includes means for taking a video of the room, means for analyzing the taken video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for recognizing the user's emotional state, means for presenting the identified tidying up points and their procedures to the user, means for displaying a message of praise according to the user's emotional state after tidying up is completed, and means for suggesting further tidying up points to the user. This makes it possible to present tidying up points that can be completed in a short time while taking the user's emotional state into consideration, reducing resistance to tidying up and increasing motivation.

[0713] "Means for capturing video in a room" is a function that allows a user to capture video of the entire room using a device such as a smartphone or tablet.

[0714] "Means of analyzing the captured video to recognize objects and understand the state of the room" refers to a function that enables the server to use computer vision and machine learning algorithms to identify objects in the video and determine the current state of the room.

[0715] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" is a function that enables the server to select tidying up tasks that the user can complete within 15 minutes based on the analysis results.

[0716] The "means for recognizing the user's emotional state" is a function in which the server analyzes the user's facial expressions and voice data collected through the device's camera and microphone to determine the user's emotional state.

[0717] The "means for presenting the identified tidying up points and the procedures to the user" is a function for displaying the tidying up tasks selected by the server and the procedures for executing them on the user's terminal.

[0718] The "means for displaying a praising message according to the user's emotional state after tidying up is completed" is a function for displaying an appropriate praising message on the user's terminal after the tidying up task is completed, based on the emotional state analyzed by the server.

[0719] The "means for suggesting further tidying up points to the user" is a function for selecting and suggesting new tidying up tasks to a user who wants to continue tidying up.

[0720] "Means for sending quotation requests to be used for matching with housekeeping services" is a function for sending quotation requests to housekeeping service providers based on the video footage and the results of sentiment analysis.

[0721] This system presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[0722] Hardware and Software Configuration

[0723] This system consists of a device such as a smartphone or tablet, a processing server, and an emotion engine. The specific hardware and software used are as follows:

[0724] Devices: Smartphones and tablets equipped with cameras and microphones

[0725] Server: GPU-equipped server

[0726] software:

[0727] Video analysis and object recognition: TensorFlow, OpenCV

[0728] Emotion recognition engine: Microsoft Azure Emotion API or other emotion recognition API

[0729] How it works

[0730] 1. Recording and sending video in the room:

[0731] User: Use the device camera to record a video of the entire room. It is recommended to capture key areas such as the desk, floor, and shelves.

[0732] Device: Temporarily stores the captured video and then uploads it to the server.

[0733] 2. Video Analysis:

[0734] Server: Analyzes the received video using computer vision and machine learning algorithms (e.g., TensorFlow and OpenCV) to recognize objects in the room.

[0735] Example: "Recognize documents and pens on a desk, the position of a chair, and objects on the floor."

[0736] 3. Recognition of emotional states:

[0737] User: Speaks and shows facial expressions into the device's camera and microphone to capture emotional state.

[0738] Terminal: Sends camera images and audio to the server.

[0739] Server: The emotion engine analyzes facial expressions and tone of voice to determine the user's current emotional state. For example, if the user is smiling, it is determined that the user is "motivated."

[0740] 4. Generate cleanup points and procedures:

[0741] Server: Based on the analysis of the video and the user's emotional state, it generates tidying up points and specific steps that the user can complete within 15 minutes.

[0742] Example: Generate a procedure such as "Classify the documents on your desk into three categories (important, temporary storage, and disposal)" and send it to the terminal.

[0743] 5. Feedback Generation:

[0744] User: Follow the steps provided to tidy up. When finished, press the "Clean up" button in the app.

[0745] Terminal: Sends a completion notification to the server.

[0746] Server: Based on the results of emotion analysis by the emotion engine, a message of praise is generated and sent to the device. A message such as "You did a great job! Let's take a short break" is displayed.

[0747] 6. Suggestions for the following tidying points:

[0748] Server: If the user wants to continue cleaning, the server suggests new cleaning points. Taking into account the user's emotional state, the server presents slightly larger tasks if the emotional state is positive, and easier tasks if the emotional state is negative.

[0749] Example: "I suggest organizing your desk drawers as your next tidying point."

[0750] Matching housekeeping services

[0751] If the user feels that cleaning up by themselves is difficult, they can select the "Ask a professional" button to use a housekeeping service. In this case, the system operates as follows:

[0752] Terminal: Sends information about the room's state and the emotion engine's analysis results to the server.

[0753] Server: Sends a quote request to affiliated housekeeping service providers and sends the quote results to the terminal.

[0754] User: Can review the quote and decide to use the service.

[0755] Prompt Sentence Examples

[0756] An example of a prompt for a generative AI model is:

[0757] "After the user takes a video of their room and sends it to the server, use computer vision and machine learning to analyze the state of the room. Then, use an emotion engine to understand the user's emotional state, generate tidying points and steps that can be completed in 15 minutes, and return appropriate praise messages."

[0758] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0759] Step 1: Record and send your video

[0760] The user uses a device such as a smartphone or tablet to record a video of the entire room, making sure to capture key areas such as the desk, floor, and shelves.

[0761] Input: A video of the room taken by the user

[0762] Specific operation: The user lifts the device and moves the camera to capture the entire room while shooting video.

[0763] Output: Recorded video file

[0764] The device temporarily stores the captured video files and then uploads them to a server using Wi-Fi or mobile data.

[0765] Input: Video files stored on the device

[0766] Specific operation: The device saves the video file and sends it to the server via the network.

[0767] Output: Video file sent to the server

[0768] Step 2: Analyze the video

[0769] The server analyzes the received video using computer vision and machine learning algorithms (TensorFlow and OpenCV), recognizes objects in the room, and determines the current state of the room.

[0770] Input: Video file sent to the server

[0771] What it does: The server extracts frames from a video file, uses computer vision algorithms to detect objects, and uses machine learning models to predict the category of each object.

[0772] Output: Object recognition results in the room (object position, type)

[0773] Step 3: Recognizing your emotional state

[0774] The user's emotional state is captured by speaking into the device's camera and microphone and showing facial expressions.

[0775] Input: User's facial expression and voice data

[0776] Specific actions: The user smiles at the camera and makes a short comment out loud.

[0777] Output: Photographed facial expressions and recorded audio

[0778] The device transmits camera images and audio to the server.

[0779] Input: Photographed facial expressions and recorded audio data

[0780] Specific operation: The device captures these data and sends them to the server.

[0781] Output: Facial expression and voice data sent to the server

[0782] The server uses an emotion engine to analyze facial expressions and tone of voice to determine the user's current emotional state.

[0783] Input: Transmitted facial expression and voice data

[0784] Specific operation: The server extracts facial and vocal features using an emotion engine and classifies the emotional state based on these features.

[0785] Output: Emotional state (positive, negative, neutral, etc.)

[0786] Step 4: Generate cleanup points and procedures

[0787] Based on the results of object recognition and emotional state analysis in the video, the server generates tidying up points and specific steps that the user can complete within 15 minutes.

[0788] Input: Object recognition results in the room and emotional state

[0789] Specific operation: The server compares the analysis results and selects the optimal tidying task. For example, it generates a procedure such as "classify the documents on the desk into three categories (important, temporary storage, and disposal)."

[0790] Output: Cleaning points and procedures

[0791] Step 5: Clean up and notify completion

[0792] The user follows the displayed steps to tidy up, and when they are done, they press the "Tidy up complete" button displayed in the app on their device.

[0793] Input: Cleaning points and procedures

[0794] Specific actions: The user follows the instructions to tidy up and presses the "tidy up complete" button when finished.

[0795] Output: Tidy up completion notification

[0796] The terminal notifies the server that the "tidy up complete" button has been pressed.

[0797] Input: Cleanup completion notification

[0798] Specific operation: The terminal sends a completion notification to the server.

[0799] Output: Completion notification sent to the server

[0800] Step 6: Generate feedback

[0801] Based on the emotion analysis results from the emotion engine, the server generates a praising message that corresponds to the user's emotional state and sends it to the terminal.

[0802] Input: Tidy-up completion notification and emotional state

[0803] Specific behavior: The server generates an appropriate praise message based on the emotional state, for example, "Good job! Let's take a break."

[0804] Output: Complimentary message

[0805] Step 7: Suggest the next step

[0806] If the user wants to continue tidying up, the server will suggest new tidying up points taking into account the user's emotional state.

[0807] Input: Emotional state and intention to continue tidying up

[0808] Specific behavior: The server checks the user's emotional state and suggests a slightly larger task if it is positive, or an easier task if it is negative. For example, it suggests, "As the next step in tidying up, we suggest organizing your desk drawers."

[0809] Output: Suggested next steps and procedures

[0810] Matching housekeeping services

[0811] If the user finds it difficult to clean up by themselves, they can use a housekeeping service by selecting the "Ask a professional" button.

[0812] Input: Select the Ask a Pro button

[0813] Specific operation: When the "Ask a Pro" button is pressed, the device sends information about the room's status and the emotion engine's analysis results to the server.

[0814] Output: Room status information and sentiment analysis results sent to the server

[0815] The server sends a request for a quote to an affiliated housekeeping service provider and sends the quote result to the terminal.

[0816] Input: Room status information and sentiment analysis results

[0817] Specific operation: The server sends this information to the housekeeping service provider and receives the estimate.

[0818] Output: Estimate results sent to the user's device

[0819] The user checks the estimate and decides to use the service if necessary.

[0820] Input: Estimate result

[0821] Specific operation: The user checks the quote and decides whether to use the service.

[0822] Output: Decision to use the service or cancellation

[0823] (Application example 2)

[0824] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0825] Conventional tidying support systems lacked sufficient means to effectively tidy up while maintaining user motivation. Furthermore, in physical stores, it was difficult to efficiently organize inventory and optimize product placement. This increased resistance to tidying up, making it difficult to organize regularly.

[0826] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video of the room, means for analyzing the captured video to recognize objects and understand the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is completed, means for recognizing the user's emotional state and generating a message according to the emotional state, means for suggesting further tidying up points based on the emotional state, and means for applying the results to product placement and inventory management in a physical store. This enables effective tidying up while maintaining the user's motivation, and enables efficient management of inventory management and product placement in a physical store.

[0827] "Means for capturing video in a room" refers to a device or method for capturing video of the entire room or a specific area and acquiring the video data.

[0828] "Means for analyzing captured video to recognize objects and understand the state of the room" refers to technology that processes captured video data, identifies objects and their placement within the room, and understands the state of the room as a whole.

[0829] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method for using the results of video analysis to determine specific locations and items of tidying up work that can be completed in a short amount of time.

[0830] The "means for presenting the identified cleanup points and procedures to the user" refers to an interface or notification mechanism for providing the user with clear cleanup instructions and guidelines.

[0831] The "means for displaying a message praising the user after the completion of tidying up" refers to a system or function for displaying a message praising the efforts of a user who has completed tidying up.

[0832] "Means for recognizing the user's emotional state and generating a message according to that emotional state" refers to technology that analyzes the user's facial expressions and tone of voice to understand their emotional state and generate an appropriate message accordingly.

[0833] The "means for suggesting further tidying up points based on the emotional state" is a method for suggesting the next tidying up task to be done, taking into account the user's current emotional state.

[0834] "Means for application to product placement and inventory management in physical stores" refers to systems and methods for applying similar video analysis and tidying procedure presentation technologies to improve the efficiency of product placement and inventory management in physical stores.

[0835] The present invention is a system that realizes efficient management of inventory organization and product placement in physical stores. Its distinctive feature is that users can take videos of the store interior using smartphones or smart glasses, and the system analyzes the video data to suggest quick tidying tips and procedures. The specific system configuration and its implementation are described below.

[0836] System program generation

[0837] The system consists of the following main components:

[0838] 1. Camera device: A camera device for capturing video inside the store, such as a smartphone or smart glasses.

[0839] 2. Server: A central processing unit for video analysis and sentiment analysis.

[0840] 3. User interface: An application that presents cleaning points and procedures to the user.

[0841] Hardware and Software Description

[0842] Hardware: Smartphone, smart glasses (camera function)

[0843] These devices are used to capture images of the conditions inside the store.

[0844] software:

[0845] OpenCV: Used for video capture and processing.

[0846] Emotion Engine Module: Used to analyze the user's emotional state.

[0847] Object Detection Module: Used to perform object recognition and identify the location and type of product.

[0848] Server: As the central processing unit, it performs video analysis, object recognition, and emotion analysis.

[0849] Data processing and calculation flow

[0850] 1. The server receives videos taken by users using their smartphones or smart glasses.

[0851] 2. The received video data is analyzed using OpenCV, and objects within the store are recognized and their locations are determined.

[0852] 3. Use the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to determine their emotional state.

[0853] 4. Based on the results of object recognition and emotion analysis, the server generates tidying up points and steps that can be completed within 15 minutes and presents them to the user.

[0854] 5. After the user has completed tidying up, a message of praise will be displayed according to the user's emotional state, and if the user wishes to continue tidying up, new tidying points will be suggested.

[0855] Adding specific examples

[0856] As a concrete example, consider the following scenario.

[0857] Specific examples

[0858] Imagine a user tidying up a product shelf in a physical store. The user takes a video of the product shelf with their smartphone and sends the video data to a server. The server analyzes the video and recognizes the types of products on the shelf and their arrangement. The Emotion Engine analyzes the user's facial expressions and recognizes that the user is a little tired. The server then generates instructions such as "Organize the three columns on the right side of this shelf and rearrange the products by category," and displays these instructions on the smartphone. When the user follows these instructions to organize the shelves and presses the "Tidying up complete" button, the server displays a message of praise saying, "Good job! The shelves are now very clean!" If the user wants to continue tidying, the server displays the message "Do you want to continue?" and if the user answers "Yes," it suggests tidying up points that can be completed within the next 15 minutes.

[0859] Prompt Sentence Examples

[0860] "Based on your current emotional state, prioritize your inventory. Evaluate the condition of your shelves and tell us where to organize next."

[0861] As described above, the system of the present invention supports effective tidying up in physical stores while taking into account the emotional state of the user.

[0862] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0863] Step 1:

[0864] Users use smartphones or smart glasses to take videos of the inside of a store. This device is used to capture the inside of the store in detail. The input is video data that represents the current state of the store. The output is a video file that has been taken.

[0865] Step 2:

[0866] The device sends the captured video data to the server. Here, the device uses a stable communication environment to quickly upload the captured video to the server. The input is the video file and communication data. The output is the video data sent to the server.

[0867] Step 3:

[0868] The server analyzes the received video data. Using video analysis software such as OpenCV, it recognizes objects and identifies their layout in the store. This process uses computer vision technology. The input is the video data sent to the server. The output is the analysis results that identify the object's location and type.

[0869] Step 4:

[0870] The server uses the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to recognize their emotional state. This analysis uses facial recognition algorithms and voice analysis technology. The input is video data containing the user's facial expressions and vocal characteristics. The output is data indicating the user's current emotional state.

[0871] Step 5:

[0872] The server identifies tidying up points that can be completed within 15 minutes based on the results of object recognition and emotion analysis. Specifically, it selects areas that are easy to tackle and will have the greatest impact. The input is the results of object recognition and emotion analysis. The output is data on tidying up points and the steps involved.

[0873] Step 6:

[0874] The server sends the identified cleanup points and procedures to the user's terminal. The user then follows the procedures to carry out the cleanup work. The input is the data on the cleanup points and procedures. The output is the cleanup instructions displayed on the user's terminal.

[0875] Step 7:

[0876] When the user finishes cleaning up and presses the "Clean Up Complete" button, the terminal sends a completion notification to the server. The input is the user's completion operation. The output is the completion notification sent to the server.

[0877] Step 8:

[0878] The server again uses the Emotion Engine to analyze the user's emotional state. Based on the results, it generates a praising message to display after the user has tidied up. The input is the user's latest emotional data. The output is a message based on the user's emotional state.

[0879] Step 9:

[0880] The server suggests further tidying up points based on the user's emotional state. If the user is positive, it suggests bigger tasks, and if negative, it suggests easier tasks. The input is the user's emotional state and the current tidying up situation. The output is new tidying up points and procedure data.

[0881] Step 10:

[0882] The server generates information to support the optimization of product placement and inventory management in physical stores and provides it to managers and staff. This enables efficient in-store management. The input is data on the placement of objects in the store and the status of organization. The output is optimized placement plans and organization procedures.

[0883] The above are the processing steps of this system.

[0884] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0885] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0886] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0887] [Third embodiment]

[0888] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0889] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0890] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0891] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0892] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0893] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0894] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0895] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0896] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0897] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0898] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0899] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0900] The present invention provides an application system that makes it easier for users to tidy up their rooms, and an embodiment of the system is specifically described below. In particular, by presenting tidying up points that can be completed in a short time and showing the steps in an easy-to-understand manner, the system allows users to tackle tidying up without feeling any resistance.

[0901] This system allows users to use devices such as smartphones or tablets to take videos of their rooms and then analyze those videos. Specifically, the user launches the app and sends the video of the room's condition to the server. The server then analyzes the received video data, recognizes objects in the room, and determines their placement and the level of clutter. This analysis uses computer vision technology and machine learning algorithms.

[0902] Based on the video analysis results, the server identifies tidying up tasks that the user can complete within 15 minutes and generates instructions for doing so. For example, if the user is trying to tidy up documents scattered on a desk, the server will classify the documents by category and present instructions for putting them away in the appropriate storage location. This information is sent to the user's device, and the user can view the specific tidying up steps on the app screen.

[0903] After the user has finished tidying up, they press the "Tidy up complete" button in the app, and the device sends a completion notification to the server. The server receives this notification and generates a message praising the user. For example, the message might read, "Great job! Your desk is so tidy now!" If the user wants to continue tidying up, the option "Do you want to continue?" is presented. If the user answers "Yes," the server identifies new tidying points and sends the instructions to the device again. Through this series of steps, the user can gradually tidy up the entire room.

[0904] In addition, the system also provides a matching function for housekeeping services. When a user selects the "Ask a Professional" button within the app, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can review this quote and decide whether to use the service.

[0905] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The server then generates instructions such as "Classify documents into three categories (work, school, and personal) and put work-related documents in a red folder," and sends these to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying, "Great job! Your desk looks so tidy now!" If the user wants to continue tidying, the server asks, "Do you want to continue?". If the user answers "yes," it presents instructions for tidying up points (for example, inside the desk drawers) that can be completed within the next 15 minutes.

[0906] In this way, the system of the present invention provides the user with appropriate means and advice to help them tidy up their room effectively and efficiently, thereby reducing their resistance to tidying up and helping them develop good tidying habits.

[0907] The processing flow will be explained below.

[0908] Step 1:

[0909] The user launches the app using a smartphone or tablet and takes a video of the room.

[0910] Step 2:

[0911] The device stores the video data captured within the app and sends it to a server via the Internet.

[0912] Step 3:

[0913] The server receives the video data and the analysis module performs object recognition and classification within the video, using computer vision and machine learning algorithms.

[0914] Step 4:

[0915] Based on the analysis results, the server evaluates the arrangement and clutter of items and furniture in the room to grasp the overall condition.

[0916] Step 5:

[0917] The server identifies areas and tasks that can be completed within 15 minutes, and selects areas and tasks that are easy for users to complete in a short time.

[0918] Step 6:

[0919] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[0920] Step 7:

[0921] The server transmits the identified cleaning points and procedures to the terminal.

[0922] Step 8:

[0923] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[0924] Step 9:

[0925] The user follows the instructions to tidy up, and when tidying up is complete, presses the "tidy up complete" button.

[0926] Step 10:

[0927] The terminal sends a "tidying up completed" notification to the server.

[0928] Step 11:

[0929] The server receives the completion notification and generates a message praising the user.

[0930] Step 12:

[0931] A server-generated compliment message is sent to the device.

[0932] Step 13:

[0933] The terminal displays a complimentary message on the user interface.

[0934] Step 14:

[0935] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[0936] Step 15:

[0937] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[0938] Step 16:

[0939] The server checks the room status again and identifies areas that can be cleaned up within the next 15 minutes.

[0940] Step 17:

[0941] The server generates a new cleanup procedure and sends it to the terminal.

[0942] Step 18:

[0943] The terminal displays the new tidying up points and procedures received from the server on the user interface, and the user can confirm them and choose whether to tidy up again.

[0944] Optional Process:

[0945] Step A1:

[0946] The user selects the "Ask a Pro" button within the app.

[0947] Step A2:

[0948] The terminal transmits the room status information to the server.

[0949] Step A3:

[0950] Based on the received room status information, the server sends a quote request to an affiliated housekeeping service provider.

[0951] Step A4:

[0952] The housekeeping service provider creates a quote and sends it to the server.

[0953] Step A5:

[0954] The server sends the quote information to the terminal.

[0955] Step A6:

[0956] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[0957] Example 1

[0958] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0959] For many people, tidying up their rooms is a tedious and time-consuming task. This can lead to resistance or a sense of burden, making it difficult to keep a room clean. Even when hiring a professional, finding the right service can be a complicated process, resulting in cost and effort. This invention aims to solve these problems by providing a means for users to efficiently tidy up their rooms in a short amount of time.

[0960] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0961] In this invention, the server includes a means for recording images of the room, a means for analyzing the recorded images to identify objects and grasp the state of the room, and a means for identifying tidying up points that can be completed in a short time based on the analysis results. This allows the user to easily know the specific steps for tidying up the room, reducing the hassle. Furthermore, by matching with a housekeeping service, it is possible to reduce the effort required to hire a professional.

[0962] The "means for recording images within the room" refers to a camera function or a video recording function for recording the state of the room.

[0963] "Means for analyzing recorded images to identify objects and understand the state of a room" refers to methods that use computer vision techniques and machine learning algorithms to detect and identify objects from image data and understand the state of a room.

[0964] "Means for identifying tidying up points that can be completed in a short time based on the analysis results" refers to a method for selecting specific locations or items that a user can complete tidying up in a short time based on the analyzed image data.

[0965] "Means for displaying the identified tidying up points and the procedures to the user" refers to the functionality of a display or application for visually or textually showing the tidying up points and the procedures to the user.

[0966] The "means for displaying an acknowledgement message to the user after tidying up" refers to a method for displaying a message praising the user after tidying up is completed.

[0967] "Means for suggesting additional tidying up points to the user" refers to a method for suggesting additional tidying up tasks to the user.

[0968] "Means for sending a quote request to be used for matching with a housekeeping service" refers to a method for sending a quote request to a housekeeping service provider based on the condition of the room and matching with a user.

[0969] This invention is a system that allows a user to efficiently and effectively tidy up a room. Specifically, the system allows a user to take a video of the room using a device such as a smartphone or tablet, and analyzes the video to present tidying tips and procedures that the user can implement in a short amount of time. An embodiment of this system is described in detail below.

[0970] First, the user launches the application on their smartphone or tablet and records a video of the room's condition. This video is then sent from the device to the server. The server then analyzes the video data using computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow). The goal of the analysis is to identify objects in the room and understand the state of the room.

[0971] As a concrete example, consider the case where a user wants to tidy up their living room. The user uses a smartphone to record video of the living room and sends the video from the device to a server. The server processes the received video data and identifies objects in the living room (e.g., magazines, remote controls, cushions, etc.). Based on the analysis results, the server generates instructions such as "classify the magazines into three categories (fashion, sports, and news) and put them away in a specific place on the bookshelf." This instruction is then sent to the device and presented to the user.

[0972] The user follows the presented steps to tidy up. After completing the tidy up, the user presses the "Tidy up complete" button in the app. A completion notification is sent from the device to the server, and the server generates a message of praise for the user and displays it on the device. For example, a message such as "Great job! Your living room is now so tidy!" may be displayed.

[0973] If the user wants to continue tidying up, the system asks the user, "Do you want to continue?" If the user answers "yes," the server identifies the next tidying point and sends the instructions to the device again. For example, the next step might be, "Clean up the items on the living room table."

[0974] The system also provides a matching function with housekeeping services. When a user selects the "Ask a Professional" button, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The quote from the provider is sent to the device via the server, and the user can review it and decide whether to use the service.

[0975] An example of a prompt sentence is, "I took a picture of the state of my desk with my smartphone and sent it to the server. Please generate easy steps to tidy up the documents on my desk, and if necessary, please also suggest the next step to tidy up."

[0976] This allows users to gradually tidy up their rooms and easily use housekeeping services. The system aims to reduce users' resistance to tidying up and help them develop the habit of tidying up.

[0977] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0978] Step 1:

[0979] The user takes a video of the room.

[0980] A user uses a smartphone or tablet to record a video of the state of a room. For example, they can record a video so that magazines, remote controls, and other items in the living room are visible. The input is video data of the current state of the room, and the output is a video file saved in the device's storage.

[0981] Step 2:

[0982] The device sends the video to the server.

[0983] The captured video is uploaded from the device to the server. At this time, the device compresses the video data before transferring it to reduce communication delays. The input is the video file stored in the device's storage, and the output is the video data uploaded to the server.

[0984] Step 3:

[0985] The server analyzes the video and identifies the object.

[0986] The server uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow) to analyze the received video data. In this process, an object detection algorithm is used to identify objects in the video (e.g., magazines, remote controls, cushions, etc.). The input is the video data uploaded to the server, and the output is the object identification results (object type and location).

[0987] Step 4:

[0988] The server identifies the tidying up point based on the analysis results.

[0989] Based on the location information of the identified objects, the server identifies specific tidying up steps that the user can complete in a short time. For example, it generates a procedure such as "sort the magazines in the living room into three categories (fashion, sports, and news) and store them in a specific place on the bookshelf." The input is the object identification result, and the output is the identification of the tidying up steps (specific tidying up steps).

[0990] Step 5:

[0991] The server sends the generated cleanup procedure to the terminal.

[0992] The specific cleanup procedure created by the server is sent to the terminal. The input is the cleanup procedure generated by the server, and the output is the cleanup procedure sent to the terminal.

[0993] Step 6:

[0994] The user follows the cleaning procedure to perform the cleaning.

[0995] The user follows the instructions displayed on the device application to tidy up. For example, the user performs a specific action such as "sorting magazines into fashion, sports, and news, and storing them on the bookshelf." The input is the tidying up instructions displayed on the device, and the output is the physical state of the room after the tidying up is completed.

[0996] Step 7:

[0997] The user sends a cleanup completion notification to the server.

[0998] After the user has finished cleaning up, they press the "Clean Up Complete" button in the application. This causes a completion notification to be sent from the device to the server. The input is the user's operation, and the output is the completion notification sent to the server.

[0999] Step 8:

[1000] The server generates a completion message and sends it to the terminal.

[1001] The server receives the completion notification and generates a message praising the user. For example, it creates a message such as "Great job! Your living room is now so tidy!" and sends it to the terminal. The input is the completion notification, and the output is the completion message sent to the terminal.

[1002] Step 9:

[1003] The user chooses whether to clean up further.

[1004] The terminal presents the user with the option "Do you want to continue?" If the user answers "Yes," the procedure to suggest the next tidying up point begins. The input is the user's selection, and the output is the start of the suggestion of the next tidying up point.

[1005] Step 10:

[1006] The server identifies the next cleanup point.

[1007] If the user answers "yes," the server identifies the next step to clean up and generates instructions for it. For example, it creates a specific step such as "Clean up the items on the living room table" and sends it to the device. The input is the user's selection, and the output is the next step to clean up that was sent to the device.

[1008] Step 11:

[1009] If the user selects a housekeeping service, the server performs the matching.

[1010] When a user selects the "Ask a Professional" button in the application, the device sends the room's condition information to the server. Based on this information, the server sends a quote request to a housekeeping service provider and displays the quote from the provider to the user. The input is the room's condition information, and the output is the quote for the housekeeping service.

[1011] (Application example 1)

[1012] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1013] The present invention relates to an application system that allows users to easily tidy up their rooms. Conventional tidying support systems have the problem that when a user starts tidying up, it is unclear which object to start with, making it difficult to tidy up effectively. A similar problem exists in tidying up physical stores, where there is a lack of specific guidance for store staff to work efficiently. The present invention aims to solve these problems.

[1014] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1015] In this invention, the server includes means for taking video of the room, means for analyzing the video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is complete, means for suggesting further tidying up points to the user, means for taking video of the tidying up state of the physical store, and means for presenting the tidying up procedures to store staff. This enables users and store staff to specifically and efficiently tidy up and organize in a short amount of time.

[1016] "Means for capturing video of the inside of a room" refers to equipment or a method for capturing video of the inside of a room using a camera and acquiring the video data.

[1017] "Means for analyzing captured video to recognize objects and grasp the state of a room" refers to equipment and methods that use video analysis technology to identify items present in a room based on acquired video data and grasp the overall situation of the room.

[1018] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method or algorithm that selects and identifies tidying up tasks that a user can complete in a short amount of time based on the results of video analysis.

[1019] "Means for presenting the identified tidying up points and procedures to the user" refers to equipment or methods for displaying or notifying the user of the identified work content and specific procedures so that the user can tidy up efficiently.

[1020] The "means for displaying a message praising the user after tidying up is completed" is a means for presenting a message praising the user's efforts when the user has finished tidying up.

[1021] The "means for suggesting further tidying up points to the user" is a means for presenting new tidying up tasks to the user when the user wants to continue tidying up.

[1022] "Means for photographing the tidiness and neatness of a physical store" refers to equipment or a method for using a camera to photograph the state of items and shelves in a physical store and obtain the video data.

[1023] "Means for presenting tidying procedures to store staff" refers to devices or methods for displaying or informing store staff of specified work content and specific procedures so that they can efficiently tidy up items.

[1024] The present invention provides a system that allows users to efficiently tidy up and organize their rooms and physical stores. Specific embodiments of the system will be described below.

[1025] This system uses devices such as smartphones and tablets to record footage of the state of a room or physical store, and then analyzes the video data to present cleaning and tidying procedures. Users use their devices to record video of the room or store and send the video to a server. The server analyzes the video data and uses computer vision technology and machine learning algorithms to recognize objects. Libraries such as OpenCV and TensorFlow can be used for this analysis.

[1026] Based on the results of the video analysis, the server identifies tidying up tasks that the user can complete within 15 minutes and automatically generates detailed instructions. The generated instructions are sent to the user's device, showing them exactly how to proceed with the tidying up. For example, specific steps such as "sort documents by category and put work-related documents in a red folder" are presented.

[1027] The user checks these steps through the app screen and performs the tidying task. After completing the tidying, the user presses the "Tidying up complete" button on the device, and the device sends a notification to the server. The server receives this notification and generates and displays a message of praise to the user, such as "Great job! Your desk is so tidy now!" If the user wants to continue tidying, it presents the option "Do you want to continue?". If the user answers "Yes," the server identifies new tidying points, generates new steps, and sends them to the device.

[1028] In the case of physical stores, the photographing and analysis procedures using the device are similar. The server identifies the state of organization of the physical store and presents specific organizational procedures to store staff. For example, it presents procedures for sorting products that are randomly placed on shelves and tidying them up by category. This allows store staff to work efficiently.

[1029] The hardware used is a smartphone, tablet, and server, and the software uses OpenCV for video analysis and TensorFlow for machine learning algorithms.

[1030] As a concrete example, if magazines are placed in a disorganized manner in a storeroom, a store clerk can use a smartphone to take a video of the situation and analyze the video. Based on the analysis results, a procedure will be generated, such as "sort the magazines by title and organize them on specific shelves." This allows the store clerk to follow the instructions and efficiently organize the magazines.

[1031] An example of a prompt for a generative AI model might be, "Please take a survey for an application that takes photos of the state of shelves in a physical store and generates the types of items and the procedures for organizing them."

[1032] This system allows users and store staff to tidy up and organize in a specific and efficient manner in a short amount of time.

[1033] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1034] Step 1:

[1035] Users use devices such as smartphones or tablets to shoot video of a room or a physical store.

[1036] Input: Video data captured using the device's camera function.

[1037] Output: The captured video file.

[1038] This video file will be used for later analysis.

[1039] Step 2:

[1040] The device sends the captured video to the server.

[1041] Input: Video files stored on your device.

[1042] Output: Video data sent to the server over the network.

[1043] The server receives this data and prepares it for analysis.

[1044] Step 3:

[1045] The server analyzes the received video data, recognizes objects, and understands the condition of the room or physical store.

[1046] Input: Video data sent to the server.

[1047] Output: Object recognition results in the video and room / store status information.

[1048] This analysis uses libraries such as OpenCV and TensorFlow to identify the location and type of object.

[1049] Step 4:

[1050] Based on the results of the video analysis, the server identifies tidying up points that the user can complete within 15 minutes.

[1051] Input: Object recognition results and room / store state information.

[1052] Output: Identified cleanup points and specific cleanup procedures.

[1053] Specific instructions are automatically generated, detailing which items to organize and how.

[1054] Step 5:

[1055] The server transmits the identified tidying up points and their procedures to the terminal and presents them to the user.

[1056] Input: Server-generated cleanup instructions.

[1057] Output: Cleanup instructions displayed on the user's terminal.

[1058] By following this procedure, the user can efficiently tidy up.

[1059] Step 6:

[1060] After the user has finished tidying up, he / she presses the "tidying up complete" button on the terminal.

[1061] Input: Cleanup completion notification action from user.

[1062] Output: Tidy-up completion notification sent to the server.

[1063] This notification lets the server know the progress of the cleanup.

[1064] Step 7:

[1065] The server receives the notification that the tidying up is complete, generates a message praising the user, such as "good job," and sends it to the terminal.

[1066] Input: Notification from user that cleanup is complete.

[1067] Output: A compliment message displayed on the terminal.

[1068] This gives the user a sense of accomplishment.

[1069] Step 8:

[1070] The server asks the user whether he / she wants to continue cleaning up, and if he / she does, it identifies a new cleaning point, generates a procedure, and sends it to the terminal.

[1071] Input: User confirms their desire to continue tidying.

[1072] Output: New cleanup procedure.

[1073] This allows the user to move on to the next tidying task.

[1074] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1075] This invention provides a system that presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[1076] The system begins when a user takes a video of their room using a device such as a smartphone or tablet and sends the video to the server. The server analyzes the received video to recognize objects and determine the state of the room. This analysis is carried out using computer vision and machine learning algorithms. Based on the results of this analysis, the server identifies tidying up points that the user can complete within 15 minutes and generates a procedure for doing so.

[1077] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotional state. This emotion engine analyzes the user's facial expressions and tone of voice via a camera and microphone to grasp the user's current emotional state. This allows the system to take the user's emotional state into consideration when presenting tidying points and procedures.

[1078] When a user tidies up through the app and presses the "Tidy up complete" button after completing the task, the device sends a completion notification to the server. The server generates a message of praise according to the user's emotional state based on the results of emotion analysis by the emotion engine. For example, if the user is tired, a message such as "You did a great job! Let's take a break" will be displayed.

[1079] Furthermore, if the user chooses to continue tidying up, the system also takes their emotional state into account when suggesting the next step: if the user is in a positive emotional state, it may suggest a larger or more complex task, while if the user is in a negative emotional state, it may suggest an easier task.

[1080] The system also provides a matching function for housekeeping services. When the user selects the "Ask a Professional" button, the device sends information about the room's condition and the results of analysis by the emotion engine to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can then review the quote and decide whether to use the service.

[1081] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The emotion engine also analyzes the user's facial expressions and recognizes that the user is motivated. The server then generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder" and sends them to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying "Good job! Your desk looks so tidy now!" If the user wants to continue tidying, the system asks "Do you want to continue?". If the user answers "yes," the system suggests tidying up points that can be completed within the next 15 minutes (for example, tasks related to the desk drawers).

[1082] In this way, the system of the present invention provides tidying up advice that takes into account the user's emotional state, thereby reducing resistance to tidying up and helping to form tidying up habits. Furthermore, by facilitating the use of housekeeping services, the system provides more efficient tidying up support.

[1083] The processing flow will be explained below.

[1084] Step 1:

[1085] The user launches the app using a smartphone or tablet and takes a video of the room.

[1086] Step 2:

[1087] The device stores the video data captured within the app and sends it to a server via the Internet.

[1088] Step 3:

[1089] The server receives the video data and the analysis module recognizes and classifies objects in the video using computer vision and machine learning algorithms. The server then determines the location, type, and overall condition of the recognized objects.

[1090] Step 4:

[1091] Based on the analysis results, the server identifies tidying up points that the user can complete within 15 minutes, and also selects areas and tasks that the user can easily accomplish in a short amount of time.

[1092] Step 5:

[1093] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[1094] Step 6:

[1095] The emotion engine analyzes the user's facial expressions and tone of voice through the device's camera and microphone to recognize the user's emotional state.

[1096] Step 7:

[1097] The server then adjusts the cleaning procedure and approach based on the emotional state of the user based on the analysis results of the emotion engine. For example, it adds positive comments to users in a positive emotional state and encouraging comments to users in a negative emotional state.

[1098] Step 8:

[1099] The server transmits the adjusted cleaning points and procedures to the terminal based on the emotional state.

[1100] Step 9:

[1101] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[1102] Step 10:

[1103] The user follows the instructions to tidy up, and when finished, presses the "tidy up complete" button.

[1104] Step 11:

[1105] The terminal sends a "tidying up completed" notification to the server.

[1106] Step 12:

[1107] The server generates a praising message according to the user's emotional state based on the emotion analysis results of the emotion engine.

[1108] Step 13:

[1109] A server-generated compliment message is sent to the device.

[1110] Step 14:

[1111] The device will display a praising message on the user interface. For example, if the user is tired, it will say "Great job! Let's take a break."

[1112] Step 15:

[1113] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[1114] Step 16:

[1115] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[1116] Step 17:

[1117] The server checks the room status again and identifies new cleanup points that can be completed within 15 minutes.

[1118] Step 18:

[1119] The server uses an emotion engine to analyze the user's current emotional state.

[1120] Step 19:

[1121] The server generates a new cleaning procedure based on the emotional state and sends it to the terminal.

[1122] Step 20:

[1123] The terminal displays the new tidying up points and procedures on the user interface, and the user can confirm this and choose whether to tidy up again.

[1124] Optional Process:

[1125] Step A1:

[1126] The user selects the "Ask a Pro" button within the app.

[1127] Step A2:

[1128] The device sends the room status and the emotion engine's analysis results to the server.

[1129] Step A3:

[1130] Based on the information received by the server, a request for an estimate is sent to an affiliated housekeeping service provider.

[1131] Step A4:

[1132] The housekeeping service provider creates a quote and sends it to the server.

[1133] Step A5:

[1134] The server sends the quote information to the terminal.

[1135] Step A6:

[1136] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[1137] As a concrete example, if a user wants to tidy up the documents on their desk, the process would proceed as follows: The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and understands the state of the desk. The emotion engine recognizes that the user is motivated. The server generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder," and sends these instructions to the device. When the user completes the tidying up by following the instructions and presses the "Tidying up complete" button, the server generates a message saying "Great job! Your desk looks so tidy now!" and displays it on the device. If the user wants to continue tidying up, the message "Do you want to continue?" will be displayed, and if the user answers "yes," the next tidying up point (for example, instructions for tidying up the desk drawers) will be suggested. This process allows the user to gradually tidy up an entire room.

[1138] Example 2

[1139] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1140] The present invention aims to provide effective tidying support for people who have difficulty tidying up and for children who want to make tidying a habit. Another objective of the present invention is to realize a system that takes into account the user's emotional state to increase motivation to tidy up and encourage tidying up in a short amount of time.

[1141] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1142] In this invention, the server includes means for taking a video of the room, means for analyzing the taken video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for recognizing the user's emotional state, means for presenting the identified tidying up points and their procedures to the user, means for displaying a message of praise according to the user's emotional state after tidying up is completed, and means for suggesting further tidying up points to the user. This makes it possible to present tidying up points that can be completed in a short time while taking the user's emotional state into consideration, reducing resistance to tidying up and increasing motivation.

[1143] "Means for capturing video in a room" is a function that allows a user to capture video of the entire room using a device such as a smartphone or tablet.

[1144] "Means of analyzing the captured video to recognize objects and understand the state of the room" refers to a function that enables the server to use computer vision and machine learning algorithms to identify objects in the video and determine the current state of the room.

[1145] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" is a function that enables the server to select tidying up tasks that the user can complete within 15 minutes based on the analysis results.

[1146] The "means for recognizing the user's emotional state" is a function in which the server analyzes the user's facial expressions and voice data collected through the device's camera and microphone to determine the user's emotional state.

[1147] The "means for presenting the identified tidying up points and the procedures to the user" is a function for displaying the tidying up tasks selected by the server and the procedures for executing them on the user's terminal.

[1148] The "means for displaying a praising message according to the user's emotional state after tidying up is completed" is a function for displaying an appropriate praising message on the user's terminal after the tidying up task is completed, based on the emotional state analyzed by the server.

[1149] The "means for suggesting further tidying up points to the user" is a function for selecting and suggesting new tidying up tasks to a user who wants to continue tidying up.

[1150] "Means for sending quotation requests to be used for matching with housekeeping services" is a function for sending quotation requests to housekeeping service providers based on the video footage and the results of sentiment analysis.

[1151] This system presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[1152] Hardware and Software Configuration

[1153] This system consists of a device such as a smartphone or tablet, a processing server, and an emotion engine. The specific hardware and software used are as follows:

[1154] Devices: Smartphones and tablets equipped with cameras and microphones

[1155] Server: GPU-equipped server

[1156] software:

[1157] Video analysis and object recognition: TensorFlow, OpenCV

[1158] Emotion recognition engine: Microsoft Azure Emotion API or other emotion recognition API

[1159] How it works

[1160] 1. Recording and sending video in the room:

[1161] User: Use the device camera to record a video of the entire room. It is recommended to capture key areas such as the desk, floor, and shelves.

[1162] Device: Temporarily stores the captured video and then uploads it to the server.

[1163] 2. Video Analysis:

[1164] Server: Analyzes the received video using computer vision and machine learning algorithms (e.g., TensorFlow and OpenCV) to recognize objects in the room.

[1165] Example: "Recognize documents and pens on a desk, the position of a chair, and objects on the floor."

[1166] 3. Recognition of emotional states:

[1167] User: Speaks and shows facial expressions into the device's camera and microphone to capture emotional state.

[1168] Terminal: Sends camera images and audio to the server.

[1169] Server: The emotion engine analyzes facial expressions and tone of voice to determine the user's current emotional state. For example, if the user is smiling, it is determined that the user is "motivated."

[1170] 4. Generate cleanup points and procedures:

[1171] Server: Based on the analysis of the video and the user's emotional state, it generates tidying up points and specific steps that the user can complete within 15 minutes.

[1172] Example: Generate a procedure such as "Classify the documents on your desk into three categories (important, temporary storage, and disposal)" and send it to the terminal.

[1173] 5. Feedback Generation:

[1174] User: Follow the steps provided to tidy up. When finished, press the "Clean up" button in the app.

[1175] Terminal: Sends a completion notification to the server.

[1176] Server: Based on the results of emotion analysis by the emotion engine, a message of praise is generated and sent to the device. A message such as "You did a great job! Let's take a short break" is displayed.

[1177] 6. Suggestions for the following tidying points:

[1178] Server: If the user wants to continue cleaning, the server suggests new cleaning points. Taking into account the user's emotional state, the server presents slightly larger tasks if the emotional state is positive, and easier tasks if the emotional state is negative.

[1179] Example: "I suggest organizing your desk drawers as your next tidying point."

[1180] Matching housekeeping services

[1181] If the user feels that cleaning up by themselves is difficult, they can select the "Ask a professional" button to use a housekeeping service. In this case, the system operates as follows:

[1182] Terminal: Sends information about the room's state and the emotion engine's analysis results to the server.

[1183] Server: Sends a quote request to affiliated housekeeping service providers and sends the quote results to the terminal.

[1184] User: Can review the quote and decide to use the service.

[1185] Prompt Sentence Examples

[1186] An example of a prompt for a generative AI model is:

[1187] "After the user takes a video of their room and sends it to the server, use computer vision and machine learning to analyze the state of the room. Then, use an emotion engine to understand the user's emotional state, generate tidying points and steps that can be completed in 15 minutes, and return appropriate praise messages."

[1188] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1189] Step 1: Record and send your video

[1190] The user uses a device such as a smartphone or tablet to record a video of the entire room, making sure to capture key areas such as the desk, floor, and shelves.

[1191] Input: A video of the room taken by the user

[1192] Specific operation: The user lifts the device and moves the camera to capture the entire room while shooting video.

[1193] Output: Recorded video file

[1194] The device temporarily stores the captured video files and then uploads them to a server using Wi-Fi or mobile data.

[1195] Input: Video files stored on the device

[1196] Specific operation: The device saves the video file and sends it to the server via the network.

[1197] Output: Video file sent to the server

[1198] Step 2: Analyze the video

[1199] The server analyzes the received video using computer vision and machine learning algorithms (TensorFlow and OpenCV), recognizes objects in the room, and determines the current state of the room.

[1200] Input: Video file sent to the server

[1201] What it does: The server extracts frames from a video file, uses computer vision algorithms to detect objects, and uses machine learning models to predict the category of each object.

[1202] Output: Object recognition results in the room (object position, type)

[1203] Step 3: Recognizing your emotional state

[1204] The user's emotional state is captured by speaking into the device's camera and microphone and showing facial expressions.

[1205] Input: User's facial expression and voice data

[1206] Specific actions: The user smiles at the camera and makes a short comment out loud.

[1207] Output: Photographed facial expressions and recorded audio

[1208] The device transmits camera images and audio to the server.

[1209] Input: Photographed facial expressions and recorded audio data

[1210] Specific operation: The device captures these data and sends them to the server.

[1211] Output: Facial expression and voice data sent to the server

[1212] The server uses an emotion engine to analyze facial expressions and tone of voice to determine the user's current emotional state.

[1213] Input: Transmitted facial expression and voice data

[1214] Specific operation: The server extracts facial and vocal features using an emotion engine and classifies the emotional state based on these features.

[1215] Output: Emotional state (positive, negative, neutral, etc.)

[1216] Step 4: Generate cleanup points and procedures

[1217] Based on the results of object recognition and emotional state analysis in the video, the server generates tidying up points and specific steps that the user can complete within 15 minutes.

[1218] Input: Object recognition results in the room and emotional state

[1219] Specific operation: The server compares the analysis results and selects the optimal tidying task. For example, it generates a procedure such as "classify the documents on the desk into three categories (important, temporary storage, and disposal)."

[1220] Output: Cleaning points and procedures

[1221] Step 5: Clean up and notify completion

[1222] The user follows the displayed steps to tidy up, and when they are done, they press the "Tidy up complete" button displayed in the app on their device.

[1223] Input: Cleaning points and procedures

[1224] Specific actions: The user follows the instructions to tidy up and presses the "tidy up complete" button when finished.

[1225] Output: Tidy up completion notification

[1226] The terminal notifies the server that the "tidy up complete" button has been pressed.

[1227] Input: Cleanup completion notification

[1228] Specific operation: The terminal sends a completion notification to the server.

[1229] Output: Completion notification sent to the server

[1230] Step 6: Generate feedback

[1231] Based on the emotion analysis results from the emotion engine, the server generates a praising message that corresponds to the user's emotional state and sends it to the terminal.

[1232] Input: Tidy-up completion notification and emotional state

[1233] Specific behavior: The server generates an appropriate praise message based on the emotional state, for example, "Good job! Let's take a break."

[1234] Output: Complimentary message

[1235] Step 7: Suggest the next step

[1236] If the user wants to continue tidying up, the server will suggest new tidying up points taking into account the user's emotional state.

[1237] Input: Emotional state and intention to continue tidying up

[1238] Specific behavior: The server checks the user's emotional state and suggests a slightly larger task if it is positive, or an easier task if it is negative. For example, it suggests, "As the next step in tidying up, we suggest organizing your desk drawers."

[1239] Output: Suggested next steps and procedures

[1240] Matching housekeeping services

[1241] If the user finds it difficult to clean up by themselves, they can use a housekeeping service by selecting the "Ask a professional" button.

[1242] Input: Select the Ask a Pro button

[1243] Specific operation: When the "Ask a Pro" button is pressed, the device sends information about the room's status and the emotion engine's analysis results to the server.

[1244] Output: Room status information and sentiment analysis results sent to the server

[1245] The server sends a request for a quote to an affiliated housekeeping service provider and sends the quote result to the terminal.

[1246] Input: Room status information and sentiment analysis results

[1247] Specific operation: The server sends this information to the housekeeping service provider and receives the estimate.

[1248] Output: Estimate results sent to the user's device

[1249] The user checks the estimate and decides to use the service if necessary.

[1250] Input: Estimate result

[1251] Specific operation: The user checks the quote and decides whether to use the service.

[1252] Output: Decision to use the service or cancellation

[1253] (Application example 2)

[1254] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1255] Conventional tidying support systems lacked sufficient means to effectively tidy up while maintaining user motivation. Furthermore, in physical stores, it was difficult to efficiently organize inventory and optimize product placement. This increased resistance to tidying up, making it difficult to organize regularly.

[1256] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video of the room, means for analyzing the captured video to recognize objects and understand the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is completed, means for recognizing the user's emotional state and generating a message according to the emotional state, means for suggesting further tidying up points based on the emotional state, and means for applying the results to product placement and inventory management in a physical store. This enables effective tidying up while maintaining the user's motivation, and enables efficient management of inventory management and product placement in a physical store.

[1257] "Means for capturing video in a room" refers to a device or method for capturing video of the entire room or a specific area and acquiring the video data.

[1258] "Means for analyzing captured video to recognize objects and understand the state of the room" refers to technology that processes captured video data, identifies objects and their placement within the room, and understands the state of the room as a whole.

[1259] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method for using the results of video analysis to determine specific locations and items of tidying up work that can be completed in a short amount of time.

[1260] The "means for presenting the identified cleanup points and procedures to the user" refers to an interface or notification mechanism for providing the user with clear cleanup instructions and guidelines.

[1261] The "means for displaying a message praising the user after the completion of tidying up" refers to a system or function for displaying a message praising the efforts of a user who has completed tidying up.

[1262] "Means for recognizing the user's emotional state and generating a message according to that emotional state" refers to technology that analyzes the user's facial expressions and tone of voice to understand their emotional state and generate an appropriate message accordingly.

[1263] The "means for suggesting further tidying up points based on the emotional state" is a method for suggesting the next tidying up task to be done, taking into account the user's current emotional state.

[1264] "Means for application to product placement and inventory management in physical stores" refers to systems and methods for applying similar video analysis and tidying procedure presentation technologies to improve the efficiency of product placement and inventory management in physical stores.

[1265] The present invention is a system that realizes efficient management of inventory organization and product placement in physical stores. Its distinctive feature is that users can take videos of the store interior using smartphones or smart glasses, and the system analyzes the video data to suggest quick tidying tips and procedures. The specific system configuration and its implementation are described below.

[1266] System program generation

[1267] The system consists of the following main components:

[1268] 1. Camera device: A camera device for capturing video inside the store, such as a smartphone or smart glasses.

[1269] 2. Server: A central processing unit for video analysis and sentiment analysis.

[1270] 3. User interface: An application that presents cleaning points and procedures to the user.

[1271] Hardware and Software Description

[1272] Hardware: Smartphone, smart glasses (camera function)

[1273] These devices are used to capture images of the conditions inside the store.

[1274] software:

[1275] OpenCV: Used for video capture and processing.

[1276] Emotion Engine Module: Used to analyze the user's emotional state.

[1277] Object Detection Module: Used to perform object recognition and identify the location and type of product.

[1278] Server: As the central processing unit, it performs video analysis, object recognition, and emotion analysis.

[1279] Data processing and calculation flow

[1280] 1. The server receives videos taken by users using their smartphones or smart glasses.

[1281] 2. The received video data is analyzed using OpenCV, and objects within the store are recognized and their locations are determined.

[1282] 3. Use the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to determine their emotional state.

[1283] 4. Based on the results of object recognition and emotion analysis, the server generates tidying up points and steps that can be completed within 15 minutes and presents them to the user.

[1284] 5. After the user has completed tidying up, a message of praise will be displayed according to the user's emotional state, and if the user wishes to continue tidying up, new tidying points will be suggested.

[1285] Adding specific examples

[1286] As a concrete example, consider the following scenario.

[1287] Specific examples

[1288] Imagine a user tidying up a product shelf in a physical store. The user takes a video of the product shelf with their smartphone and sends the video data to a server. The server analyzes the video and recognizes the types of products on the shelf and their arrangement. The Emotion Engine analyzes the user's facial expressions and recognizes that the user is a little tired. The server then generates instructions such as "Organize the three columns on the right side of this shelf and rearrange the products by category," and displays these instructions on the smartphone. When the user follows these instructions to organize the shelves and presses the "Tidying up complete" button, the server displays a message of praise saying, "Good job! The shelves are now very clean!" If the user wants to continue tidying, the server displays the message "Do you want to continue?" and if the user answers "Yes," it suggests tidying up points that can be completed within the next 15 minutes.

[1289] Prompt Sentence Examples

[1290] "Based on your current emotional state, prioritize your inventory. Evaluate the condition of your shelves and tell us where to organize next."

[1291] As described above, the system of the present invention supports effective tidying up in physical stores while taking into account the emotional state of the user.

[1292] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1293] Step 1:

[1294] Users use smartphones or smart glasses to take videos of the inside of a store. This device is used to capture the inside of the store in detail. The input is video data that represents the current state of the store. The output is a video file that has been taken.

[1295] Step 2:

[1296] The device sends the captured video data to the server. Here, the device uses a stable communication environment to quickly upload the captured video to the server. The input is the video file and communication data. The output is the video data sent to the server.

[1297] Step 3:

[1298] The server analyzes the received video data. Using video analysis software such as OpenCV, it recognizes objects and identifies their layout in the store. This process uses computer vision technology. The input is the video data sent to the server. The output is the analysis results that identify the object's location and type.

[1299] Step 4:

[1300] The server uses the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to recognize their emotional state. This analysis uses facial recognition algorithms and voice analysis technology. The input is video data containing the user's facial expressions and vocal characteristics. The output is data indicating the user's current emotional state.

[1301] Step 5:

[1302] The server identifies tidying up points that can be completed within 15 minutes based on the results of object recognition and emotion analysis. Specifically, it selects areas that are easy to tackle and will have the greatest impact. The input is the results of object recognition and emotion analysis. The output is data on tidying up points and the steps involved.

[1303] Step 6:

[1304] The server sends the identified cleanup points and procedures to the user's terminal. The user then follows the procedures to carry out the cleanup work. The input is the data on the cleanup points and procedures. The output is the cleanup instructions displayed on the user's terminal.

[1305] Step 7:

[1306] When the user finishes cleaning up and presses the "Clean Up Complete" button, the terminal sends a completion notification to the server. The input is the user's completion operation. The output is the completion notification sent to the server.

[1307] Step 8:

[1308] The server again uses the Emotion Engine to analyze the user's emotional state. Based on the results, it generates a praising message to display after the user has tidied up. The input is the user's latest emotional data. The output is a message based on the user's emotional state.

[1309] Step 9:

[1310] The server suggests further tidying up points based on the user's emotional state. If the user is positive, it suggests bigger tasks, and if negative, it suggests easier tasks. The input is the user's emotional state and the current tidying up situation. The output is new tidying up points and procedure data.

[1311] Step 10:

[1312] The server generates information to support the optimization of product placement and inventory management in physical stores and provides it to managers and staff. This enables efficient in-store management. The input is data on the placement of objects in the store and the status of organization. The output is optimized placement plans and organization procedures.

[1313] The above are the processing steps of this system.

[1314] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1315] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1316] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1317] [Fourth embodiment]

[1318] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1319] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1320] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1321] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1322] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1323] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1324] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1325] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1326] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1327] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1328] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1329] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1330] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1331] The present invention provides an application system that makes it easier for users to tidy up their rooms, and an embodiment of the system is specifically described below. In particular, by presenting tidying up points that can be completed in a short time and showing the steps in an easy-to-understand manner, the system allows users to tackle tidying up without feeling any resistance.

[1332] This system allows users to use devices such as smartphones or tablets to take videos of their rooms and then analyze those videos. Specifically, the user launches the app and sends the video of the room's condition to the server. The server then analyzes the received video data, recognizes objects in the room, and determines their placement and the level of clutter. This analysis uses computer vision technology and machine learning algorithms.

[1333] Based on the video analysis results, the server identifies tidying up tasks that the user can complete within 15 minutes and generates instructions for doing so. For example, if the user is trying to tidy up documents scattered on a desk, the server will classify the documents by category and present instructions for putting them away in the appropriate storage location. This information is sent to the user's device, and the user can view the specific tidying up steps on the app screen.

[1334] After the user has finished tidying up, they press the "Tidy up complete" button in the app, and the device sends a completion notification to the server. The server receives this notification and generates a message praising the user. For example, the message might read, "Great job! Your desk is so tidy now!" If the user wants to continue tidying up, the option "Do you want to continue?" is presented. If the user answers "Yes," the server identifies new tidying points and sends the instructions to the device again. Through this series of steps, the user can gradually tidy up the entire room.

[1335] In addition, the system also provides a matching function for housekeeping services. When a user selects the "Ask a Professional" button within the app, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can review this quote and decide whether to use the service.

[1336] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The server then generates instructions such as "Classify documents into three categories (work, school, and personal) and put work-related documents in a red folder," and sends these to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying, "Great job! Your desk looks so tidy now!" If the user wants to continue tidying, the server asks, "Do you want to continue?". If the user answers "yes," it presents instructions for tidying up points (for example, inside the desk drawers) that can be completed within the next 15 minutes.

[1337] In this way, the system of the present invention provides the user with appropriate means and advice to help them tidy up their room effectively and efficiently, thereby reducing their resistance to tidying up and helping them develop good tidying habits.

[1338] The processing flow will be explained below.

[1339] Step 1:

[1340] The user launches the app using a smartphone or tablet and takes a video of the room.

[1341] Step 2:

[1342] The device stores the video data captured within the app and sends it to a server via the Internet.

[1343] Step 3:

[1344] The server receives the video data and the analysis module performs object recognition and classification within the video, using computer vision and machine learning algorithms.

[1345] Step 4:

[1346] Based on the analysis results, the server evaluates the arrangement and clutter of items and furniture in the room to grasp the overall condition.

[1347] Step 5:

[1348] The server identifies areas and tasks that can be completed within 15 minutes, and selects areas and tasks that are easy for users to complete in a short time.

[1349] Step 6:

[1350] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[1351] Step 7:

[1352] The server transmits the identified cleaning points and procedures to the terminal.

[1353] Step 8:

[1354] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[1355] Step 9:

[1356] The user follows the instructions to tidy up, and when tidying up is complete, presses the "tidy up complete" button.

[1357] Step 10:

[1358] The terminal sends a "tidying up completed" notification to the server.

[1359] Step 11:

[1360] The server receives the completion notification and generates a message praising the user.

[1361] Step 12:

[1362] A server-generated compliment message is sent to the device.

[1363] Step 13:

[1364] The terminal displays a complimentary message on the user interface.

[1365] Step 14:

[1366] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[1367] Step 15:

[1368] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[1369] Step 16:

[1370] The server checks the room status again and identifies areas that can be cleaned up within the next 15 minutes.

[1371] Step 17:

[1372] The server generates a new cleanup procedure and sends it to the terminal.

[1373] Step 18:

[1374] The terminal displays the new tidying up points and procedures received from the server on the user interface, and the user can confirm them and choose whether to tidy up again.

[1375] Optional Process:

[1376] Step A1:

[1377] The user selects the "Ask a Pro" button within the app.

[1378] Step A2:

[1379] The terminal transmits the room status information to the server.

[1380] Step A3:

[1381] Based on the received room status information, the server sends a quote request to an affiliated housekeeping service provider.

[1382] Step A4:

[1383] The housekeeping service provider creates a quote and sends it to the server.

[1384] Step A5:

[1385] The server sends the quote information to the terminal.

[1386] Step A6:

[1387] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[1388] Example 1

[1389] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1390] For many people, tidying up their rooms is a tedious and time-consuming task. This can lead to resistance or a sense of burden, making it difficult to keep a room clean. Even when hiring a professional, finding the right service can be a complicated process, resulting in cost and effort. This invention aims to solve these problems by providing a means for users to efficiently tidy up their rooms in a short amount of time.

[1391] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1392] In this invention, the server includes a means for recording images of the room, a means for analyzing the recorded images to identify objects and grasp the state of the room, and a means for identifying tidying up points that can be completed in a short time based on the analysis results. This allows the user to easily know the specific steps for tidying up the room, reducing the hassle. Furthermore, by matching with a housekeeping service, it is possible to reduce the effort required to hire a professional.

[1393] The "means for recording images within the room" refers to a camera function or a video recording function for recording the state of the room.

[1394] "Means for analyzing recorded images to identify objects and understand the state of a room" refers to methods that use computer vision techniques and machine learning algorithms to detect and identify objects from image data and understand the state of a room.

[1395] "Means for identifying tidying up points that can be completed in a short time based on the analysis results" refers to a method for selecting specific locations or items that a user can complete tidying up in a short time based on the analyzed image data.

[1396] "Means for displaying the identified tidying up points and the procedures to the user" refers to the functionality of a display or application for visually or textually showing the tidying up points and the procedures to the user.

[1397] The "means for displaying an acknowledgement message to the user after tidying up" refers to a method for displaying a message praising the user after tidying up is completed.

[1398] "Means for suggesting additional tidying up points to the user" refers to a method for suggesting additional tidying up tasks to the user.

[1399] "Means for sending a quote request to be used for matching with a housekeeping service" refers to a method for sending a quote request to a housekeeping service provider based on the condition of the room and matching with a user.

[1400] This invention is a system that allows a user to efficiently and effectively tidy up a room. Specifically, the system allows a user to take a video of the room using a device such as a smartphone or tablet, and analyzes the video to present tidying tips and procedures that the user can implement in a short amount of time. An embodiment of this system is described in detail below.

[1401] First, the user launches the application on their smartphone or tablet and records a video of the room's condition. This video is then sent from the device to the server. The server then analyzes the video data using computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow). The goal of the analysis is to identify objects in the room and understand the state of the room.

[1402] As a concrete example, consider the case where a user wants to tidy up their living room. The user uses a smartphone to record video of the living room and sends the video from the device to a server. The server processes the received video data and identifies objects in the living room (e.g., magazines, remote controls, cushions, etc.). Based on the analysis results, the server generates instructions such as "classify the magazines into three categories (fashion, sports, and news) and put them away in a specific place on the bookshelf." This instruction is then sent to the device and presented to the user.

[1403] The user follows the presented steps to tidy up. After completing the tidy up, the user presses the "Tidy up complete" button in the app. A completion notification is sent from the device to the server, and the server generates a message of praise for the user and displays it on the device. For example, a message such as "Great job! Your living room is now so tidy!" may be displayed.

[1404] If the user wants to continue tidying up, the system asks the user, "Do you want to continue?" If the user answers "yes," the server identifies the next tidying point and sends the instructions to the device again. For example, the next step might be, "Clean up the items on the living room table."

[1405] The system also provides a matching function with housekeeping services. When a user selects the "Ask a Professional" button, the device sends information about the room's condition to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The quote from the provider is sent to the device via the server, and the user can review it and decide whether to use the service.

[1406] An example of a prompt sentence is, "I took a picture of the state of my desk with my smartphone and sent it to the server. Please generate easy steps to tidy up the documents on my desk, and if necessary, please also suggest the next step to tidy up."

[1407] This allows users to gradually tidy up their rooms and easily use housekeeping services. The system aims to reduce users' resistance to tidying up and help them develop the habit of tidying up.

[1408] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1409] Step 1:

[1410] The user takes a video of the room.

[1411] A user uses a smartphone or tablet to record a video of the state of a room. For example, they can record a video so that magazines, remote controls, and other items in the living room are visible. The input is video data of the current state of the room, and the output is a video file saved in the device's storage.

[1412] Step 2:

[1413] The device sends the video to the server.

[1414] The captured video is uploaded from the device to the server. At this time, the device compresses the video data before transferring it to reduce communication delays. The input is the video file stored in the device's storage, and the output is the video data uploaded to the server.

[1415] Step 3:

[1416] The server analyzes the video and identifies the object.

[1417] The server uses computer vision techniques (e.g., OpenCV) and machine learning algorithms (e.g., TensorFlow) to analyze the received video data. In this process, an object detection algorithm is used to identify objects in the video (e.g., magazines, remote controls, cushions, etc.). The input is the video data uploaded to the server, and the output is the object identification results (object type and location).

[1418] Step 4:

[1419] The server identifies the tidying up point based on the analysis results.

[1420] Based on the location information of the identified objects, the server identifies specific tidying up steps that the user can complete in a short time. For example, it generates a procedure such as "sort the magazines in the living room into three categories (fashion, sports, and news) and store them in a specific place on the bookshelf." The input is the object identification result, and the output is the identification of the tidying up steps (specific tidying up steps).

[1421] Step 5:

[1422] The server sends the generated cleanup procedure to the terminal.

[1423] The specific cleanup procedure created by the server is sent to the terminal. The input is the cleanup procedure generated by the server, and the output is the cleanup procedure sent to the terminal.

[1424] Step 6:

[1425] The user follows the cleaning procedure to perform the cleaning.

[1426] The user follows the instructions displayed on the device application to tidy up. For example, the user performs a specific action such as "sorting magazines into fashion, sports, and news, and storing them on the bookshelf." The input is the tidying up instructions displayed on the device, and the output is the physical state of the room after the tidying up is completed.

[1427] Step 7:

[1428] The user sends a cleanup completion notification to the server.

[1429] After the user has finished cleaning up, they press the "Clean Up Complete" button in the application. This causes a completion notification to be sent from the device to the server. The input is the user's operation, and the output is the completion notification sent to the server.

[1430] Step 8:

[1431] The server generates a completion message and sends it to the terminal.

[1432] The server receives the completion notification and generates a message praising the user. For example, it creates a message such as "Great job! Your living room is now so tidy!" and sends it to the terminal. The input is the completion notification, and the output is the completion message sent to the terminal.

[1433] Step 9:

[1434] The user chooses whether to clean up further.

[1435] The terminal presents the user with the option "Do you want to continue?" If the user answers "Yes," the procedure to suggest the next tidying up point begins. The input is the user's selection, and the output is the start of the suggestion of the next tidying up point.

[1436] Step 10:

[1437] The server identifies the next cleanup point.

[1438] If the user answers "yes," the server identifies the next step to clean up and generates instructions for it. For example, it creates a specific step such as "Clean up the items on the living room table" and sends it to the device. The input is the user's selection, and the output is the next step to clean up that was sent to the device.

[1439] Step 11:

[1440] If the user selects a housekeeping service, the server performs the matching.

[1441] When a user selects the "Ask a Professional" button in the application, the device sends the room's condition information to the server. Based on this information, the server sends a quote request to a housekeeping service provider and displays the quote from the provider to the user. The input is the room's condition information, and the output is the quote for the housekeeping service.

[1442] (Application example 1)

[1443] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1444] The present invention relates to an application system that allows users to easily tidy up their rooms. Conventional tidying support systems have the problem that when a user starts tidying up, it is unclear which object to start with, making it difficult to tidy up effectively. A similar problem exists in tidying up physical stores, where there is a lack of specific guidance for store staff to work efficiently. The present invention aims to solve these problems.

[1445] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1446] In this invention, the server includes means for taking video of the room, means for analyzing the video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is complete, means for suggesting further tidying up points to the user, means for taking video of the tidying up state of the physical store, and means for presenting the tidying up procedures to store staff. This enables users and store staff to specifically and efficiently tidy up and organize in a short amount of time.

[1447] "Means for capturing video of the inside of a room" refers to equipment or a method for capturing video of the inside of a room using a camera and acquiring the video data.

[1448] "Means for analyzing captured video to recognize objects and grasp the state of a room" refers to equipment and methods that use video analysis technology to identify items present in a room based on acquired video data and grasp the overall situation of the room.

[1449] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method or algorithm that selects and identifies tidying up tasks that a user can complete in a short amount of time based on the results of video analysis.

[1450] "Means for presenting the identified tidying up points and procedures to the user" refers to equipment or methods for displaying or notifying the user of the identified work content and specific procedures so that the user can tidy up efficiently.

[1451] The "means for displaying a message praising the user after tidying up is completed" is a means for presenting a message praising the user's efforts when the user has finished tidying up.

[1452] The "means for suggesting further tidying up points to the user" is a means for presenting new tidying up tasks to the user when the user wants to continue tidying up.

[1453] "Means for photographing the tidiness and neatness of a physical store" refers to equipment or a method for using a camera to photograph the state of items and shelves in a physical store and obtain the video data.

[1454] "Means for presenting tidying procedures to store staff" refers to devices or methods for displaying or informing store staff of specified work content and specific procedures so that they can efficiently tidy up items.

[1455] The present invention provides a system that allows users to efficiently tidy up and organize their rooms and physical stores. Specific embodiments of the system will be described below.

[1456] This system uses devices such as smartphones and tablets to record footage of the state of a room or physical store, and then analyzes the video data to present cleaning and tidying procedures. Users use their devices to record video of the room or store and send the video to a server. The server analyzes the video data and uses computer vision technology and machine learning algorithms to recognize objects. Libraries such as OpenCV and TensorFlow can be used for this analysis.

[1457] Based on the results of the video analysis, the server identifies tidying up tasks that the user can complete within 15 minutes and automatically generates detailed instructions. The generated instructions are sent to the user's device, showing them exactly how to proceed with the tidying up. For example, specific steps such as "sort documents by category and put work-related documents in a red folder" are presented.

[1458] The user checks these steps through the app screen and performs the tidying task. After completing the tidying, the user presses the "Tidying up complete" button on the device, and the device sends a notification to the server. The server receives this notification and generates and displays a message of praise to the user, such as "Great job! Your desk is so tidy now!" If the user wants to continue tidying, it presents the option "Do you want to continue?". If the user answers "Yes," the server identifies new tidying points, generates new steps, and sends them to the device.

[1459] In the case of physical stores, the photographing and analysis procedures using the device are similar. The server identifies the state of organization of the physical store and presents specific organizational procedures to store staff. For example, it presents procedures for sorting products that are randomly placed on shelves and tidying them up by category. This allows store staff to work efficiently.

[1460] The hardware used is a smartphone, tablet, and server, and the software uses OpenCV for video analysis and TensorFlow for machine learning algorithms.

[1461] As a concrete example, if magazines are placed in a disorganized manner in a storeroom, a store clerk can use a smartphone to take a video of the situation and analyze the video. Based on the analysis results, a procedure will be generated, such as "sort the magazines by title and organize them on specific shelves." This allows the store clerk to follow the instructions and efficiently organize the magazines.

[1462] An example of a prompt for a generative AI model might be, "Please take a survey for an application that takes photos of the state of shelves in a physical store and generates the types of items and the procedures for organizing them."

[1463] This system allows users and store staff to tidy up and organize in a specific and efficient manner in a short amount of time.

[1464] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1465] Step 1:

[1466] Users use devices such as smartphones or tablets to shoot video of a room or a physical store.

[1467] Input: Video data captured using the device's camera function.

[1468] Output: The captured video file.

[1469] This video file will be used for later analysis.

[1470] Step 2:

[1471] The device sends the captured video to the server.

[1472] Input: Video files stored on your device.

[1473] Output: Video data sent to the server over the network.

[1474] The server receives this data and prepares it for analysis.

[1475] Step 3:

[1476] The server analyzes the received video data, recognizes objects, and understands the condition of the room or physical store.

[1477] Input: Video data sent to the server.

[1478] Output: Object recognition results in the video and room / store status information.

[1479] This analysis uses libraries such as OpenCV and TensorFlow to identify the location and type of object.

[1480] Step 4:

[1481] Based on the results of the video analysis, the server identifies tidying up points that the user can complete within 15 minutes.

[1482] Input: Object recognition results and room / store state information.

[1483] Output: Identified cleanup points and specific cleanup procedures.

[1484] Specific instructions are automatically generated, detailing which items to organize and how.

[1485] Step 5:

[1486] The server transmits the identified tidying up points and their procedures to the terminal and presents them to the user.

[1487] Input: Server-generated cleanup instructions.

[1488] Output: Cleanup instructions displayed on the user's terminal.

[1489] By following this procedure, the user can efficiently tidy up.

[1490] Step 6:

[1491] After the user has finished tidying up, he / she presses the "tidying up complete" button on the terminal.

[1492] Input: Cleanup completion notification action from user.

[1493] Output: Tidy-up completion notification sent to the server.

[1494] This notification lets the server know the progress of the cleanup.

[1495] Step 7:

[1496] The server receives the notification that the tidying up is complete, generates a message praising the user, such as "good job," and sends it to the terminal.

[1497] Input: Notification from user that cleanup is complete.

[1498] Output: A compliment message displayed on the terminal.

[1499] This gives the user a sense of accomplishment.

[1500] Step 8:

[1501] The server asks the user whether he / she wants to continue cleaning up, and if he / she does, it identifies a new cleaning point, generates a procedure, and sends it to the terminal.

[1502] Input: User confirms their desire to continue tidying.

[1503] Output: New cleanup procedure.

[1504] This allows the user to move on to the next tidying task.

[1505] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1506] This invention provides a system that presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[1507] The system begins when a user takes a video of their room using a device such as a smartphone or tablet and sends the video to the server. The server analyzes the received video to recognize objects and determine the state of the room. This analysis is carried out using computer vision and machine learning algorithms. Based on the results of this analysis, the server identifies tidying up points that the user can complete within 15 minutes and generates a procedure for doing so.

[1508] Furthermore, the system is equipped with an emotion engine that recognizes the user's emotional state. This emotion engine analyzes the user's facial expressions and tone of voice via a camera and microphone to grasp the user's current emotional state. This allows the system to take the user's emotional state into consideration when presenting tidying points and procedures.

[1509] When a user tidies up through the app and presses the "Tidy up complete" button after completing the task, the device sends a completion notification to the server. The server generates a message of praise according to the user's emotional state based on the results of emotion analysis by the emotion engine. For example, if the user is tired, a message such as "You did a great job! Let's take a break" will be displayed.

[1510] Furthermore, if the user chooses to continue tidying up, the system also takes their emotional state into account when suggesting the next step: if the user is in a positive emotional state, it may suggest a larger or more complex task, while if the user is in a negative emotional state, it may suggest an easier task.

[1511] The system also provides a matching function for housekeeping services. When the user selects the "Ask a Professional" button, the device sends information about the room's condition and the results of analysis by the emotion engine to the server. Based on this information, the server sends a quote request to affiliated housekeeping service providers. The housekeeping service provider creates a quote and sends it to the device via the server. The user can then review the quote and decide whether to use the service.

[1512] As a concrete example, consider the case where a user wants to tidy up the documents on their desk. The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and recognizes the location and type of documents on the desk. The emotion engine also analyzes the user's facial expressions and recognizes that the user is motivated. The server then generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder" and sends them to the device. The user follows these instructions to tidy up, and when they press the "Tidy up complete" button, the server displays a message praising them, saying "Good job! Your desk looks so tidy now!" If the user wants to continue tidying, the system asks "Do you want to continue?". If the user answers "yes," the system suggests tidying up points that can be completed within the next 15 minutes (for example, tasks related to the desk drawers).

[1513] In this way, the system of the present invention provides tidying up advice that takes into account the user's emotional state, thereby reducing resistance to tidying up and helping to form tidying up habits. Furthermore, by facilitating the use of housekeeping services, the system provides more efficient tidying up support.

[1514] The processing flow will be explained below.

[1515] Step 1:

[1516] The user launches the app using a smartphone or tablet and takes a video of the room.

[1517] Step 2:

[1518] The device stores the video data captured within the app and sends it to a server via the Internet.

[1519] Step 3:

[1520] The server receives the video data and the analysis module recognizes and classifies objects in the video using computer vision and machine learning algorithms. The server then determines the location, type, and overall condition of the recognized objects.

[1521] Step 4:

[1522] Based on the analysis results, the server identifies tidying up points that the user can complete within 15 minutes, and also selects areas and tasks that the user can easily accomplish in a short amount of time.

[1523] Step 5:

[1524] The server generates specific cleaning procedures for each cleaning point, such as "sort documents into three categories and put work-related documents in a red folder."

[1525] Step 6:

[1526] The emotion engine analyzes the user's facial expressions and tone of voice through the device's camera and microphone to recognize the user's emotional state.

[1527] Step 7:

[1528] The server then adjusts the cleaning procedure and approach based on the emotional state of the user based on the analysis results of the emotion engine. For example, it adds positive comments to users in a positive emotional state and encouraging comments to users in a negative emotional state.

[1529] Step 8:

[1530] The server transmits the adjusted cleaning points and procedures to the terminal based on the emotional state.

[1531] Step 9:

[1532] The terminal displays the tidying up points and procedures received from the server on the user interface, and the user confirms them.

[1533] Step 10:

[1534] The user follows the instructions to tidy up, and when finished, presses the "tidy up complete" button.

[1535] Step 11:

[1536] The terminal sends a "tidying up completed" notification to the server.

[1537] Step 12:

[1538] The server generates a praising message according to the user's emotional state based on the emotion analysis results of the emotion engine.

[1539] Step 13:

[1540] A server-generated compliment message is sent to the device.

[1541] Step 14:

[1542] The device will display a praising message on the user interface. For example, if the user is tired, it will say "Great job! Let's take a break."

[1543] Step 15:

[1544] The device displays a praise message followed by a "Do you want to continue?" option to the user.

[1545] Step 16:

[1546] If the user answers "yes," the terminal sends a request for the next tidy-up point to the server.

[1547] Step 17:

[1548] The server checks the room status again and identifies new cleanup points that can be completed within 15 minutes.

[1549] Step 18:

[1550] The server uses an emotion engine to analyze the user's current emotional state.

[1551] Step 19:

[1552] The server generates a new cleaning procedure based on the emotional state and sends it to the terminal.

[1553] Step 20:

[1554] The terminal displays the new tidying up points and procedures on the user interface, and the user can confirm this and choose whether to tidy up again.

[1555] Optional Process:

[1556] Step A1:

[1557] The user selects the "Ask a Pro" button within the app.

[1558] Step A2:

[1559] The device sends the room status and the emotion engine's analysis results to the server.

[1560] Step A3:

[1561] Based on the information received by the server, a request for an estimate is sent to an affiliated housekeeping service provider.

[1562] Step A4:

[1563] The housekeeping service provider creates a quote and sends it to the server.

[1564] Step A5:

[1565] The server sends the quote information to the terminal.

[1566] Step A6:

[1567] The terminal displays the estimate information received from the server on the user interface. The user confirms the estimate and selects whether or not to request the housekeeping service.

[1568] As a concrete example, if a user wants to tidy up the documents on their desk, the process would proceed as follows: The user takes a video of the state of their desk with their smartphone and sends this video to the server. The server analyzes the video and understands the state of the desk. The emotion engine recognizes that the user is motivated. The server generates instructions such as "Classify the documents into three categories and put work-related documents in a red folder," and sends these instructions to the device. When the user completes the tidying up by following the instructions and presses the "Tidying up complete" button, the server generates a message saying "Great job! Your desk looks so tidy now!" and displays it on the device. If the user wants to continue tidying up, the message "Do you want to continue?" will be displayed, and if the user answers "yes," the next tidying up point (for example, instructions for tidying up the desk drawers) will be suggested. This process allows the user to gradually tidy up an entire room.

[1569] Example 2

[1570] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1571] The present invention aims to provide effective tidying support for people who have difficulty tidying up and for children who want to make tidying a habit. Another objective of the present invention is to realize a system that takes into account the user's emotional state to increase motivation to tidy up and encourage tidying up in a short amount of time.

[1572] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1573] In this invention, the server includes means for taking a video of the room, means for analyzing the taken video to recognize objects and grasp the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for recognizing the user's emotional state, means for presenting the identified tidying up points and their procedures to the user, means for displaying a message of praise according to the user's emotional state after tidying up is completed, and means for suggesting further tidying up points to the user. This makes it possible to present tidying up points that can be completed in a short time while taking the user's emotional state into consideration, reducing resistance to tidying up and increasing motivation.

[1574] "Means for capturing video in a room" is a function that allows a user to capture video of the entire room using a device such as a smartphone or tablet.

[1575] "Means of analyzing the captured video to recognize objects and understand the state of the room" refers to a function that enables the server to use computer vision and machine learning algorithms to identify objects in the video and determine the current state of the room.

[1576] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" is a function that enables the server to select tidying up tasks that the user can complete within 15 minutes based on the analysis results.

[1577] The "means for recognizing the user's emotional state" is a function in which the server analyzes the user's facial expressions and voice data collected through the device's camera and microphone to determine the user's emotional state.

[1578] The "means for presenting the identified tidying up points and the procedures to the user" is a function for displaying the tidying up tasks selected by the server and the procedures for executing them on the user's terminal.

[1579] The "means for displaying a praising message according to the user's emotional state after tidying up is completed" is a function for displaying an appropriate praising message on the user's terminal after the tidying up task is completed, based on the emotional state analyzed by the server.

[1580] The "means for suggesting further tidying up points to the user" is a function for selecting and suggesting new tidying up tasks to a user who wants to continue tidying up.

[1581] "Means for sending quotation requests to be used for matching with housekeeping services" is a function for sending quotation requests to housekeeping service providers based on the video footage and the results of sentiment analysis.

[1582] This system presents tidying tips that can be completed in a short time for people who have difficulty tidying up or for children who want to make tidying a habit. In particular, by combining it with an emotion engine that recognizes the user's emotional state, we aim to maximize the effectiveness of tidying up and increase the user's motivation.

[1583] Hardware and Software Configuration

[1584] This system consists of a device such as a smartphone or tablet, a processing server, and an emotion engine. The specific hardware and software used are as follows:

[1585] Devices: Smartphones and tablets equipped with cameras and microphones

[1586] Server: GPU-equipped server

[1587] software:

[1588] Video analysis and object recognition: TensorFlow, OpenCV

[1589] Emotion recognition engine: Microsoft Azure Emotion API or other emotion recognition API

[1590] How it works

[1591] 1. Recording and sending video in the room:

[1592] User: Use the device camera to record a video of the entire room. It is recommended to capture key areas such as the desk, floor, and shelves.

[1593] Device: Temporarily stores the captured video and then uploads it to the server.

[1594] 2. Video Analysis:

[1595] Server: Analyzes the received video using computer vision and machine learning algorithms (e.g., TensorFlow and OpenCV) to recognize objects in the room.

[1596] Example: "Recognize documents and pens on a desk, the position of a chair, and objects on the floor."

[1597] 3. Recognition of emotional states:

[1598] User: Speaks and shows facial expressions into the device's camera and microphone to capture emotional state.

[1599] Terminal: Sends camera images and audio to the server.

[1600] Server: The emotion engine analyzes facial expressions and tone of voice to determine the user's current emotional state. For example, if the user is smiling, it is determined that the user is "motivated."

[1601] 4. Generate cleanup points and procedures:

[1602] Server: Based on the analysis of the video and the user's emotional state, it generates tidying up points and specific steps that the user can complete within 15 minutes.

[1603] Example: Generate a procedure such as "Classify the documents on your desk into three categories (important, temporary storage, and disposal)" and send it to the terminal.

[1604] 5. Feedback Generation:

[1605] User: Follow the steps provided to tidy up. When finished, press the "Clean up" button in the app.

[1606] Terminal: Sends a completion notification to the server.

[1607] Server: Based on the results of emotion analysis by the emotion engine, a message of praise is generated and sent to the device. A message such as "You did a great job! Let's take a short break" is displayed.

[1608] 6. Suggestions for the following tidying points:

[1609] Server: If the user wants to continue cleaning, the server suggests new cleaning points. Taking into account the user's emotional state, the server presents slightly larger tasks if the emotional state is positive, and easier tasks if the emotional state is negative.

[1610] Example: "I suggest organizing your desk drawers as your next tidying point."

[1611] Matching housekeeping services

[1612] If the user feels that cleaning up by themselves is difficult, they can select the "Ask a professional" button to use a housekeeping service. In this case, the system operates as follows:

[1613] Terminal: Sends information about the room's state and the emotion engine's analysis results to the server.

[1614] Server: Sends a quote request to affiliated housekeeping service providers and sends the quote results to the terminal.

[1615] User: Can review the quote and decide to use the service.

[1616] Prompt Sentence Examples

[1617] An example of a prompt for a generative AI model is:

[1618] "After the user takes a video of their room and sends it to the server, use computer vision and machine learning to analyze the state of the room. Then, use an emotion engine to understand the user's emotional state, generate tidying points and steps that can be completed in 15 minutes, and return appropriate praise messages."

[1619] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1620] Step 1: Record and send your video

[1621] The user uses a device such as a smartphone or tablet to record a video of the entire room, making sure to capture key areas such as the desk, floor, and shelves.

[1622] Input: A video of the room taken by the user

[1623] Specific operation: The user lifts the device and moves the camera to capture the entire room while shooting video.

[1624] Output: Recorded video file

[1625] The device temporarily stores the captured video files and then uploads them to a server using Wi-Fi or mobile data.

[1626] Input: Video files stored on the device

[1627] Specific operation: The device saves the video file and sends it to the server via the network.

[1628] Output: Video file sent to the server

[1629] Step 2: Analyze the video

[1630] The server analyzes the received video using computer vision and machine learning algorithms (TensorFlow and OpenCV), recognizes objects in the room, and determines the current state of the room.

[1631] Input: Video file sent to the server

[1632] What it does: The server extracts frames from a video file, uses computer vision algorithms to detect objects, and uses machine learning models to predict the category of each object.

[1633] Output: Object recognition results in the room (object position, type)

[1634] Step 3: Recognizing your emotional state

[1635] The user's emotional state is captured by speaking into the device's camera and microphone and showing facial expressions.

[1636] Input: User's facial expression and voice data

[1637] Specific actions: The user smiles at the camera and makes a short comment out loud.

[1638] Output: Photographed facial expressions and recorded audio

[1639] The device transmits camera images and audio to the server.

[1640] Input: Photographed facial expressions and recorded audio data

[1641] Specific operation: The device captures these data and sends them to the server.

[1642] Output: Facial expression and voice data sent to the server

[1643] The server uses an emotion engine to analyze facial expressions and tone of voice to determine the user's current emotional state.

[1644] Input: Transmitted facial expression and voice data

[1645] Specific operation: The server extracts facial and vocal features using an emotion engine and classifies the emotional state based on these features.

[1646] Output: Emotional state (positive, negative, neutral, etc.)

[1647] Step 4: Generate cleanup points and procedures

[1648] Based on the results of object recognition and emotional state analysis in the video, the server generates tidying up points and specific steps that the user can complete within 15 minutes.

[1649] Input: Object recognition results in the room and emotional state

[1650] Specific operation: The server compares the analysis results and selects the optimal tidying task. For example, it generates a procedure such as "classify the documents on the desk into three categories (important, temporary storage, and disposal)."

[1651] Output: Cleaning points and procedures

[1652] Step 5: Clean up and notify completion

[1653] The user follows the displayed steps to tidy up, and when they are done, they press the "Tidy up complete" button displayed in the app on their device.

[1654] Input: Cleaning points and procedures

[1655] Specific actions: The user follows the instructions to tidy up and presses the "tidy up complete" button when finished.

[1656] Output: Tidy up completion notification

[1657] The terminal notifies the server that the "tidy up complete" button has been pressed.

[1658] Input: Cleanup completion notification

[1659] Specific operation: The terminal sends a completion notification to the server.

[1660] Output: Completion notification sent to the server

[1661] Step 6: Generate feedback

[1662] Based on the emotion analysis results from the emotion engine, the server generates a praising message that corresponds to the user's emotional state and sends it to the terminal.

[1663] Input: Tidy-up completion notification and emotional state

[1664] Specific behavior: The server generates an appropriate praise message based on the emotional state, for example, "Good job! Let's take a break."

[1665] Output: Complimentary message

[1666] Step 7: Suggest the next step

[1667] If the user wants to continue tidying up, the server will suggest new tidying up points taking into account the user's emotional state.

[1668] Input: Emotional state and intention to continue tidying up

[1669] Specific behavior: The server checks the user's emotional state and suggests a slightly larger task if it is positive, or an easier task if it is negative. For example, it suggests, "As the next step in tidying up, we suggest organizing your desk drawers."

[1670] Output: Suggested next steps and procedures

[1671] Matching housekeeping services

[1672] If the user finds it difficult to clean up by themselves, they can use a housekeeping service by selecting the "Ask a professional" button.

[1673] Input: Select the Ask a Pro button

[1674] Specific operation: When the "Ask a Pro" button is pressed, the device sends information about the room's status and the emotion engine's analysis results to the server.

[1675] Output: Room status information and sentiment analysis results sent to the server

[1676] The server sends a request for a quote to an affiliated housekeeping service provider and sends the quote result to the terminal.

[1677] Input: Room status information and sentiment analysis results

[1678] Specific operation: The server sends this information to the housekeeping service provider and receives the estimate.

[1679] Output: Estimate results sent to the user's device

[1680] The user checks the estimate and decides to use the service if necessary.

[1681] Input: Estimate result

[1682] Specific operation: The user checks the quote and decides whether to use the service.

[1683] Output: Decision to use the service or cancellation

[1684] (Application example 2)

[1685] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1686] Conventional tidying support systems lacked sufficient means to effectively tidy up while maintaining user motivation. Furthermore, in physical stores, it was difficult to efficiently organize inventory and optimize product placement. This increased resistance to tidying up, making it difficult to organize regularly.

[1687] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing video of the room, means for analyzing the captured video to recognize objects and understand the state of the room, means for identifying tidying up points that can be completed within 15 minutes based on the analysis results, means for presenting the identified tidying up points and procedures to the user, means for displaying a message praising the user after the tidying up is completed, means for recognizing the user's emotional state and generating a message according to the emotional state, means for suggesting further tidying up points based on the emotional state, and means for applying the results to product placement and inventory management in a physical store. This enables effective tidying up while maintaining the user's motivation, and enables efficient management of inventory management and product placement in a physical store.

[1688] "Means for capturing video in a room" refers to a device or method for capturing video of the entire room or a specific area and acquiring the video data.

[1689] "Means for analyzing captured video to recognize objects and understand the state of the room" refers to technology that processes captured video data, identifies objects and their placement within the room, and understands the state of the room as a whole.

[1690] "Means for identifying tidying up points that can be completed within 15 minutes based on the analysis results" refers to a method for using the results of video analysis to determine specific locations and items of tidying up work that can be completed in a short amount of time.

[1691] The "means for presenting the identified cleanup points and procedures to the user" refers to an interface or notification mechanism for providing the user with clear cleanup instructions and guidelines.

[1692] The "means for displaying a message praising the user after the completion of tidying up" refers to a system or function for displaying a message praising the efforts of a user who has completed tidying up.

[1693] "Means for recognizing the user's emotional state and generating a message according to that emotional state" refers to technology that analyzes the user's facial expressions and tone of voice to understand their emotional state and generate an appropriate message accordingly.

[1694] The "means for suggesting further tidying up points based on the emotional state" is a method for suggesting the next tidying up task to be done, taking into account the user's current emotional state.

[1695] "Means for application to product placement and inventory management in physical stores" refers to systems and methods for applying similar video analysis and tidying procedure presentation technologies to improve the efficiency of product placement and inventory management in physical stores.

[1696] The present invention is a system that realizes efficient management of inventory organization and product placement in physical stores. Its distinctive feature is that users can take videos of the store interior using smartphones or smart glasses, and the system analyzes the video data to suggest quick tidying tips and procedures. The specific system configuration and its implementation are described below.

[1697] System program generation

[1698] The system consists of the following main components:

[1699] 1. Camera device: A camera device for capturing video inside the store, such as a smartphone or smart glasses.

[1700] 2. Server: A central processing unit for video analysis and sentiment analysis.

[1701] 3. User interface: An application that presents cleaning points and procedures to the user.

[1702] Hardware and Software Description

[1703] Hardware: Smartphone, smart glasses (camera function)

[1704] These devices are used to capture images of the conditions inside the store.

[1705] software:

[1706] OpenCV: Used for video capture and processing.

[1707] Emotion Engine Module: Used to analyze the user's emotional state.

[1708] Object Detection Module: Used to perform object recognition and identify the location and type of product.

[1709] Server: As the central processing unit, it performs video analysis, object recognition, and emotion analysis.

[1710] Data processing and calculation flow

[1711] 1. The server receives videos taken by users using their smartphones or smart glasses.

[1712] 2. The received video data is analyzed using OpenCV, and objects within the store are recognized and their locations are determined.

[1713] 3. Use the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to determine their emotional state.

[1714] 4. Based on the results of object recognition and emotion analysis, the server generates tidying up points and steps that can be completed within 15 minutes and presents them to the user.

[1715] 5. After the user has completed tidying up, a message of praise will be displayed according to the user's emotional state, and if the user wishes to continue tidying up, new tidying points will be suggested.

[1716] Adding specific examples

[1717] As a concrete example, consider the following scenario.

[1718] Specific examples

[1719] Imagine a user tidying up a product shelf in a physical store. The user takes a video of the product shelf with their smartphone and sends the video data to a server. The server analyzes the video and recognizes the types of products on the shelf and their arrangement. The Emotion Engine analyzes the user's facial expressions and recognizes that the user is a little tired. The server then generates instructions such as "Organize the three columns on the right side of this shelf and rearrange the products by category," and displays these instructions on the smartphone. When the user follows these instructions to organize the shelves and presses the "Tidying up complete" button, the server displays a message of praise saying, "Good job! The shelves are now very clean!" If the user wants to continue tidying, the server displays the message "Do you want to continue?" and if the user answers "Yes," it suggests tidying up points that can be completed within the next 15 minutes.

[1720] Prompt Sentence Examples

[1721] "Based on your current emotional state, prioritize your inventory. Evaluate the condition of your shelves and tell us where to organize next."

[1722] As described above, the system of the present invention supports effective tidying up in physical stores while taking into account the emotional state of the user.

[1723] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1724] Step 1:

[1725] Users use smartphones or smart glasses to take videos of the inside of a store. This device is used to capture the inside of the store in detail. The input is video data that represents the current state of the store. The output is a video file that has been taken.

[1726] Step 2:

[1727] The device sends the captured video data to the server. Here, the device uses a stable communication environment to quickly upload the captured video to the server. The input is the video file and communication data. The output is the video data sent to the server.

[1728] Step 3:

[1729] The server analyzes the received video data. Using video analysis software such as OpenCV, it recognizes objects and identifies their layout in the store. This process uses computer vision technology. The input is the video data sent to the server. The output is the analysis results that identify the object's location and type.

[1730] Step 4:

[1731] The server uses the Emotion Engine to analyze the user's facial expressions and tone of voice in the video to recognize their emotional state. This analysis uses facial recognition algorithms and voice analysis technology. The input is video data containing the user's facial expressions and vocal characteristics. The output is data indicating the user's current emotional state.

[1732] Step 5:

[1733] The server identifies tidying up points that can be completed within 15 minutes based on the results of object recognition and emotion analysis. Specifically, it selects areas that are easy to tackle and will have the greatest impact. The input is the results of object recognition and emotion analysis. The output is data on tidying up points and the steps involved.

[1734] Step 6:

[1735] The server sends the identified cleanup points and procedures to the user's terminal. The user then follows the procedures to carry out the cleanup work. The input is the data on the cleanup points and procedures. The output is the cleanup instructions displayed on the user's terminal.

[1736] Step 7:

[1737] When the user finishes cleaning up and presses the "Clean Up Complete" button, the terminal sends a completion notification to the server. The input is the user's completion operation. The output is the completion notification sent to the server.

[1738] Step 8:

[1739] The server again uses the Emotion Engine to analyze the user's emotional state. Based on the results, it generates a praising message to display after the user has tidied up. The input is the user's latest emotional data. The output is a message based on the user's emotional state.

[1740] Step 9:

[1741] The server suggests further tidying up points based on the user's emotional state. If the user is positive, it suggests bigger tasks, and if negative, it suggests easier tasks. The input is the user's emotional state and the current tidying up situation. The output is new tidying up points and procedure data.

[1742] Step 10:

[1743] The server generates information to support the optimization of product placement and inventory management in physical stores and provides it to managers and staff. This enables efficient in-store management. The input is data on the placement of objects in the store and the status of organization. The output is optimized placement plans and organization procedures.

[1744] The above are the processing steps of this system.

[1745] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1746] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1747] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1748] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1749] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1750] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1751] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1752] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1753] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1754] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1755] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1756] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1757] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1758] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1759] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1760] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1761] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1762] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1763] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1764] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1765] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1766] The following is further disclosed regarding the above embodiment.

[1767] (Claim 1)

[1768] A means of taking video of the room,

[1769] A means of analyzing the captured video to recognize objects and understand the state of the room,

[1770] Based on the analysis results, a method for identifying tidying up points that can be completed within 15 minutes,

[1771] a means for presenting the identified tidying up points and procedures to the user;

[1772] A means for displaying a message praising the user after the tidying up is completed;

[1773] A means for suggesting further tidying points to the user;

[1774] A system including:

[1775] (Claim 2)

[1776] The system according to claim 1, further comprising means for sending a quote request that uses the state of the room based on the captured video to match with a housekeeping service.

[1777] (Claim 3)

[1778] The system according to claim 1, wherein, when the user selects to continue tidying up, new tidying up points that can be completed within 15 minutes are identified and steps are presented.

[1779] "Example 1"

[1780] (Claim 1)

[1781] means for recording images within the room;

[1782] A means for analyzing the recorded images to identify objects and grasp the state of the room;

[1783] A means for identifying tidying up points that can be completed in a short time based on the analysis results;

[1784] a means for displaying the identified tidying up points and the procedures therefor to the user;

[1785] a means for displaying an acknowledgement message to the user after tidying up;

[1786] ...

Claims

1. A means of taking video of the room, A means of analyzing the captured video to recognize objects and understand the state of the room, Based on the analysis results, a method for identifying tidying up points that can be completed within 15 minutes, a means for presenting the identified tidying up points and procedures to the user; A means for displaying a message praising the user after the tidying up is completed; A means for suggesting further tidying points to the user; A system including:

2. The system according to claim 1, further comprising means for sending an estimate request that uses the state of the room based on the captured video for matching with a housekeeping service.

3. 2. The system according to claim 1, wherein, when the user selects to continue tidying up, new tidying up points that can be completed within 15 minutes are identified and a procedure is presented.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A