system
A system using image capture and generative AI to manage and locate objects in rooms addresses the challenge of finding items efficiently, reducing time and stress.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Managing and quickly finding objects in environments with many items is difficult, particularly for busy individuals, the elderly, and those who are not good at tidying up, leading to wasted time and stress.
A system that uses image acquisition means to capture room environments, analyzes objects with a generative AI model, records their locations, and allows users to search via voice input for efficient object management.
Reduces time spent searching for items and enhances object management efficiency by providing quick and accurate location information.
Smart Images

Figure 2026071682000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In many modern households, especially in environments with a lot of items, it is difficult to manage objects and it is hard to quickly find the necessary items. This problem is particularly prominent for busy housewives, people who are not good at tidying up, the elderly, etc., causing a waste of time and stress in daily life. Against such a background, there is a demand for a system that can visually manage the locations of items in a room and quickly obtain their location information when needed.
Means for Solving the Problems
[0005] This invention recognizes and classifies objects by acquiring images of them using image acquisition means installed in the room environment and analyzing them using a generative artificial intelligence model. It then records and stores the location information of the analyzed objects and provides a search means that allows users to search using natural language via voice input, enabling them to quickly find specific objects. This reduces the time users spend searching for things in their daily lives and enables efficient object management.
[0006] "Image acquisition means" refers to a device or function for acquiring images of objects within a room environment.
[0007] "Analysis means" refers to a device or process that has the function of recognizing and classifying objects based on acquired images.
[0008] "Recording means" refers to a device or system for storing and managing the location information of analyzed objects in a database or similar format.
[0009] "Search means" refers to a device or function for searching for location information recorded based on user input and providing a response.
[0010] A "generative artificial intelligence model" refers to an algorithm or model that uses machine learning or deep learning to analyze and recognize the features of objects.
[0011] "Natural language processing" refers to the technology of analyzing human language input as speech or text and converting it into a format that a computer can understand and process. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0018] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] As an embodiment for carrying out the present invention, the object management system is specifically configured as follows. First, an image acquisition means installed in the room environment periodically photographs the room's conditions and acquires images of objects. These images are transmitted to a server via wireless communication means.
[0034] The server processes the received images using analysis tools and recognizes and classifies objects in the images using a generative artificial intelligence model. This process identifies the type and characteristics of each object and tags them. It also extracts the object's location information as coordinates within the image.
[0035] Next, the server uses recording devices to record the analyzed object information and its location information in a database. This makes it possible to track the object's latest location and also use it as historical information.
[0036] When a user wants to know the location of an object, they send a question to the server via voice or text through their device (smartphone or voice assistant). The server analyzes this input using natural language processing technology to determine which object to search for.
[0037] The server uses a search mechanism to retrieve the latest location information of the target object from the database and sends the search results back to the terminal. The terminal then communicates the search results to the user via screen display or audio output. This allows the user to quickly learn the object's current location.
[0038] For example, if a user uses their smartphone to voice-input "Where is the remote control?", the server recognizes the keyword "remote control" and searches the database for the relevant record. If the search result is, for example, "It's on the living room table," it will communicate this to the user via voice output. In this way, the present invention aims to streamline the process of finding objects in daily life.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The camera acquires images of the room. The camera takes pictures of the room's conditions at regular intervals and transmits the image data to a server via wireless communication.
[0042] Step 2:
[0043] The server passes the received image data to the analysis unit. The analysis unit uses a generative artificial intelligence model to recognize objects in the image, identify the type and characteristics of each object, and classify them.
[0044] Step 3:
[0045] The server stores the analysis results in a database using recording devices. The information stored includes the name of the object, identified features, position coordinates in the image, and the date and time of acquisition.
[0046] Step 4:
[0047] The user uses their device to search for the location of an object. The user enters a question into the device's application or voice assistant, asking about the location of a specific object.
[0048] Step 5:
[0049] The device sends the user's question to the server. The question is sent to the server as text or audio data.
[0050] Step 6:
[0051] The server analyzes the question using natural language processing technology. The purpose of the analysis is to identify the object the user is looking for. Based on the analysis results, the server searches for information in the database.
[0052] Step 7:
[0053] The server generates search results and retrieves the latest location information for the relevant objects. It then constructs a response based on this location information.
[0054] Step 8:
[0055] The server sends a response to the terminal. The response contains location information of the object, and this information is conveyed to the user.
[0056] Step 9:
[0057] The device notifies the user of the search results. The user can then confirm the object's location information via the device's screen display or audio output.
[0058] (Example 1)
[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0060] Conventional inventory management systems struggled to accurately track the location of items and provide that information quickly. Furthermore, they lacked efficient methods for responding to user inquiries.
[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0062] In this invention, the server includes an acquisition means for acquiring images of objects installed in a room environment, a means for analyzing the acquired images to recognize and classify the objects, and a registration means for recording and storing spatial information of the analyzed objects. This enables accurate location tracking of objects and allows for quick and accurate responses to user inquiries.
[0063] "Acquisition means" refers to a device installed within the room environment for acquiring images of objects.
[0064] "Analysis means" refers to the technology and processes used to recognize and classify objects using acquired images.
[0065] A "registration mechanism" is a system for recording and storing the spatial information of an analyzed item.
[0066] A "search tool" is a system equipped with the function of searching for recorded spatial information based on inquiries from users and providing responses.
[0067] "Wireless communication means" refers to wireless technology and equipment used to transmit acquired image data to a server.
[0068] A "generative artificial intelligence model" is a form of AI technology used to recognize objects in images.
[0069] "Spatial information" refers to data about the position of objects obtained through image analysis, indicating their arrangement within a room.
[0070] The present invention is configured as follows in an embodiment for specifically implementing an item management system. The room environment is equipped with devices such as cameras and sensors as acquisition means. These acquisition means periodically acquire images of items in the room and transmit the image data to a server using wireless communication means.
[0071] The server uses analysis software incorporating a generative AI model to analyze the received image data. This analysis allows the server to recognize objects in the image and classify their type and location. The AI model uses machine learning algorithms to tag the objects.
[0072] After analysis, the server uses a registration mechanism to save the spatial information of the extracted items to a database. This saved data allows the server to track the location of the items.
[0073] When a user wants to know the location of a specific item, they send a question to the server via voice or text through their device (e.g., a smartphone or voice assistant). For example, if a user voice-inputs "Where is the remote control?", the device sends this information to the server.
[0074] The server uses natural language processing technology to analyze the user's query and identify keywords for the search. Based on this, the server retrieves the latest spatial information of the relevant item from the database and provides a response to the user using search tools.
[0075] Finally, the terminal provides the user with a response from the server via screen display or audio output. This allows the user to quickly find the location of the desired item. In this way, the system of the present invention streamlines the search for items in daily life and improves convenience.
[0076] Example of a prompt:
[0077] "Identify the items in the room, record their locations, and report them."
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] An acquisition device installed in the room environment takes images of objects. The input for this step is the current state of the room, and the output is the captured image data. Specifically, the camera automatically takes pictures at regular time intervals.
[0081] Step 2:
[0082] The terminal uses wireless communication to send captured image data to the server. The input in this step is the acquired image data, and the output is the image data transferred to the server. The terminal transfers the image data to the server using a specific protocol, such as Wi-Fi.
[0083] Step 3:
[0084] The server processes the received image using an analysis tool. The input to this step is the image data sent to the server, and the output is the data of the analyzed objects. Specifically, a generative AI model is used to recognize objects in the image and classify their type and location.
[0085] Step 4:
[0086] The server processes the analyzed data and extracts spatial information. The input for this step is the data of the analyzed objects, and the output is the coordinates of the objects and other spatial data. The AI model identifies the location of objects in the image using pixel coordinates and converts them to coordinates in real space.
[0087] Step 5:
[0088] The server saves the analysis results to the database using a registration mechanism. The input for this step is item data including spatial information, and the output is the updated database. The server establishes a database connection and inserts or updates data using SQL or similar methods.
[0089] Step 6:
[0090] The user inquires about the location of an item via a terminal. The input is a voice or text inquiry from the user, and the output is the request. The terminal performs speech recognition and sends the inquiry to the server in text format.
[0091] Step 7:
[0092] The server uses natural language processing to analyze the user's inquiry. The input is the user's request, and the output is keywords for the search. The server uses an NLP engine to convert the voice query into an appropriate search query.
[0093] Step 8:
[0094] The server searches the database to retrieve the latest spatial information of an item. The input is a search query, and the output is the item's location information. SQL queries are used to retrieve the latest relevant information from the database.
[0095] Step 9:
[0096] The server transmits the acquired information to the terminal. The input is the location information of an item, and the output is an information packet containing that information. The server encodes the data via a communication protocol and transmits it to the terminal.
[0097] Step 10:
[0098] The device provides the user with location information for an item. The input in this step is information from the server, and the output is visual or audible feedback to the user. Specifically, the device may display text on the screen, or a voice assistant may read the content aloud.
[0099] (Application Example 1)
[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0101] Managing and tracking objects within a factory is fraught with human error and inefficiency issues. Especially when dealing with a large number of parts or products, identifying their locations and managing inventory can be time-consuming and labor-intensive, leading to decreased production efficiency. There is a need to solve this problem and achieve efficient and automated object management and location tracking.
[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0103] In this invention, the server includes an acquisition means installed in the environment to acquire visual information of objects; an analysis means for analyzing the acquired visual information to recognize and classify objects; a recording means for recording and storing the location information of the analyzed objects; a search means for searching the recorded location information based on input from a user and providing a response; and a control means for detecting shortages of objects and controlling corresponding operations. This enables automatic object recognition, real-time tracking of location information, efficient inventory management, and replenishment activities.
[0104] "Means of acquisition" refers to devices installed in the environment to acquire visual information.
[0105] "Analysis means" refers to a device or system that has the function of analyzing acquired visual information and recognizing and classifying objects.
[0106] "Recording means" refers to a device or database system for recording and storing location information of an analyzed object.
[0107] "Search means" refers to a device or system that has the function of searching for location information recorded based on input from a user and providing a response.
[0108] A "control means" is a device or system that has the function of detecting a shortage of an object and controlling the corresponding work automatically.
[0109] The system for carrying out this invention mainly consists of a complex group of devices for acquiring, analyzing, recording, retrieving, and controlling visual information of objects. First, acquisition means installed in the environment periodically acquire visual information of objects using cameras and sensors. The acquired data is transmitted to a server via wireless communication.
[0110] The server processes the received data using analysis tools and recognizes and classifies objects using a generative AI model. Specifically, it analyzes visual data using the OpenCV image processing library, and a machine learning model using TENSORFLOW® identifies the type of each object. At this time, recognized objects are tagged, and their location information is extracted as coordinate data.
[0111] The recording mechanism is responsible for storing these analysis results in a database, for example, using MySQL (registered trademark) to record the latest location information of objects. This information can also be used later as historical data.
[0112] The search mechanism analyzes voice or text input transmitted from the user via the terminal and retrieves the location information of the specified object from a database. The NLTK library is used for natural language processing to understand the user's questions and generate appropriate responses. Location information is presented to the user's terminal in real time, either audibly or visually.
[0113] The control system monitors the shortage of objects identified by the analysis system and automatically controls the necessary tasks. For example, when there is a shortage of a particular type or quantity of objects in a factory, a robot automatically replenishes the necessary items from the stockroom. In this case, ROS (Robotics Operating System) is used as the robotics operating system.
[0114] As a concrete example, in response to a user's question, "Where are the bolts?", a specific answer such as "They are on biomaterial shelf 3" is provided. An example of a prompt for the generating AI model is, "Image of factory line with parts and tools. Identify and classify items, then record their positions in the database." In this way, object management is automated through a combination of devices and software.
[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0116] Step 1:
[0117] The device acquires visual information about objects through cameras and sensors installed in the environment. This visual data is generated in real time from the camera and becomes input data transmitted to the server via wireless communication.
[0118] Step 2:
[0119] The server receives the acquired data and analyzes the images using OpenCV. First, preprocessing such as noise reduction and grayscale conversion is performed, and then object recognition and classification are carried out using the generative AI model TensorFlow. The output obtained here is the type of recognized object and its position coordinates within the visual field.
[0120] Step 3:
[0121] The server saves the analysis results to a database via a recording mechanism. The inputs used are object type, location coordinates, and tagging information, and this data is recorded in a MySQL database. The output is a database entry for future searches.
[0122] Step 4:
[0123] The user asks a question from their device using voice or text input. For example, a query such as "Where are the screws?" is sent to the server. The input here is a natural language request from the user.
[0124] Step 5:
[0125] The server uses the NLTK library to analyze user queries. It breaks down the natural language input into intent and constructs a search query. The output consists of object names and related information.
[0126] Step 6:
[0127] The server uses a search mechanism to retrieve the latest object location information from the database. The input is the object name obtained in step 5, and the database query is executed to output the location information.
[0128] Step 7:
[0129] The terminal receives location information sent back from the server and notifies the user visually or audibly. The output is specific location information, such as "It's on shelf 2 in warehouse C." This allows the user to quickly determine the object's location.
[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0131] As an embodiment of the present invention, a system is presented that combines an emotion engine with an object management system to recognize the user's emotions and change its response accordingly. First, an image acquisition means installed in the room environment captures images of the room's conditions at regular intervals. The acquired image data is transmitted to a server via wireless communication means.
[0132] The server uses analysis tools to process the provided images using a generating artificial intelligence model. This model recognizes and classifies objects in the images, identifying their characteristics and locations. The classified information is stored in a database by recording tools, allowing for object referencing within the database.
[0133] When a user wants to know the location of an object, they send a search request to the server using their device via voice or text input. The server analyzes the received input using natural language processing and then searches the database for the corresponding object. A new feature introduced here is the emotion engine, which analyzes the user's input voice data and recognizes the user's emotional state.
[0134] The emotion engine assesses the user's emotional state and adjusts its response accordingly. For example, if it detects that the user is anxious, it might soften the response or emphasize urgency. It also selects a method for presenting object location information based on the user's emotions. One example is whether to provide detailed information calmly or to quickly convey only the key points.
[0135] The server generates search results and a response tailored to the user's mood and sends it to the terminal. The terminal displays the results to the user visually or outputs them audibly. If the user asks "Where is the remote control?" through the terminal, the server searches the database for the remote control's location and notifies the user audibly, "It's on the living room table." If the terminal detects from the user's tone of voice that they are in a hurry, the terminal's notification will be expedited.
[0136] Thus, the present invention aims to further improve the efficiency of object searching by providing a more user-friendly interface that takes into account user emotions in addition to object position recognition.
[0137] The following describes the processing flow.
[0138] Step 1:
[0139] The camera captures images of the room. The camera installed in the room periodically takes pictures and sends the image data to the server.
[0140] Step 2:
[0141] The server analyzes the received image. Using analysis tools, the server analyzes the image based on a generative artificial intelligence model to recognize and classify objects.
[0142] Step 3:
[0143] The server records the object's location information in a database. The characteristics and location information of the recognized object are saved in the database for later retrieval.
[0144] Step 4:
[0145] The user operates the device to request an object search. The user asks for the location of a specific object via voice or text through the device's interface.
[0146] Step 5:
[0147] The terminal sends user input to the server. User questions are sent from the terminal to the server as data.
[0148] Step 6:
[0149] The server analyzes the user's question. The server uses natural language processing to understand the user's intent and searches the database for the relevant object.
[0150] Step 7:
[0151] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's voice data to determine their emotional state.
[0152] Step 8:
[0153] The server adjusts its response based on the user's emotional state. For example, if the user is irritated, the response will be changed to a more calming tone.
[0154] Step 9:
[0155] The server sends the search results to the terminal. The adjusted response, along with the object's location information, is transmitted to the terminal.
[0156] Step 10:
[0157] The device notifies the user of the results. The device displays the location information on the screen or reports it to the user via voice. The user receives the notification from the device and confirms the object's location.
[0158] (Example 2)
[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0160] Conventional object management systems have the problem of failing to adequately reduce the user's psychological burden because they do not take into account the user's emotional state when locating objects in an environment. Furthermore, there is a lack of methods to provide a more user-friendly interface in addition to simply providing location information for objects.
[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0162] In this invention, the server includes data acquisition means for acquiring information about objects installed in the environment, analysis means for analyzing the acquired information and identifying and classifying objects, recording means for storing location information of the analyzed objects, search means for finding the recorded location information based on user information and providing a response, and emotion analysis means for analyzing the user's emotional state and adjusting the response. This makes it possible to provide the location of objects in the environment appropriately and with consideration for the user's emotions.
[0163] "Environment" refers to the place or situation in which an object exists within a specific space or set of conditions.
[0164] "Data acquisition means" refers to a function or device for collecting and recording target information or data.
[0165] "Analysis means" refers to a function or device for understanding acquired data and identifying, classifying, and evaluating necessary information.
[0166] "Recording means" refers to a function or device for organizing information and data and storing them so that they can be used later.
[0167] "Search method" refers to a function or device that finds the necessary content from recorded information and provides it to the user.
[0168] "Emotional analysis means" refers to a function or device that determines the emotional state of a user from their input and actions and adjusts the system's response accordingly.
[0169] A "generative model" refers to an algorithm that uses machine learning or artificial intelligence to learn data patterns and then performs inferences on new data.
[0170] "Natural language processing" refers to the technology that enables computers to understand and process human language appropriately.
[0171] This invention provides a novel user interface that combines an object management system with emotion analysis functionality. The system mainly consists of a server, terminals, and an image acquisition device.
[0172] First, the terminal uses a data acquisition device installed in the environment, such as a room, to capture images of objects and the environment at specific time intervals. This image acquisition device consists of hardware such as a camera, and transmits the acquired image data to a server wirelessly.
[0173] The server analyzes the received image data using a generative artificial intelligence model (generative AI model). This model has the function of identifying and classifying objects contained in the image. The analysis results are stored in a database through recording means and can be searched as needed.
[0174] When a user needs information about an object, they can send a natural language prompt to the server via their device, for example, "Tell me where the remote control is." The server uses natural language processing technology to analyze this prompt and retrieve information about the corresponding object from its database.
[0175] Furthermore, this system incorporates emotion analysis capabilities. The server can analyze the user's voice input and recognize their emotional state. For example, if it determines that the user is anxious, it can adjust the format and pace of the information provided.
[0176] In this way, users not only receive information but also convenient responses tailored to their psychological state. This system achieves a more user-friendly experience by incorporating emotional states into object position recognition.
[0177] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0178] Step 1:
[0179] The terminal periodically captures images of the environment using image acquisition devices such as cameras installed in the room. The input is image data captured by the camera. The output is image data transmitted to a server via wireless communication. Specifically, this process acquires visual information about objects within the environment.
[0180] Step 2:
[0181] The server passes the received image data to a generative AI model for analysis. The input here is image data sent from the terminal. The generative AI model uses this data to identify and classify objects within the image. The output is data that identifies the features and location information of the objects. Specifically, the model extracts object labels and location coordinates.
[0182] Step 3:
[0183] The server stores the analyzed object features and location information in a database. The input is the object feature information created by the generative AI model. The output is information recorded in a format that can be searched later. Specifically, the process involves saving the information and adding it to the database.
[0184] Step 4:
[0185] The user sends prompts to the server via voice or text using a terminal to find the location of an object. The input is a natural language query from the user. The output is a search request to the server. For example, the user might give instructions such as "Tell me where the remote control is."
[0186] Step 5:
[0187] The server parses the received prompt message using natural language processing and searches the database for related objects. The input is the prompt message sent by the user. This prompt message is parsed to retrieve information about the corresponding object in the database. The output is the object's location information. Specifically, the parsing engine interprets the intent of the query and retrieves the relevant data.
[0188] Step 6:
[0189] The server analyzes the user's voice data using emotion analysis tools to recognize their emotional state. The input is the user's voice data. Through analysis, it determines the user's emotional state. The output is response adjustment information based on the emotional state. For example, it might detect anxiety from the user's voice.
[0190] Step 7:
[0191] The server generates a response based on the location and emotional state of the acquired object and sends it to the terminal. The input is adjustment information based on the object's location and emotional state. The output is the adjusted response provided to the user. Specifically, the information is presented in audio or visual form, and the tone and amount of information are changed as needed.
[0192] Step 8:
[0193] The terminal provides the user with the response received from the server. Input is the pre-arranged response sent from the server. Output is information presented to the user visually or audibly. Specifically, it provides notifications in a timely and appropriate manner depending on the user's status.
[0194] (Application Example 2)
[0195] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0196] Conventional object management systems, when providing location information for objects, often respond mechanically without considering the user's emotional state, making it difficult to provide a user-satisfying interface. Furthermore, while real-time, rapid, and accurate information is required for product searches within stores, there is a lack of consideration for the customer's emotions. This can lead to increased customer stress and reduced purchasing intent.
[0197] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0198] In this invention, the server is installed in a room environment and includes an image acquisition means for acquiring images of objects, an analysis means for analyzing the images acquired by the acquisition means and recognizing and classifying objects, a recording means for recording and storing the location information of the analyzed objects, a search means for searching the recorded location information based on input from the user and providing a response, and an emotion analysis means for analyzing the user's emotional state and optimizing the response. This enables the provision of appropriate information according to the user's emotions, realizes a more user-friendly interface, and allows for efficient product searching within a store.
[0199] "Image acquisition means" refers to devices or functions installed in a room environment for acquiring images of objects.
[0200] "Analysis means" refers to devices or functions that analyze acquired image data and recognize and classify objects.
[0201] "Recording means" refers to a device or function for recording and storing the positional information of an analyzed object.
[0202] A "search means" refers to a device or function that searches for location information recorded based on input from a user and provides a response.
[0203] "Emotional analysis tools" refer to devices or functions that analyze a user's emotional state and optimize their response based on that evaluation.
[0204] A "generative artificial intelligence model" is a model based on artificial intelligence technology used for recognizing and analyzing image data, and it improves accuracy using machine learning techniques.
[0205] "Natural language processing" is a technology that analyzes the user's voice input, understands its meaning, and responds accordingly.
[0206] This system is a technology for managing objects within rooms and stores and providing information tailored to the user's emotions. The system mainly consists of a server, user terminals (smart glasses or smartphones), and image acquisition devices installed in the room environment.
[0207] The server utilizes a generative AI model to recognize and classify objects from received image data. This analysis employs computer vision technologies such as OpenCV and TensorFlow. The location information of the analyzed objects is stored in a SQL database. When a user wants to know the location of an object, they can query the server using voice or text via their device. Natural language processing, such as Google Cloud Natural Language API, is used to analyze the user's query.
[0208] Furthermore, the server uses emotion analysis tools to determine the user's emotional state. It analyzes whether the user is in a hurry or relaxed based on their tone of voice and text expression. Based on this information, it provides responses not simply as data, but in a way that is appropriate to the situation.
[0209] As a concrete example, consider a scenario where a user asks, "Where is the shaved ice machine?" The server identifies the location of the product from the analyzed database information. If the emotion analysis determines that the user is anxious, it immediately responds with a voice message; if the user is calm, it displays detailed map information on the smart glasses. This allows the user to easily find the product within the store.
[0210] Possible prompts to input into a generative AI model include the following:
[0211] "I'm looking for this product. Please tell me its location."
[0212] "I'm in a hurry. Please show me the way quickly."
[0213] This system will allow users to have a more efficient and satisfying experience.
[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0215] Step 1:
[0216] The terminal sends instructions to the image acquisition device to acquire images of the room environment or store interior. The image acquisition device uses a camera to take still images at predetermined time intervals. The captured image data is transmitted to the server via wireless communication. The input is the capture instruction, and the output is the acquired image data.
[0217] Step 2:
[0218] The server feeds the received image data into an AI image analysis model (for example, using TensorFlow). Here, objects in the image are recognized and classified. The generative AI model extracts the features of the objects and determines their type based on these features. The input is image data, and the output is the type of object and its location information.
[0219] Step 3:
[0220] The server saves the object location information obtained as a result of the analysis to an SQL database. Before saving, the data integrity is checked, and data processing is performed to remove duplicates and other errors. The input is the analyzed object information, and the output is the updated database.
[0221] Step 4:
[0222] The user asks about the location of an object via voice or text through the terminal. The voice data is initially processed within the terminal and then sent to the server in text format. The input is the user's voice or text, and the output is the question in text format.
[0223] Step 5:
[0224] The server analyzes the received question using a natural language processing (NLP) engine (such as the Google Cloud Natural Language API) and retrieves the necessary object information from a SQL database. The input is a text question from the user, and the output is the location information of the searched object.
[0225] Step 6:
[0226] The server uses sentiment analysis to identify the user's emotional state based on the content of the user's input voice and text. It analyzes the user's emotional state, such as whether they are anxious or calm, based on the tone of their voice and the content of their text. The input is the user's voice / text data, and the output is their emotional state.
[0227] Step 7:
[0228] The server generates an optimized response based on the user's emotional state. For example, if the user is in a hurry, it will quickly provide the object's location via voice. If the user is calm, it will provide a detailed explanation or map information. The input is the emotional state and the object's location, and the output is the response message.
[0229] Step 8:
[0230] The server generates a response, which is then sent to the terminal and presented to the user. The terminal communicates the result through speech synthesis or display. The input is the response message, and the output is the provision of visual or auditory information to the user.
[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0234] [Second Embodiment]
[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0247] As an embodiment for carrying out the present invention, the object management system is specifically configured as follows. First, an image acquisition means installed in the room environment periodically photographs the room's conditions and acquires images of objects. These images are transmitted to a server via wireless communication means.
[0248] The server processes the received images using analysis tools and recognizes and classifies objects in the images using a generative artificial intelligence model. This process identifies the type and characteristics of each object and tags them. It also extracts the object's location information as coordinates within the image.
[0249] Next, the server uses recording devices to record the analyzed object information and its location information in a database. This makes it possible to track the object's latest location and also use it as historical information.
[0250] When a user wants to know the location of an object, they send a question to the server via voice or text through their device (smartphone or voice assistant). The server analyzes this input using natural language processing technology to determine which object to search for.
[0251] The server uses a search mechanism to retrieve the latest location information of the target object from the database and sends the search results back to the terminal. The terminal then communicates the search results to the user via screen display or audio output. This allows the user to quickly learn the object's current location.
[0252] For example, if a user uses their smartphone to voice-input "Where is the remote control?", the server recognizes the keyword "remote control" and searches the database for the relevant record. If the search result is, for example, "It's on the living room table," it will communicate this to the user via voice output. In this way, the present invention aims to streamline the process of finding objects in daily life.
[0253] The following describes the processing flow.
[0254] Step 1:
[0255] The camera acquires images of the room. The camera takes pictures of the room's conditions at regular intervals and transmits the image data to a server via wireless communication.
[0256] Step 2:
[0257] The server passes the received image data to the analysis unit. The analysis unit uses a generative artificial intelligence model to recognize objects in the image, identify the type and characteristics of each object, and classify them.
[0258] Step 3:
[0259] The server stores the analysis results in a database using recording devices. The information stored includes the name of the object, identified features, position coordinates in the image, and the date and time of acquisition.
[0260] Step 4:
[0261] The user uses their device to search for the location of an object. The user enters a question into the device's application or voice assistant, asking about the location of a specific object.
[0262] Step 5:
[0263] The device sends the user's question to the server. The question is sent to the server as text or audio data.
[0264] Step 6:
[0265] The server analyzes the question using natural language processing technology. The purpose of the analysis is to identify the object the user is looking for. Based on the analysis results, the server searches for information in the database.
[0266] Step 7:
[0267] The server generates search results and retrieves the latest location information for the relevant objects. It then constructs a response based on this location information.
[0268] Step 8:
[0269] The server sends a response to the terminal. The response contains location information of the object, and this information is conveyed to the user.
[0270] Step 9:
[0271] The device notifies the user of the search results. The user can then confirm the object's location information via the device's screen display or audio output.
[0272] (Example 1)
[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0274] Conventional inventory management systems struggled to accurately track the location of items and provide that information quickly. Furthermore, they lacked efficient methods for responding to user inquiries.
[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0276] In this invention, the server includes an acquisition means for acquiring images of objects installed in a room environment, a means for analyzing the acquired images to recognize and classify the objects, and a registration means for recording and storing spatial information of the analyzed objects. This enables accurate location tracking of objects and allows for quick and accurate responses to user inquiries.
[0277] "Acquisition means" refers to a device installed within the room environment for acquiring images of objects.
[0278] "Analysis means" refers to the technology and processes used to recognize and classify objects using acquired images.
[0279] A "registration mechanism" is a system for recording and storing the spatial information of an analyzed item.
[0280] A "search tool" is a system equipped with the function of searching for recorded spatial information based on inquiries from users and providing responses.
[0281] "Wireless communication means" refers to wireless technology and equipment used to transmit acquired image data to a server.
[0282] The "generative AI model" is a form of AI technology used to recognize objects in images.
[0283] "Spatial information" is data related to the position of an object obtained by image analysis, indicating its placement within a room.
[0284] The present invention is configured as follows as a form for specifically implementing an object management system. In the room environment, devices such as cameras and sensors are installed as acquisition means. This acquisition means periodically acquires the objects in the room as images and transmits the image data to the server using wireless communication means.
[0285] The server uses analysis software incorporating a generative AI model to analyze the received image data. By this analysis means, the server recognizes the objects in the image and classifies their types and positions. The AI model tags the objects using machine learning algorithms.
[0286] After analysis, the server uses registration means to save the spatial information of the extracted objects in the database. Based on this saved data, the server can track the position information of the objects.
[0287] When the user wants to know the position of a specific object, they send a query to the server via a terminal (e.g., smartphone or voice assistant) either by voice or text. For example, when the user inputs "Where is the remote control?" by voice, the terminal sends this information to the server.
[0288] The server analyzes the user's query using natural language processing technology to identify keywords for searching. Thereby, the server searches the database for the latest spatial information of the corresponding object and uses exploration means to provide a response to the user.
[0289] Finally, the terminal provides the user with a response from the server via screen display or audio output. This allows the user to quickly find the location of the desired item. In this way, the system of the present invention streamlines the search for items in daily life and improves convenience.
[0290] Example of a prompt:
[0291] "Identify the items in the room, record their locations, and report them."
[0292] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0293] Step 1:
[0294] An acquisition device installed in the room environment takes images of objects. The input for this step is the current state of the room, and the output is the captured image data. Specifically, the camera automatically takes pictures at regular time intervals.
[0295] Step 2:
[0296] The terminal uses wireless communication to send captured image data to the server. The input in this step is the acquired image data, and the output is the image data transferred to the server. The terminal transfers the image data to the server using a specific protocol, such as Wi-Fi.
[0297] Step 3:
[0298] The server processes the received image using an analysis tool. The input to this step is the image data sent to the server, and the output is the data of the analyzed objects. Specifically, a generative AI model is used to recognize objects in the image and classify their type and location.
[0299] Step 4:
[0300] The server processes the analyzed data and extracts spatial information. The input for this step is the data of the analyzed item, and the output is the coordinates of the item and other spatial data. The AI model identifies the position of the item in the image in pixel coordinates and converts it to real-space coordinates.
[0301] Step 5:
[0302] The server saves the analysis results in the database using the registration means. The input for this step is the item data including spatial information, and the output is the updated database. The server establishes a database connection and uses SQL or the like to insert or update the data.
[0303] Step 6:
[0304] The user inquires about the position of the item via the terminal. The input is the voice or text query by the user, and the output is the request. The terminal performs voice recognition and sends the query to the server in text format.
[0305] Step 7:
[0306] The server analyzes the user's query using natural language processing. The input is the user's request, and the output is the keyword for searching. The server uses the NLP engine to convert the voice query into an appropriate search query.
[0307] Step 8:
[0308] The server searches the database and obtains the latest spatial information of the item. The input is the search query, and the output is the position information of the item. Use the SQL query to obtain the latest relevant information in the database.
[0309] Step 9:
[0310] The server transmits the acquired information to the terminal. The input is the location information of an item, and the output is an information packet containing that information. The server encodes the data via a communication protocol and transmits it to the terminal.
[0311] Step 10:
[0312] The device provides the user with location information for an item. The input in this step is information from the server, and the output is visual or audible feedback to the user. Specifically, the device may display text on the screen, or a voice assistant may read the content aloud.
[0313] (Application Example 1)
[0314] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0315] Managing and tracking objects within a factory is fraught with human error and inefficiency issues. Especially when dealing with a large number of parts or products, identifying their locations and managing inventory can be time-consuming and labor-intensive, leading to decreased production efficiency. There is a need to solve this problem and achieve efficient and automated object management and location tracking.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0317] In this invention, the server includes an acquisition means installed in the environment to acquire visual information of objects; an analysis means for analyzing the acquired visual information to recognize and classify objects; a recording means for recording and storing the location information of the analyzed objects; a search means for searching the recorded location information based on input from a user and providing a response; and a control means for detecting shortages of objects and controlling corresponding operations. This enables automatic object recognition, real-time tracking of location information, efficient inventory management, and replenishment activities.
[0318] "Means of acquisition" refers to devices installed in the environment to acquire visual information.
[0319] "Analysis means" refers to a device or system that has the function of analyzing acquired visual information and recognizing and classifying objects.
[0320] "Recording means" refers to a device or database system for recording and storing location information of an analyzed object.
[0321] "Search means" refers to a device or system that has the function of searching for location information recorded based on input from a user and providing a response.
[0322] A "control means" is a device or system that has the function of detecting a shortage of an object and controlling the corresponding work automatically.
[0323] The system for carrying out this invention mainly consists of a complex group of devices for acquiring, analyzing, recording, retrieving, and controlling visual information of objects. First, acquisition means installed in the environment periodically acquire visual information of objects using cameras and sensors. The acquired data is transmitted to a server via wireless communication.
[0324] The server processes the received data using analysis tools and recognizes and classifies objects using a generative AI model. Specifically, it analyzes visual data using the OpenCV image processing library, and a machine learning model using TensorFlow identifies the type of each object. At this time, recognized objects are tagged, and their location information is extracted as coordinate data.
[0325] The recording mechanism is responsible for saving these analysis results to a database, such as using MySQL to record the latest location information of objects. This information can later be used as historical data.
[0326] The search mechanism analyzes voice or text input transmitted from the user via the terminal and retrieves the location information of the specified object from a database. The NLTK library is used for natural language processing to understand the user's questions and generate appropriate responses. Location information is presented to the user's terminal in real time, either audibly or visually.
[0327] The control system monitors the shortage of objects identified by the analysis system and automatically controls the necessary tasks. For example, when there is a shortage of a particular type or quantity of objects in a factory, a robot automatically replenishes the necessary items from the stockroom. In this case, ROS (Robotics Operating System) is used as the robotics operating system.
[0328] As a concrete example, in response to a user's question, "Where are the bolts?", a specific answer such as "They are on biomaterial shelf 3" is provided. An example of a prompt for the generating AI model is, "Image of factory line with parts and tools. Identify and classify items, then record their positions in the database." In this way, object management is automated through a combination of devices and software.
[0329] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0330] Step 1:
[0331] The device acquires visual information about objects through cameras and sensors installed in the environment. This visual data is generated in real time from the camera and becomes input data transmitted to the server via wireless communication.
[0332] Step 2:
[0333] The server receives the acquired data and analyzes the images using OpenCV. First, preprocessing such as noise reduction and grayscale conversion is performed, and then object recognition and classification are carried out using the generative AI model TensorFlow. The output obtained here is the type of recognized object and its position coordinates within the visual field.
[0334] Step 3:
[0335] The server saves the analysis results to a database via a recording mechanism. The inputs used are object type, location coordinates, and tagging information, and this data is recorded in a MySQL database. The output is a database entry for future searches.
[0336] Step 4:
[0337] The user asks a question from their device using voice or text input. For example, a query such as "Where are the screws?" is sent to the server. The input here is a natural language request from the user.
[0338] Step 5:
[0339] The server uses the NLTK library to analyze user queries. It breaks down the natural language input into intent and constructs a search query. The output consists of object names and related information.
[0340] Step 6:
[0341] The server uses a search mechanism to retrieve the latest object location information from the database. The input is the object name obtained in step 5, and the database query is executed to output the location information.
[0342] Step 7:
[0343] The terminal receives location information sent back from the server and notifies the user visually or audibly. The output is specific location information, such as "It's on shelf 2 in warehouse C." This allows the user to quickly determine the object's location.
[0344] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0345] As an embodiment of the present invention, a system is presented that combines an emotion engine with an object management system to recognize the user's emotions and change its response accordingly. First, an image acquisition means installed in the room environment captures images of the room's conditions at regular intervals. The acquired image data is transmitted to a server via wireless communication means.
[0346] The server uses analysis tools to process the provided images using a generating artificial intelligence model. This model recognizes and classifies objects in the images, identifying their characteristics and locations. The classified information is stored in a database by recording tools, allowing for object referencing within the database.
[0347] When a user wants to know the location of an object, they send a search request to the server using their device via voice or text input. The server analyzes the received input using natural language processing and then searches the database for the corresponding object. A new feature introduced here is the emotion engine, which analyzes the user's input voice data and recognizes the user's emotional state.
[0348] The emotion engine assesses the user's emotional state and adjusts its response accordingly. For example, if it detects that the user is anxious, it might soften the response or emphasize urgency. It also selects a method for presenting object location information based on the user's emotions. One example is whether to provide detailed information calmly or to quickly convey only the key points.
[0349] The server generates search results and a response tailored to the user's mood and sends it to the terminal. The terminal displays the results to the user visually or outputs them audibly. If the user asks "Where is the remote control?" through the terminal, the server searches the database for the remote control's location and notifies the user audibly, "It's on the living room table." If the terminal detects from the user's tone of voice that they are in a hurry, the terminal's notification will be expedited.
[0350] Thus, the present invention aims to further improve the efficiency of object searching by providing a more user-friendly interface that takes into account user emotions in addition to object position recognition.
[0351] The following describes the processing flow.
[0352] Step 1:
[0353] The camera captures images of the room. The camera installed in the room periodically takes pictures and sends the image data to the server.
[0354] Step 2:
[0355] The server analyzes the received image. Using analysis tools, the server analyzes the image based on a generative artificial intelligence model to recognize and classify objects.
[0356] Step 3:
[0357] The server records the object's location information in a database. The characteristics and location information of the recognized object are saved in the database for later retrieval.
[0358] Step 4:
[0359] The user operates the device to request an object search. The user asks for the location of a specific object via voice or text through the device's interface.
[0360] Step 5:
[0361] The terminal sends user input to the server. User questions are sent from the terminal to the server as data.
[0362] Step 6:
[0363] The server analyzes the user's question. The server uses natural language processing to understand the user's intent and searches the database for the relevant object.
[0364] Step 7:
[0365] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's voice data to determine their emotional state.
[0366] Step 8:
[0367] The server adjusts its response based on the user's emotional state. For example, if the user is irritated, the response will be changed to a more calming tone.
[0368] Step 9:
[0369] The server sends the search results to the terminal. The adjusted response, along with the object's location information, is transmitted to the terminal.
[0370] Step 10:
[0371] The device notifies the user of the results. The device displays the location information on the screen or reports it to the user via voice. The user receives the notification from the device and confirms the object's location.
[0372] (Example 2)
[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0374] Conventional object management systems have the problem of failing to adequately reduce the user's psychological burden because they do not take into account the user's emotional state when locating objects in an environment. Furthermore, there is a lack of methods to provide a more user-friendly interface in addition to simply providing location information for objects.
[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0376] In this invention, the server includes data acquisition means for acquiring information about objects installed in the environment, analysis means for analyzing the acquired information and identifying and classifying objects, recording means for storing location information of the analyzed objects, search means for finding the recorded location information based on user information and providing a response, and emotion analysis means for analyzing the user's emotional state and adjusting the response. This makes it possible to provide the location of objects in the environment appropriately and with consideration for the user's emotions.
[0377] "Environment" refers to the place or situation in which an object exists within a specific space or set of conditions.
[0378] "Data acquisition means" refers to a function or device for collecting and recording target information or data.
[0379] "Analysis means" refers to a function or device for understanding acquired data and identifying, classifying, and evaluating necessary information.
[0380] "Recording means" refers to a function or device for organizing information and data and storing them so that they can be used later.
[0381] "Search method" refers to a function or device that finds the necessary content from recorded information and provides it to the user.
[0382] "Emotional analysis means" refers to a function or device that determines the emotional state of a user from their input and actions and adjusts the system's response accordingly.
[0383] A "generative model" refers to an algorithm that uses machine learning or artificial intelligence to learn data patterns and then performs inferences on new data.
[0384] "Natural language processing" refers to the technology that enables computers to understand and process human language appropriately.
[0385] This invention provides a novel user interface that combines an object management system with emotion analysis functionality. The system mainly consists of a server, terminals, and an image acquisition device.
[0386] First, the terminal uses a data acquisition device installed in the environment, such as a room, to capture images of objects and the environment at specific time intervals. This image acquisition device consists of hardware such as a camera, and transmits the acquired image data to a server wirelessly.
[0387] The server analyzes the received image data using a generative artificial intelligence model (generative AI model). This model has the function of identifying and classifying objects contained in the image. The analysis results are stored in a database through recording means and can be searched as needed.
[0388] When a user needs information about an object, they can send a natural language prompt to the server via their device, for example, "Tell me where the remote control is." The server uses natural language processing technology to analyze this prompt and retrieve information about the corresponding object from its database.
[0389] Furthermore, this system incorporates emotion analysis capabilities. The server can analyze the user's voice input and recognize their emotional state. For example, if it determines that the user is anxious, it can adjust the format and pace of the information provided.
[0390] In this way, users not only receive information but also convenient responses tailored to their psychological state. This system achieves a more user-friendly experience by incorporating emotional states into object position recognition.
[0391] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0392] Step 1:
[0393] The terminal periodically captures images of the environment using image acquisition devices such as cameras installed in the room. The input is image data captured by the camera. The output is image data transmitted to a server via wireless communication. Specifically, this process acquires visual information about objects within the environment.
[0394] Step 2:
[0395] The server passes the received image data to a generative AI model for analysis. The input here is image data sent from the terminal. The generative AI model uses this data to identify and classify objects within the image. The output is data that identifies the features and location information of the objects. Specifically, the model extracts object labels and location coordinates.
[0396] Step 3:
[0397] The server stores the analyzed object features and location information in a database. The input is the object feature information created by the generative AI model. The output is information recorded in a format that can be searched later. Specifically, the process involves saving the information and adding it to the database.
[0398] Step 4:
[0399] The user sends prompts to the server via voice or text using a terminal to find the location of an object. The input is a natural language query from the user. The output is a search request to the server. For example, the user might give instructions such as "Tell me where the remote control is."
[0400] Step 5:
[0401] The server parses the received prompt message using natural language processing and searches the database for related objects. The input is the prompt message sent by the user. This prompt message is parsed to retrieve information about the corresponding object in the database. The output is the object's location information. Specifically, the parsing engine interprets the intent of the query and retrieves the relevant data.
[0402] Step 6:
[0403] The server analyzes the user's voice data using emotion analysis tools to recognize their emotional state. The input is the user's voice data. Through analysis, it determines the user's emotional state. The output is response adjustment information based on the emotional state. For example, it might detect anxiety from the user's voice.
[0404] Step 7:
[0405] The server generates a response based on the location and emotional state of the acquired object and sends it to the terminal. The input is adjustment information based on the object's location and emotional state. The output is the adjusted response provided to the user. Specifically, the information is presented in audio or visual form, and the tone and amount of information are changed as needed.
[0406] Step 8:
[0407] The terminal provides the user with the response received from the server. Input is the pre-arranged response sent from the server. Output is information presented to the user visually or audibly. Specifically, it provides notifications in a timely and appropriate manner depending on the user's status.
[0408] (Application Example 2)
[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0410] Conventional object management systems, when providing location information for objects, often respond mechanically without considering the user's emotional state, making it difficult to provide a user-satisfying interface. Furthermore, while real-time, rapid, and accurate information is required for product searches within stores, there is a lack of consideration for the customer's emotions. This can lead to increased customer stress and reduced purchasing intent.
[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0412] In this invention, the server is installed in a room environment and includes an image acquisition means for acquiring images of objects, an analysis means for analyzing the images acquired by the acquisition means and recognizing and classifying objects, a recording means for recording and storing the location information of the analyzed objects, a search means for searching the recorded location information based on input from the user and providing a response, and an emotion analysis means for analyzing the user's emotional state and optimizing the response. This enables the provision of appropriate information according to the user's emotions, realizes a more user-friendly interface, and allows for efficient product searching within a store.
[0413] "Image acquisition means" refers to devices or functions installed in a room environment for acquiring images of objects.
[0414] "Analysis means" refers to devices or functions that analyze acquired image data and recognize and classify objects.
[0415] "Recording means" refers to a device or function for recording and storing the positional information of an analyzed object.
[0416] A "search means" refers to a device or function that searches for location information recorded based on input from a user and provides a response.
[0417] "Emotional analysis tools" refer to devices or functions that analyze a user's emotional state and optimize their response based on that evaluation.
[0418] A "generative artificial intelligence model" is a model based on artificial intelligence technology used for recognizing and analyzing image data, and it improves accuracy using machine learning techniques.
[0419] "Natural language processing" is a technology that analyzes the user's voice input, understands its meaning, and responds accordingly.
[0420] This system is a technology for managing objects within rooms and stores and providing information tailored to the user's emotions. The system mainly consists of a server, user terminals (smart glasses or smartphones), and image acquisition devices installed in the room environment.
[0421] The server utilizes a generative AI model to recognize and classify objects from received image data. This analysis employs computer vision technologies such as OpenCV and TensorFlow. The location information of the analyzed objects is stored in a SQL database. When a user wants to know the location of an object, they query the server using voice or text via their device. Natural language processing, such as the Google Cloud Natural Language API, is used to analyze the user's query.
[0422] Furthermore, the server uses emotion analysis tools to determine the user's emotional state. It analyzes whether the user is in a hurry or relaxed based on their tone of voice and text expression. Based on this information, it provides responses not simply as data, but in a way that is appropriate to the situation.
[0423] As a concrete example, consider a scenario where a user asks, "Where is the shaved ice machine?" The server identifies the location of the product from the analyzed database information. If the emotion analysis determines that the user is anxious, it immediately responds with a voice message; if the user is calm, it displays detailed map information on the smart glasses. This allows the user to easily find the product within the store.
[0424] Possible prompts to input into a generative AI model include the following:
[0425] "I'm looking for this product. Please tell me its location."
[0426] "I'm in a hurry. Please show me the way quickly."
[0427] This system will allow users to have a more efficient and satisfying experience.
[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0429] Step 1:
[0430] The terminal sends instructions to the image acquisition device to acquire images of the room environment or store interior. The image acquisition device uses a camera to take still images at predetermined time intervals. The captured image data is transmitted to the server via wireless communication. The input is the capture instruction, and the output is the acquired image data.
[0431] Step 2:
[0432] The server feeds the received image data into an AI image analysis model (for example, using TensorFlow). Here, objects in the image are recognized and classified. The generative AI model extracts the features of the objects and determines their type based on these features. The input is image data, and the output is the type of object and its location information.
[0433] Step 3:
[0434] The server saves the object location information obtained as a result of the analysis to an SQL database. Before saving, the data integrity is checked, and data processing is performed to remove duplicates and other errors. The input is the analyzed object information, and the output is the updated database.
[0435] Step 4:
[0436] The user asks about the location of an object via voice or text through the terminal. The voice data is initially processed within the terminal and then sent to the server in text format. The input is the user's voice or text, and the output is the question in text format.
[0437] Step 5:
[0438] The server analyzes the received question using a natural language processing (NLP) engine (such as the Google Cloud Natural Language API) and retrieves the necessary object information from a SQL database. The input is a text question from the user, and the output is the location information of the searched object.
[0439] Step 6:
[0440] The server uses sentiment analysis to identify the user's emotional state based on the content of the user's input voice and text. It analyzes the user's emotional state, such as whether they are anxious or calm, based on the tone of their voice and the content of their text. The input is the user's voice / text data, and the output is their emotional state.
[0441] Step 7:
[0442] The server generates an optimized response based on the user's emotional state. For example, if the user is in a hurry, it will quickly provide the object's location via voice. If the user is calm, it will provide a detailed explanation or map information. The input is the emotional state and the object's location, and the output is the response message.
[0443] Step 8:
[0444] The server generates a response, which is then sent to the terminal and presented to the user. The terminal communicates the result through speech synthesis or display. The input is the response message, and the output is the provision of visual or auditory information to the user.
[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0448] [Third Embodiment]
[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0461] As an embodiment for carrying out the present invention, the object management system is specifically configured as follows. First, an image acquisition means installed in the room environment periodically photographs the room's conditions and acquires images of objects. These images are transmitted to a server via wireless communication means.
[0462] The server processes the received images using analysis tools and recognizes and classifies objects in the images using a generative artificial intelligence model. This process identifies the type and characteristics of each object and tags them. It also extracts the object's location information as coordinates within the image.
[0463] Next, the server uses recording devices to record the analyzed object information and its location information in a database. This makes it possible to track the object's latest location and also use it as historical information.
[0464] When a user wants to know the location of an object, they send a question to the server via voice or text through their device (smartphone or voice assistant). The server analyzes this input using natural language processing technology to determine which object to search for.
[0465] The server uses a search mechanism to retrieve the latest location information of the target object from the database and sends the search results back to the terminal. The terminal then communicates the search results to the user via screen display or audio output. This allows the user to quickly learn the object's current location.
[0466] For example, if a user uses their smartphone to voice-input "Where is the remote control?", the server recognizes the keyword "remote control" and searches the database for the relevant record. If the search result is, for example, "It's on the living room table," it will communicate this to the user via voice output. In this way, the present invention aims to streamline the process of finding objects in daily life.
[0467] The following describes the processing flow.
[0468] Step 1:
[0469] The camera acquires images of the room. The camera takes pictures of the room's conditions at regular intervals and transmits the image data to a server via wireless communication.
[0470] Step 2:
[0471] The server passes the received image data to the analysis unit. The analysis unit uses a generative artificial intelligence model to recognize objects in the image, identify the type and characteristics of each object, and classify them.
[0472] Step 3:
[0473] The server stores the analysis results in a database using recording devices. The information stored includes the name of the object, identified features, position coordinates in the image, and the date and time of acquisition.
[0474] Step 4:
[0475] The user uses their device to search for the location of an object. The user enters a question into the device's application or voice assistant, asking about the location of a specific object.
[0476] Step 5:
[0477] The device sends the user's question to the server. The question is sent to the server as text or audio data.
[0478] Step 6:
[0479] The server analyzes the question using natural language processing technology. The purpose of the analysis is to identify the object the user is looking for. Based on the analysis results, the server searches for information in the database.
[0480] Step 7:
[0481] The server generates search results and retrieves the latest location information for the relevant objects. It then constructs a response based on this location information.
[0482] Step 8:
[0483] The server sends a response to the terminal. The response contains location information of the object, and this information is conveyed to the user.
[0484] Step 9:
[0485] The device notifies the user of the search results. The user can then confirm the object's location information via the device's screen display or audio output.
[0486] (Example 1)
[0487] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0488] Conventional inventory management systems struggled to accurately track the location of items and provide that information quickly. Furthermore, they lacked efficient methods for responding to user inquiries.
[0489] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0490] In this invention, the server includes an acquisition means for acquiring images of objects installed in a room environment, a means for analyzing the acquired images to recognize and classify the objects, and a registration means for recording and storing spatial information of the analyzed objects. This enables accurate location tracking of objects and allows for quick and accurate responses to user inquiries.
[0491] "Acquisition means" refers to a device installed within the room environment for acquiring images of objects.
[0492] "Analysis means" refers to the technology and processes used to recognize and classify objects using acquired images.
[0493] A "registration mechanism" is a system for recording and storing the spatial information of an analyzed item.
[0494] A "search tool" is a system equipped with the function of searching for recorded spatial information based on inquiries from users and providing responses.
[0495] "Wireless communication means" refers to wireless technology and equipment used to transmit acquired image data to a server.
[0496] A "generative artificial intelligence model" is a form of AI technology used to recognize objects in images.
[0497] "Spatial information" refers to data about the position of objects obtained through image analysis, indicating their arrangement within a room.
[0498] The present invention is configured as follows in an embodiment for specifically implementing an item management system. The room environment is equipped with devices such as cameras and sensors as acquisition means. These acquisition means periodically acquire images of items in the room and transmit the image data to a server using wireless communication means.
[0499] The server uses analysis software incorporating a generative AI model to analyze the received image data. This analysis allows the server to recognize objects in the image and classify their type and location. The AI model uses machine learning algorithms to tag the objects.
[0500] After analysis, the server uses a registration mechanism to save the spatial information of the extracted items to a database. This saved data allows the server to track the location of the items.
[0501] When a user wants to know the location of a specific item, they send a question to the server via voice or text through their device (e.g., a smartphone or voice assistant). For example, if a user voice-inputs "Where is the remote control?", the device sends this information to the server.
[0502] The server uses natural language processing technology to analyze the user's query and identify keywords for the search. Based on this, the server retrieves the latest spatial information of the relevant item from the database and provides a response to the user using search tools.
[0503] Finally, the terminal provides the user with a response from the server via screen display or audio output. This allows the user to quickly find the location of the desired item. In this way, the system of the present invention streamlines the search for items in daily life and improves convenience.
[0504] Example of a prompt:
[0505] "Identify the items in the room, record their locations, and report them."
[0506] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0507] Step 1:
[0508] An acquisition device installed in the room environment takes images of objects. The input for this step is the current state of the room, and the output is the captured image data. Specifically, the camera automatically takes pictures at regular time intervals.
[0509] Step 2:
[0510] The terminal uses wireless communication to send captured image data to the server. The input in this step is the acquired image data, and the output is the image data transferred to the server. The terminal transfers the image data to the server using a specific protocol, such as Wi-Fi.
[0511] Step 3:
[0512] The server processes the received image using an analysis tool. The input to this step is the image data sent to the server, and the output is the data of the analyzed objects. Specifically, a generative AI model is used to recognize objects in the image and classify their type and location.
[0513] Step 4:
[0514] The server processes the analyzed data and extracts spatial information. The input for this step is the data of the analyzed objects, and the output is the coordinates of the objects and other spatial data. The AI model identifies the location of objects in the image using pixel coordinates and converts them to coordinates in real space.
[0515] Step 5:
[0516] The server saves the analysis results to the database using a registration mechanism. The input for this step is item data including spatial information, and the output is the updated database. The server establishes a database connection and inserts or updates data using SQL or similar methods.
[0517] Step 6:
[0518] The user inquires about the location of an item via a terminal. The input is a voice or text inquiry from the user, and the output is the request. The terminal performs speech recognition and sends the inquiry to the server in text format.
[0519] Step 7:
[0520] The server uses natural language processing to analyze the user's inquiry. The input is the user's request, and the output is keywords for the search. The server uses an NLP engine to convert the voice query into an appropriate search query.
[0521] Step 8:
[0522] The server searches the database to retrieve the latest spatial information of an item. The input is a search query, and the output is the item's location information. SQL queries are used to retrieve the latest relevant information from the database.
[0523] Step 9:
[0524] The server transmits the acquired information to the terminal. The input is the location information of an item, and the output is an information packet containing that information. The server encodes the data via a communication protocol and transmits it to the terminal.
[0525] Step 10:
[0526] The device provides the user with location information for an item. The input in this step is information from the server, and the output is visual or audible feedback to the user. Specifically, the device may display text on the screen, or a voice assistant may read the content aloud.
[0527] (Application Example 1)
[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] Managing and tracking objects within a factory is fraught with human error and inefficiency issues. Especially when dealing with a large number of parts or products, identifying their locations and managing inventory can be time-consuming and labor-intensive, leading to decreased production efficiency. There is a need to solve this problem and achieve efficient and automated object management and location tracking.
[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0531] In this invention, the server includes an acquisition means installed in the environment to acquire visual information of objects; an analysis means for analyzing the acquired visual information to recognize and classify objects; a recording means for recording and storing the location information of the analyzed objects; a search means for searching the recorded location information based on input from a user and providing a response; and a control means for detecting shortages of objects and controlling corresponding operations. This enables automatic object recognition, real-time tracking of location information, efficient inventory management, and replenishment activities.
[0532] "Means of acquisition" refers to devices installed in the environment to acquire visual information.
[0533] "Analysis means" refers to a device or system that has the function of analyzing acquired visual information and recognizing and classifying objects.
[0534] "Recording means" refers to a device or database system for recording and storing location information of an analyzed object.
[0535] "Search means" refers to a device or system that has the function of searching for location information recorded based on input from a user and providing a response.
[0536] A "control means" is a device or system that has the function of detecting a shortage of an object and controlling the corresponding work automatically.
[0537] The system for carrying out this invention mainly consists of a complex group of devices for acquiring, analyzing, recording, retrieving, and controlling visual information of objects. First, acquisition means installed in the environment periodically acquire visual information of objects using cameras and sensors. The acquired data is transmitted to a server via wireless communication.
[0538] The server processes the received data using analysis tools and recognizes and classifies objects using a generative AI model. Specifically, it analyzes visual data using the OpenCV image processing library, and a machine learning model using TensorFlow identifies the type of each object. At this time, recognized objects are tagged, and their location information is extracted as coordinate data.
[0539] The recording mechanism is responsible for saving these analysis results to a database, such as using MySQL to record the latest location information of objects. This information can later be used as historical data.
[0540] The search mechanism analyzes voice or text input transmitted from the user via the terminal and retrieves the location information of the specified object from a database. The NLTK library is used for natural language processing to understand the user's questions and generate appropriate responses. Location information is presented to the user's terminal in real time, either audibly or visually.
[0541] The control system monitors the shortage of objects identified by the analysis system and automatically controls the necessary tasks. For example, when there is a shortage of a particular type or quantity of objects in a factory, a robot automatically replenishes the necessary items from the stockroom. In this case, ROS (Robotics Operating System) is used as the robotics operating system.
[0542] As a concrete example, in response to a user's question, "Where are the bolts?", a specific answer such as "They are on biomaterial shelf 3" is provided. An example of a prompt for the generating AI model is, "Image of factory line with parts and tools. Identify and classify items, then record their positions in the database." In this way, object management is automated through a combination of devices and software.
[0543] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0544] Step 1:
[0545] The device acquires visual information about objects through cameras and sensors installed in the environment. This visual data is generated in real time from the camera and becomes input data transmitted to the server via wireless communication.
[0546] Step 2:
[0547] The server receives the acquired data and analyzes the images using OpenCV. First, preprocessing such as noise reduction and grayscale conversion is performed, and then object recognition and classification are carried out using the generative AI model TensorFlow. The output obtained here is the type of recognized object and its position coordinates within the visual field.
[0548] Step 3:
[0549] The server saves the analysis results to a database via a recording mechanism. The inputs used are object type, location coordinates, and tagging information, and this data is recorded in a MySQL database. The output is a database entry for future searches.
[0550] Step 4:
[0551] The user asks a question from their device using voice or text input. For example, a query such as "Where are the screws?" is sent to the server. The input here is a natural language request from the user.
[0552] Step 5:
[0553] The server uses the NLTK library to analyze user queries. It breaks down the natural language input into intent and constructs a search query. The output consists of object names and related information.
[0554] Step 6:
[0555] The server uses a search mechanism to retrieve the latest object location information from the database. The input is the object name obtained in step 5, and the database query is executed to output the location information.
[0556] Step 7:
[0557] The terminal receives location information sent back from the server and notifies the user visually or audibly. The output is specific location information, such as "It's on shelf 2 in warehouse C." This allows the user to quickly determine the object's location.
[0558] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0559] As an embodiment of the present invention, a system is presented that combines an emotion engine with an object management system to recognize the user's emotions and change its response accordingly. First, an image acquisition means installed in the room environment captures images of the room's conditions at regular intervals. The acquired image data is transmitted to a server via wireless communication means.
[0560] The server uses analysis tools to process the provided images using a generating artificial intelligence model. This model recognizes and classifies objects in the images, identifying their characteristics and locations. The classified information is stored in a database by recording tools, allowing for object referencing within the database.
[0561] When a user wants to know the location of an object, they send a search request to the server using their device via voice or text input. The server analyzes the received input using natural language processing and then searches the database for the corresponding object. A new feature introduced here is the emotion engine, which analyzes the user's input voice data and recognizes the user's emotional state.
[0562] The emotion engine assesses the user's emotional state and adjusts its response accordingly. For example, if it detects that the user is anxious, it might soften the response or emphasize urgency. It also selects a method for presenting object location information based on the user's emotions. One example is whether to provide detailed information calmly or to quickly convey only the key points.
[0563] The server generates search results and a response tailored to the user's mood and sends it to the terminal. The terminal displays the results to the user visually or outputs them audibly. If the user asks "Where is the remote control?" through the terminal, the server searches the database for the remote control's location and notifies the user audibly, "It's on the living room table." If the terminal detects from the user's tone of voice that they are in a hurry, the terminal's notification will be expedited.
[0564] Thus, the present invention aims to further improve the efficiency of object searching by providing a more user-friendly interface that takes into account user emotions in addition to object position recognition.
[0565] The following describes the processing flow.
[0566] Step 1:
[0567] The camera captures images of the room. The camera installed in the room periodically takes pictures and sends the image data to the server.
[0568] Step 2:
[0569] The server analyzes the received image. Using analysis tools, the server analyzes the image based on a generative artificial intelligence model to recognize and classify objects.
[0570] Step 3:
[0571] The server records the object's location information in a database. The characteristics and location information of the recognized object are saved in the database for later retrieval.
[0572] Step 4:
[0573] The user operates the device to request an object search. The user asks for the location of a specific object via voice or text through the device's interface.
[0574] Step 5:
[0575] The terminal sends user input to the server. User questions are sent from the terminal to the server as data.
[0576] Step 6:
[0577] The server analyzes the user's question. The server uses natural language processing to understand the user's intent and searches the database for the relevant object.
[0578] Step 7:
[0579] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's voice data to determine their emotional state.
[0580] Step 8:
[0581] The server adjusts its response based on the user's emotional state. For example, if the user is irritated, the response will be changed to a more calming tone.
[0582] Step 9:
[0583] The server sends the search results to the terminal. The adjusted response, along with the object's location information, is transmitted to the terminal.
[0584] Step 10:
[0585] The device notifies the user of the results. The device displays the location information on the screen or reports it to the user via voice. The user receives the notification from the device and confirms the object's location.
[0586] (Example 2)
[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0588] Conventional object management systems have the problem of failing to adequately reduce the user's psychological burden because they do not take into account the user's emotional state when locating objects in an environment. Furthermore, there is a lack of methods to provide a more user-friendly interface in addition to simply providing location information for objects.
[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0590] In this invention, the server includes data acquisition means for acquiring information about objects installed in the environment, analysis means for analyzing the acquired information and identifying and classifying objects, recording means for storing location information of the analyzed objects, search means for finding the recorded location information based on user information and providing a response, and emotion analysis means for analyzing the user's emotional state and adjusting the response. This makes it possible to provide the location of objects in the environment appropriately and with consideration for the user's emotions.
[0591] "Environment" refers to the place or situation in which an object exists within a specific space or set of conditions.
[0592] "Data acquisition means" refers to a function or device for collecting and recording target information or data.
[0593] "Analysis means" refers to a function or device for understanding acquired data and identifying, classifying, and evaluating necessary information.
[0594] "Recording means" refers to a function or device for organizing information and data and storing them so that they can be used later.
[0595] "Search method" refers to a function or device that finds the necessary content from recorded information and provides it to the user.
[0596] "Emotional analysis means" refers to a function or device that determines the emotional state of a user from their input and actions and adjusts the system's response accordingly.
[0597] A "generative model" refers to an algorithm that uses machine learning or artificial intelligence to learn data patterns and then performs inferences on new data.
[0598] "Natural language processing" refers to the technology that enables computers to understand and process human language appropriately.
[0599] This invention provides a novel user interface that combines an object management system with emotion analysis functionality. The system mainly consists of a server, terminals, and an image acquisition device.
[0600] First, the terminal uses a data acquisition device installed in the environment, such as a room, to capture images of objects and the environment at specific time intervals. This image acquisition device consists of hardware such as a camera, and transmits the acquired image data to a server wirelessly.
[0601] The server analyzes the received image data using a generative artificial intelligence model (generative AI model). This model has the function of identifying and classifying objects contained in the image. The analysis results are stored in a database through recording means and can be searched as needed.
[0602] When a user needs information about an object, they can send a natural language prompt to the server via their device, for example, "Tell me where the remote control is." The server uses natural language processing technology to analyze this prompt and retrieve information about the corresponding object from its database.
[0603] Furthermore, this system incorporates emotion analysis capabilities. The server can analyze the user's voice input and recognize their emotional state. For example, if it determines that the user is anxious, it can adjust the format and pace of the information provided.
[0604] In this way, users not only receive information but also convenient responses tailored to their psychological state. This system achieves a more user-friendly experience by incorporating emotional states into object position recognition.
[0605] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0606] Step 1:
[0607] The terminal periodically captures images of the environment using image acquisition devices such as cameras installed in the room. The input is image data captured by the camera. The output is image data transmitted to a server via wireless communication. Specifically, this process acquires visual information about objects within the environment.
[0608] Step 2:
[0609] The server passes the received image data to a generative AI model for analysis. The input here is image data sent from the terminal. The generative AI model uses this data to identify and classify objects within the image. The output is data that identifies the features and location information of the objects. Specifically, the model extracts object labels and location coordinates.
[0610] Step 3:
[0611] The server stores the analyzed object features and location information in a database. The input is the object feature information created by the generative AI model. The output is information recorded in a format that can be searched later. Specifically, the process involves saving the information and adding it to the database.
[0612] Step 4:
[0613] The user sends prompts to the server via voice or text using a terminal to find the location of an object. The input is a natural language query from the user. The output is a search request to the server. For example, the user might give instructions such as "Tell me where the remote control is."
[0614] Step 5:
[0615] The server parses the received prompt message using natural language processing and searches the database for related objects. The input is the prompt message sent by the user. This prompt message is parsed to retrieve information about the corresponding object in the database. The output is the object's location information. Specifically, the parsing engine interprets the intent of the query and retrieves the relevant data.
[0616] Step 6:
[0617] The server analyzes the user's voice data using emotion analysis tools to recognize their emotional state. The input is the user's voice data. Through analysis, it determines the user's emotional state. The output is response adjustment information based on the emotional state. For example, it might detect anxiety from the user's voice.
[0618] Step 7:
[0619] The server generates a response based on the location and emotional state of the acquired object and sends it to the terminal. The input is adjustment information based on the object's location and emotional state. The output is the adjusted response provided to the user. Specifically, the information is presented in audio or visual form, and the tone and amount of information are changed as needed.
[0620] Step 8:
[0621] The terminal provides the user with the response received from the server. Input is the pre-arranged response sent from the server. Output is information presented to the user visually or audibly. Specifically, it provides notifications in a timely and appropriate manner depending on the user's status.
[0622] (Application Example 2)
[0623] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0624] Conventional object management systems, when providing location information for objects, often respond mechanically without considering the user's emotional state, making it difficult to provide a user-satisfying interface. Furthermore, while real-time, rapid, and accurate information is required for product searches within stores, there is a lack of consideration for the customer's emotions. This can lead to increased customer stress and reduced purchasing intent.
[0625] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0626] In this invention, the server is installed in a room environment and includes an image acquisition means for acquiring images of objects, an analysis means for analyzing the images acquired by the acquisition means and recognizing and classifying objects, a recording means for recording and storing the location information of the analyzed objects, a search means for searching the recorded location information based on input from the user and providing a response, and an emotion analysis means for analyzing the user's emotional state and optimizing the response. This enables the provision of appropriate information according to the user's emotions, realizes a more user-friendly interface, and allows for efficient product searching within a store.
[0627] "Image acquisition means" refers to devices or functions installed in a room environment for acquiring images of objects.
[0628] "Analysis means" refers to devices or functions that analyze acquired image data and recognize and classify objects.
[0629] "Recording means" refers to a device or function for recording and storing the positional information of an analyzed object.
[0630] A "search means" refers to a device or function that searches for location information recorded based on input from a user and provides a response.
[0631] "Emotional analysis tools" refer to devices or functions that analyze a user's emotional state and optimize their response based on that evaluation.
[0632] A "generative artificial intelligence model" is a model based on artificial intelligence technology used for recognizing and analyzing image data, and it improves accuracy using machine learning techniques.
[0633] "Natural language processing" is a technology that analyzes the user's voice input, understands its meaning, and responds accordingly.
[0634] This system is a technology for managing objects within rooms and stores and providing information tailored to the user's emotions. The system mainly consists of a server, user terminals (smart glasses or smartphones), and image acquisition devices installed in the room environment.
[0635] The server utilizes a generative AI model to recognize and classify objects from received image data. This analysis employs computer vision technologies such as OpenCV and TensorFlow. The location information of the analyzed objects is stored in a SQL database. When a user wants to know the location of an object, they query the server using voice or text via their device. Natural language processing, such as the Google Cloud Natural Language API, is used to analyze the user's query.
[0636] Furthermore, the server uses emotion analysis tools to determine the user's emotional state. It analyzes whether the user is in a hurry or relaxed based on their tone of voice and text expression. Based on this information, it provides responses not simply as data, but in a way that is appropriate to the situation.
[0637] As a concrete example, consider a scenario where a user asks, "Where is the shaved ice machine?" The server identifies the location of the product from the analyzed database information. If the emotion analysis determines that the user is anxious, it immediately responds with a voice message; if the user is calm, it displays detailed map information on the smart glasses. This allows the user to easily find the product within the store.
[0638] Possible prompts to input into a generative AI model include the following:
[0639] "I'm looking for this product. Please tell me its location."
[0640] "I'm in a hurry. Please show me the way quickly."
[0641] This system will allow users to have a more efficient and satisfying experience.
[0642] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0643] Step 1:
[0644] The terminal sends instructions to the image acquisition device to acquire images of the room environment or store interior. The image acquisition device uses a camera to take still images at predetermined time intervals. The captured image data is transmitted to the server via wireless communication. The input is the capture instruction, and the output is the acquired image data.
[0645] Step 2:
[0646] The server feeds the received image data into an AI image analysis model (for example, using TensorFlow). Here, objects in the image are recognized and classified. The generative AI model extracts the features of the objects and determines their type based on these features. The input is image data, and the output is the type of object and its location information.
[0647] Step 3:
[0648] The server saves the object location information obtained as a result of the analysis to an SQL database. Before saving, the data integrity is checked, and data processing is performed to remove duplicates and other errors. The input is the analyzed object information, and the output is the updated database.
[0649] Step 4:
[0650] The user asks about the location of an object via voice or text through the terminal. The voice data is initially processed within the terminal and then sent to the server in text format. The input is the user's voice or text, and the output is the question in text format.
[0651] Step 5:
[0652] The server analyzes the received question using a natural language processing (NLP) engine (such as the Google Cloud Natural Language API) and retrieves the necessary object information from a SQL database. The input is a text question from the user, and the output is the location information of the searched object.
[0653] Step 6:
[0654] The server uses sentiment analysis to identify the user's emotional state based on the content of the user's input voice and text. It analyzes the user's emotional state, such as whether they are anxious or calm, based on the tone of their voice and the content of their text. The input is the user's voice / text data, and the output is their emotional state.
[0655] Step 7:
[0656] The server generates a response optimized according to the user's emotional state. For example, if the user is in a hurry, it will quickly provide the object's location via voice. If the user is calm, it will provide a detailed explanation or map information. The input is the emotional state and the object's location, and the output is the response message.
[0657] Step 8:
[0658] The server generates a response, which is then sent to the terminal and presented to the user. The terminal communicates the result through speech synthesis or display. The input is the response message, and the output is the provision of visual or auditory information to the user.
[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0662] [Fourth Embodiment]
[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0676] As an embodiment for carrying out the present invention, the object management system is specifically configured as follows. First, an image acquisition means installed in the room environment periodically photographs the room's conditions and acquires images of objects. These images are transmitted to a server via wireless communication means.
[0677] The server processes the received images using analysis tools and recognizes and classifies objects in the images using a generative artificial intelligence model. This process identifies the type and characteristics of each object and tags them. It also extracts the object's location information as coordinates within the image.
[0678] Next, the server uses recording devices to record the analyzed object information and its location information in a database. This makes it possible to track the object's latest location and also use it as historical information.
[0679] When a user wants to know the location of an object, they send a question to the server via voice or text through their device (smartphone or voice assistant). The server analyzes this input using natural language processing technology to determine which object to search for.
[0680] The server uses a search mechanism to retrieve the latest location information of the target object from the database and sends the search results back to the terminal. The terminal then communicates the search results to the user via screen display or audio output. This allows the user to quickly learn the object's current location.
[0681] For example, if a user uses their smartphone to voice-input "Where is the remote control?", the server recognizes the keyword "remote control" and searches the database for the relevant record. If the search result is, for example, "It's on the living room table," it will communicate this to the user via voice output. In this way, the present invention aims to streamline the process of finding objects in daily life.
[0682] The following describes the processing flow.
[0683] Step 1:
[0684] The camera acquires images of the room. The camera takes pictures of the room's conditions at regular intervals and transmits the image data to a server via wireless communication.
[0685] Step 2:
[0686] The server passes the received image data to the analysis unit. The analysis unit uses a generative artificial intelligence model to recognize objects in the image, identify the type and characteristics of each object, and classify them.
[0687] Step 3:
[0688] The server stores the analysis results in a database using recording devices. The information stored includes the name of the object, identified features, position coordinates in the image, and the date and time of acquisition.
[0689] Step 4:
[0690] The user uses their device to search for the location of an object. The user enters a question into the device's application or voice assistant, asking about the location of a specific object.
[0691] Step 5:
[0692] The device sends the user's question to the server. The question is sent to the server as text or audio data.
[0693] Step 6:
[0694] The server analyzes the question using natural language processing technology. The purpose of the analysis is to identify the object the user is looking for. Based on the analysis results, the server searches for information in the database.
[0695] Step 7:
[0696] The server generates search results and retrieves the latest location information for the relevant objects. It then constructs a response based on this location information.
[0697] Step 8:
[0698] The server sends a response to the terminal. The response contains location information of the object, and this information is conveyed to the user.
[0699] Step 9:
[0700] The device notifies the user of the search results. The user can then confirm the object's location information via the device's screen display or audio output.
[0701] (Example 1)
[0702] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0703] Conventional inventory management systems struggled to accurately track the location of items and provide that information quickly. Furthermore, they lacked efficient methods for responding to user inquiries.
[0704] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0705] In this invention, the server includes an acquisition means for acquiring images of objects installed in a room environment, a means for analyzing the acquired images to recognize and classify the objects, and a registration means for recording and storing spatial information of the analyzed objects. This enables accurate location tracking of objects and allows for quick and accurate responses to user inquiries.
[0706] "Acquisition means" refers to a device installed within the room environment for acquiring images of objects.
[0707] "Analysis means" refers to the technology and processes used to recognize and classify objects using acquired images.
[0708] A "registration mechanism" is a system for recording and storing the spatial information of an analyzed item.
[0709] A "search tool" is a system equipped with the function of searching for recorded spatial information based on inquiries from users and providing responses.
[0710] "Wireless communication means" refers to wireless technology and equipment used to transmit acquired image data to a server.
[0711] A "generative artificial intelligence model" is a form of AI technology used to recognize objects in images.
[0712] "Spatial information" refers to data about the position of objects obtained through image analysis, indicating their arrangement within a room.
[0713] The present invention is configured as follows in an embodiment for specifically implementing an item management system. The room environment is equipped with devices such as cameras and sensors as acquisition means. These acquisition means periodically acquire images of items in the room and transmit the image data to a server using wireless communication means.
[0714] The server uses analysis software incorporating a generative AI model to analyze the received image data. This analysis allows the server to recognize objects in the image and classify their type and location. The AI model uses machine learning algorithms to tag the objects.
[0715] After analysis, the server uses a registration mechanism to save the spatial information of the extracted items to a database. This saved data allows the server to track the location of the items.
[0716] When a user wants to know the location of a specific item, they send a question to the server via voice or text through their device (e.g., a smartphone or voice assistant). For example, if a user voice-inputs "Where is the remote control?", the device sends this information to the server.
[0717] The server uses natural language processing technology to analyze the user's query and identify keywords for the search. Based on this, the server retrieves the latest spatial information of the relevant item from the database and provides a response to the user using search tools.
[0718] Finally, the terminal provides the user with a response from the server via screen display or audio output. This allows the user to quickly find the location of the desired item. In this way, the system of the present invention streamlines the search for items in daily life and improves convenience.
[0719] Example of a prompt:
[0720] "Identify the items in the room, record their locations, and report them."
[0721] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0722] Step 1:
[0723] An acquisition device installed in the room environment takes images of objects. The input for this step is the current state of the room, and the output is the captured image data. Specifically, the camera automatically takes pictures at regular time intervals.
[0724] Step 2:
[0725] The terminal uses wireless communication to send captured image data to the server. The input in this step is the acquired image data, and the output is the image data transferred to the server. The terminal transfers the image data to the server using a specific protocol, such as Wi-Fi.
[0726] Step 3:
[0727] The server processes the received image using an analysis tool. The input to this step is the image data sent to the server, and the output is the data of the analyzed objects. Specifically, a generative AI model is used to recognize objects in the image and classify their type and location.
[0728] Step 4:
[0729] The server processes the analyzed data and extracts spatial information. The input for this step is the data of the analyzed objects, and the output is the coordinates of the objects and other spatial data. The AI model identifies the location of objects in the image using pixel coordinates and converts them to coordinates in real space.
[0730] Step 5:
[0731] The server saves the analysis results to the database using a registration mechanism. The input for this step is item data including spatial information, and the output is the updated database. The server establishes a database connection and inserts or updates data using SQL or similar methods.
[0732] Step 6:
[0733] The user inquires about the location of an item via a terminal. The input is a voice or text inquiry from the user, and the output is the request. The terminal performs speech recognition and sends the inquiry to the server in text format.
[0734] Step 7:
[0735] The server uses natural language processing to analyze the user's inquiry. The input is the user's request, and the output is keywords for the search. The server uses an NLP engine to convert the voice query into an appropriate search query.
[0736] Step 8:
[0737] The server searches the database to retrieve the latest spatial information of an item. The input is a search query, and the output is the item's location information. SQL queries are used to retrieve the latest relevant information from the database.
[0738] Step 9:
[0739] The server transmits the acquired information to the terminal. The input is the location information of an item, and the output is an information packet containing that information. The server encodes the data via a communication protocol and transmits it to the terminal.
[0740] Step 10:
[0741] The device provides the user with location information for an item. The input in this step is information from the server, and the output is visual or audible feedback to the user. Specifically, the device may display text on the screen, or a voice assistant may read the content aloud.
[0742] (Application Example 1)
[0743] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] Managing and tracking objects within a factory is fraught with human error and inefficiency issues. Especially when dealing with a large number of parts or products, identifying their locations and managing inventory can be time-consuming and labor-intensive, leading to decreased production efficiency. There is a need to solve this problem and achieve efficient and automated object management and location tracking.
[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0746] In this invention, the server includes an acquisition means installed in the environment to acquire visual information of objects; an analysis means for analyzing the acquired visual information to recognize and classify objects; a recording means for recording and storing the location information of the analyzed objects; a search means for searching the recorded location information based on input from a user and providing a response; and a control means for detecting shortages of objects and controlling corresponding operations. This enables automatic object recognition, real-time tracking of location information, efficient inventory management, and replenishment activities.
[0747] "Means of acquisition" refers to devices installed in the environment to acquire visual information.
[0748] "Analysis means" refers to a device or system that has the function of analyzing acquired visual information and recognizing and classifying objects.
[0749] "Recording means" refers to a device or database system for recording and storing location information of an analyzed object.
[0750] "Search means" refers to a device or system that has the function of searching for location information recorded based on input from a user and providing a response.
[0751] A "control means" is a device or system that has the function of detecting a shortage of an object and controlling the corresponding work automatically.
[0752] The system for carrying out this invention mainly consists of a complex group of devices for acquiring, analyzing, recording, retrieving, and controlling visual information of objects. First, acquisition means installed in the environment periodically acquire visual information of objects using cameras and sensors. The acquired data is transmitted to a server via wireless communication.
[0753] The server processes the received data using analysis tools and recognizes and classifies objects using a generative AI model. Specifically, it analyzes visual data using the OpenCV image processing library, and a machine learning model using TensorFlow identifies the type of each object. At this time, recognized objects are tagged, and their location information is extracted as coordinate data.
[0754] The recording mechanism is responsible for saving these analysis results to a database, such as using MySQL to record the latest location information of objects. This information can later be used as historical data.
[0755] The search mechanism analyzes voice or text input transmitted from the user via the terminal and retrieves the location information of the specified object from a database. The NLTK library is used for natural language processing to understand the user's questions and generate appropriate responses. Location information is presented to the user's terminal in real time, either audibly or visually.
[0756] The control system monitors the shortage of objects identified by the analysis system and automatically controls the necessary tasks. For example, when there is a shortage of a particular type or quantity of objects in a factory, a robot automatically replenishes the necessary items from the stockroom. In this case, ROS (Robotics Operating System) is used as the robotics operating system.
[0757] As a concrete example, in response to a user's question, "Where are the bolts?", a specific answer such as "They are on biomaterial shelf 3" is provided. An example of a prompt for the generating AI model is, "Image of factory line with parts and tools. Identify and classify items, then record their positions in the database." In this way, object management is automated through a combination of devices and software.
[0758] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0759] Step 1:
[0760] The device acquires visual information about objects through cameras and sensors installed in the environment. This visual data is generated in real time from the camera and becomes input data transmitted to the server via wireless communication.
[0761] Step 2:
[0762] The server receives the acquired data and analyzes the images using OpenCV. First, preprocessing such as noise reduction and grayscale conversion is performed, and then object recognition and classification are carried out using the generative AI model TensorFlow. The output obtained here is the type of recognized object and its position coordinates within the visual field.
[0763] Step 3:
[0764] The server saves the analysis results to a database via a recording mechanism. The inputs used are object type, location coordinates, and tagging information, and this data is recorded in a MySQL database. The output is a database entry for future searches.
[0765] Step 4:
[0766] The user asks a question from their device using voice or text input. For example, a query such as "Where are the screws?" is sent to the server. The input here is a natural language request from the user.
[0767] Step 5:
[0768] The server uses the NLTK library to analyze user queries. It breaks down the natural language input into intent and constructs a search query. The output consists of object names and related information.
[0769] Step 6:
[0770] The server uses a search mechanism to retrieve the latest object location information from the database. The input is the object name obtained in step 5, and the database query is executed to output the location information.
[0771] Step 7:
[0772] The terminal receives location information sent back from the server and notifies the user visually or audibly. The output is specific location information, such as "It's on shelf 2 in warehouse C." This allows the user to quickly determine the object's location.
[0773] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0774] As an embodiment of the present invention, a system is presented that combines an emotion engine with an object management system to recognize the user's emotions and change its response accordingly. First, an image acquisition means installed in the room environment captures images of the room's conditions at regular intervals. The acquired image data is transmitted to a server via wireless communication means.
[0775] The server uses analysis tools to process the provided images using a generating artificial intelligence model. This model recognizes and classifies objects in the images, identifying their characteristics and locations. The classified information is stored in a database by recording tools, allowing for object referencing within the database.
[0776] When a user wants to know the location of an object, they send a search request to the server using their device via voice or text input. The server analyzes the received input using natural language processing and then searches the database for the corresponding object. A new feature introduced here is the emotion engine, which analyzes the user's input voice data and recognizes the user's emotional state.
[0777] The emotion engine assesses the user's emotional state and adjusts its response accordingly. For example, if it detects that the user is anxious, it might soften the response or emphasize urgency. It also selects a method for presenting object location information based on the user's emotions. One example is whether to provide detailed information calmly or to quickly convey only the key points.
[0778] The server generates search results and a response tailored to the user's mood and sends it to the terminal. The terminal displays the results to the user visually or outputs them audibly. If the user asks "Where is the remote control?" through the terminal, the server searches the database for the remote control's location and notifies the user audibly, "It's on the living room table." If the terminal detects from the user's tone of voice that they are in a hurry, the terminal's notification will be expedited.
[0779] Thus, the present invention aims to further improve the efficiency of object searching by providing a more user-friendly interface that takes into account user emotions in addition to object position recognition.
[0780] The following describes the processing flow.
[0781] Step 1:
[0782] The camera captures images of the room. The camera installed in the room periodically takes pictures and sends the image data to the server.
[0783] Step 2:
[0784] The server analyzes the received image. Using analysis tools, the server analyzes the image based on a generative artificial intelligence model to recognize and classify objects.
[0785] Step 3:
[0786] The server records the object's location information in a database. The characteristics and location information of the recognized object are saved in the database for later retrieval.
[0787] Step 4:
[0788] The user operates the device to request an object search. The user asks for the location of a specific object via voice or text through the device's interface.
[0789] Step 5:
[0790] The terminal sends user input to the server. User questions are sent from the terminal to the server as data.
[0791] Step 6:
[0792] The server analyzes the user's question. The server uses natural language processing to understand the user's intent and searches the database for the relevant object.
[0793] Step 7:
[0794] The server uses an emotion engine to recognize the user's emotions. It analyzes the user's voice data to determine their emotional state.
[0795] Step 8:
[0796] The server adjusts its response based on the user's emotional state. For example, if the user is irritated, the response will be changed to a more calming tone.
[0797] Step 9:
[0798] The server sends the search results to the terminal. The adjusted response, along with the object's location information, is transmitted to the terminal.
[0799] Step 10:
[0800] The device notifies the user of the results. The device displays the location information on the screen or reports it to the user via voice. The user receives the notification from the device and confirms the object's location.
[0801] (Example 2)
[0802] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0803] Conventional object management systems have the problem of failing to adequately reduce the user's psychological burden because they do not take into account the user's emotional state when locating objects in an environment. Furthermore, there is a lack of methods to provide a more user-friendly interface in addition to simply providing location information for objects.
[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0805] In this invention, the server includes data acquisition means for acquiring information about objects installed in the environment, analysis means for analyzing the acquired information and identifying and classifying objects, recording means for storing location information of the analyzed objects, search means for finding the recorded location information based on user information and providing a response, and emotion analysis means for analyzing the user's emotional state and adjusting the response. This makes it possible to provide the location of objects in the environment appropriately and with consideration for the user's emotions.
[0806] "Environment" refers to the place or situation in which an object exists within a specific space or set of conditions.
[0807] "Data acquisition means" refers to a function or device for collecting and recording target information or data.
[0808] "Analysis means" refers to a function or device for understanding acquired data and identifying, classifying, and evaluating necessary information.
[0809] "Recording means" refers to a function or device for organizing information and data and storing them so that they can be used later.
[0810] "Search method" refers to a function or device that finds the necessary content from recorded information and provides it to the user.
[0811] "Emotional analysis means" refers to a function or device that determines the emotional state of a user from their input and actions and adjusts the system's response accordingly.
[0812] A "generative model" refers to an algorithm that uses machine learning or artificial intelligence to learn data patterns and then performs inferences on new data.
[0813] "Natural language processing" refers to the technology that enables computers to understand and process human language appropriately.
[0814] This invention provides a novel user interface that combines an object management system with emotion analysis functionality. The system mainly consists of a server, terminals, and an image acquisition device.
[0815] First, the terminal uses a data acquisition device installed in the environment, such as a room, to capture images of objects and the environment at specific time intervals. This image acquisition device consists of hardware such as a camera, and transmits the acquired image data to a server wirelessly.
[0816] The server analyzes the received image data using a generative artificial intelligence model (generative AI model). This model has the function of identifying and classifying objects contained in the image. The analysis results are stored in a database through recording means and can be searched as needed.
[0817] When a user needs information about an object, they can send a natural language prompt to the server via their device, for example, "Tell me where the remote control is." The server uses natural language processing technology to analyze this prompt and retrieve information about the corresponding object from its database.
[0818] Furthermore, this system incorporates emotion analysis capabilities. The server can analyze the user's voice input and recognize their emotional state. For example, if it determines that the user is anxious, it can adjust the format and pace of the information provided.
[0819] In this way, users not only receive information but also convenient responses tailored to their psychological state. This system achieves a more user-friendly experience by incorporating emotional states into object position recognition.
[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0821] Step 1:
[0822] The terminal periodically captures images of the environment using image acquisition devices such as cameras installed in the room. The input is image data captured by the camera. The output is image data transmitted to a server via wireless communication. Specifically, this process acquires visual information about objects within the environment.
[0823] Step 2:
[0824] The server passes the received image data to a generative AI model for analysis. The input here is image data sent from the terminal. The generative AI model uses this data to identify and classify objects within the image. The output is data that identifies the features and location information of the objects. Specifically, the model extracts object labels and location coordinates.
[0825] Step 3:
[0826] The server stores the analyzed object features and location information in a database. The input is the object feature information created by the generative AI model. The output is information recorded in a format that can be searched later. Specifically, the process involves saving the information and adding it to the database.
[0827] Step 4:
[0828] The user sends prompts to the server via voice or text using a terminal to find the location of an object. The input is a natural language query from the user. The output is a search request to the server. For example, the user might give instructions such as "Tell me where the remote control is."
[0829] Step 5:
[0830] The server parses the received prompt message using natural language processing and searches the database for related objects. The input is the prompt message sent by the user. This prompt message is parsed to retrieve information about the corresponding object in the database. The output is the object's location information. Specifically, the parsing engine interprets the intent of the query and retrieves the relevant data.
[0831] Step 6:
[0832] The server analyzes the user's voice data using emotion analysis tools to recognize their emotional state. The input is the user's voice data. Through analysis, it determines the user's emotional state. The output is response adjustment information based on the emotional state. For example, it might detect anxiety from the user's voice.
[0833] Step 7:
[0834] The server generates a response based on the location and emotional state of the acquired object and sends it to the terminal. The input is adjustment information based on the object's location and emotional state. The output is the adjusted response provided to the user. Specifically, the information is presented in audio or visual form, and the tone and amount of information are changed as needed.
[0835] Step 8:
[0836] The terminal provides the user with the response received from the server. Input is the pre-arranged response sent from the server. Output is information presented to the user visually or audibly. Specifically, it provides notifications in a timely and appropriate manner depending on the user's status.
[0837] (Application Example 2)
[0838] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0839] Conventional object management systems, when providing location information for objects, often respond mechanically without considering the user's emotional state, making it difficult to provide a user-satisfying interface. Furthermore, while real-time, rapid, and accurate information is required for product searches within stores, there is a lack of consideration for the customer's emotions. This can lead to increased customer stress and reduced purchasing intent.
[0840] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0841] In this invention, the server is installed in a room environment and includes an image acquisition means for acquiring images of objects, an analysis means for analyzing the images acquired by the acquisition means and recognizing and classifying objects, a recording means for recording and storing the location information of the analyzed objects, a search means for searching the recorded location information based on input from the user and providing a response, and an emotion analysis means for analyzing the user's emotional state and optimizing the response. This enables the provision of appropriate information according to the user's emotions, realizes a more user-friendly interface, and allows for efficient product searching within a store.
[0842] "Image acquisition means" refers to devices or functions installed in a room environment for acquiring images of objects.
[0843] "Analysis means" refers to devices or functions that analyze acquired image data and recognize and classify objects.
[0844] "Recording means" refers to a device or function for recording and storing the positional information of an analyzed object.
[0845] A "search means" refers to a device or function that searches for location information recorded based on input from a user and provides a response.
[0846] "Emotional analysis tools" refer to devices or functions that analyze a user's emotional state and optimize their response based on that evaluation.
[0847] A "generative artificial intelligence model" is a model based on artificial intelligence technology used for recognizing and analyzing image data, and it improves accuracy using machine learning techniques.
[0848] "Natural language processing" is a technology that analyzes the user's voice input, understands its meaning, and responds accordingly.
[0849] This system is a technology for managing objects within rooms and stores and providing information tailored to the user's emotions. The system mainly consists of a server, user terminals (smart glasses or smartphones), and image acquisition devices installed in the room environment.
[0850] The server utilizes a generative AI model to recognize and classify objects from received image data. This analysis employs computer vision technologies such as OpenCV and TensorFlow. The location information of the analyzed objects is stored in a SQL database. When a user wants to know the location of an object, they query the server using voice or text via their device. Natural language processing, such as the Google Cloud Natural Language API, is used to analyze the user's query.
[0851] Furthermore, the server uses emotion analysis tools to determine the user's emotional state. It analyzes whether the user is in a hurry or relaxed based on their tone of voice and text expression. Based on this information, it provides responses not simply as data, but in a way that is appropriate to the situation.
[0852] As a concrete example, consider a scenario where a user asks, "Where is the shaved ice machine?" The server identifies the location of the product from the analyzed database information. If the emotion analysis determines that the user is anxious, it immediately responds with a voice message; if the user is calm, it displays detailed map information on the smart glasses. This allows the user to easily find the product within the store.
[0853] Possible prompts to input into a generative AI model include the following:
[0854] "I'm looking for this product. Please tell me its location."
[0855] "I'm in a hurry. Please show me the way quickly."
[0856] This system will allow users to have a more efficient and satisfying experience.
[0857] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0858] Step 1:
[0859] The terminal sends instructions to the image acquisition device to acquire images of the room environment or store interior. The image acquisition device uses a camera to take still images at predetermined time intervals. The captured image data is transmitted to the server via wireless communication. The input is the capture instruction, and the output is the acquired image data.
[0860] Step 2:
[0861] The server feeds the received image data into an AI image analysis model (for example, using TensorFlow). Here, objects in the image are recognized and classified. The generative AI model extracts the features of the objects and determines their type based on these features. The input is image data, and the output is the type of object and its location information.
[0862] Step 3:
[0863] The server saves the object location information obtained as a result of the analysis to an SQL database. Before saving, the data integrity is checked, and data processing is performed to remove duplicates and other errors. The input is the analyzed object information, and the output is the updated database.
[0864] Step 4:
[0865] The user asks about the location of an object via voice or text through the terminal. The voice data is initially processed within the terminal and then sent to the server in text format. The input is the user's voice or text, and the output is the question in text format.
[0866] Step 5:
[0867] The server analyzes the received question using a natural language processing (NLP) engine (such as the Google Cloud Natural Language API) and retrieves the necessary object information from a SQL database. The input is a text question from the user, and the output is the location information of the searched object.
[0868] Step 6:
[0869] The server uses sentiment analysis to identify the user's emotional state based on the content of the user's input voice and text. It analyzes the user's emotional state, such as whether they are anxious or calm, based on the tone of their voice and the content of their text. The input is the user's voice / text data, and the output is their emotional state.
[0870] Step 7:
[0871] The server generates an optimized response based on the user's emotional state. For example, if the user is in a hurry, it will quickly provide the object's location via voice. If the user is calm, it will provide a detailed explanation or map information. The input is the emotional state and the object's location, and the output is the response message.
[0872] Step 8:
[0873] The server generates a response, which is then sent to the terminal and presented to the user. The terminal communicates the result through speech synthesis or display. The input is the response message, and the output is the provision of visual or auditory information to the user.
[0874] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0875] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0876] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0877] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0878] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0879] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0880] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0881] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0882] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0883] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0884] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0885] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0886] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0887] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0888] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0889] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0890] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0891] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0892] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0893] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0894] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0895] The following is further disclosed regarding the embodiments described above.
[0896] (Claim 1)
[0897] An image acquisition means installed in a room environment to acquire images of objects,
[0898] An analysis means for analyzing images acquired by the acquisition means and recognizing and classifying objects,
[0899] A recording means for recording and storing the positional information of the analyzed object,
[0900] A search means that searches for the recorded location information based on input from the user and provides a response,
[0901] An object management system that includes this.
[0902] (Claim 2)
[0903] The object management system according to claim 1, characterized in that the analysis means has a function to recognize objects using a generative artificial intelligence model.
[0904] (Claim 3)
[0905] The object management system according to claim 1, characterized in that the search means performs object searches by analyzing the user's voice input using natural language processing.
[0906] "Example 1"
[0907] (Claim 1)
[0908] A means of acquiring images of objects, installed in the room environment,
[0909] A means for analyzing images acquired by the aforementioned acquisition means and recognizing and classifying items,
[0910] A registration means for recording and storing the spatial information of the analyzed article,
[0911] A search means that searches the recorded spatial information based on an inquiry from a user and provides a response,
[0912] Means including a function for transmitting the acquired image via wireless communication means,
[0913] A system that includes this.
[0914] (Claim 2)
[0915] The system according to claim 1, characterized in that the analysis means has a function to recognize an object using a generative artificial intelligence model and tags the object in the image.
[0916] (Claim 3)
[0917] The system according to claim 1, characterized in that the search means analyzes the user's voice input using natural language processing to search for items and quickly provides the latest spatial information.
[0918] "Application Example 1"
[0919] (Claim 1)
[0920] An acquisition means installed in the environment to acquire visual information of an object,
[0921] An analysis means for analyzing the visual information acquired by the acquisition means and recognizing and classifying objects,
[0922] A recording means for recording and storing the positional information of the analyzed object,
[0923] A search means that searches for the recorded location information based on input from the user and provides a response,
[0924] A control means for detecting the shortage of an object and controlling the corresponding operation,
[0925] A system that includes this.
[0926] (Claim 2)
[0927] The system according to claim 1, characterized in that the analysis means has the function of recognizing objects using a generative artificial intelligence model.
[0928] (Claim 3)
[0929] The system according to claim 1, characterized in that the search means performs object searches by analyzing the user's voice input using natural language processing.
[0930] "Example 2 of combining an emotion engine"
[0931] (Claim 1)
[0932] A data acquisition means installed in the environment to acquire information about objects,
[0933] An analysis means for analyzing the information obtained by the acquisition means and identifying and classifying objects,
[0934] A recording means for storing the position information of the analyzed object,
[0935] A search means that finds the recorded location information based on information from the user and provides a response,
[0936] An emotion analysis tool that analyzes the user's emotional state and adjusts their response,
[0937] A system that includes this.
[0938] (Claim 2)
[0939] The system according to claim 1, characterized in that the analysis means has the function of identifying objects using a generative model.
[0940] (Claim 3)
[0941] The system according to claim 1, characterized in that the search means analyzes the user's voice information using natural language processing to locate objects.
[0942] "Application example 2 when combining with an emotional engine"
[0943] (Claim 1)
[0944] An image acquisition means installed in a room environment to acquire images of objects,
[0945] An analysis means for analyzing images acquired by the acquisition means and recognizing and classifying objects,
[0946] A recording means for recording and storing the positional information of the analyzed object,
[0947] A search means that searches for the recorded location information based on input from the user and provides a response,
[0948] An emotion analysis means that analyzes the user's emotional state and optimizes the response,
[0949] A system that includes this.
[0950] (Claim 2)
[0951] The system according to claim 1, characterized in that the analysis means has the function of recognizing objects using a generative artificial intelligence model.
[0952] (Claim 3)
[0953] The system according to claim 1, characterized in that the search means performs object searches by analyzing the user's voice input using natural language processing. [Explanation of symbols]
[0954] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An image acquisition means installed in a room environment to acquire images of objects, An analysis means for analyzing the image acquired by the acquisition means and recognizing and classifying objects, A recording means for recording and storing the positional information of the analyzed object, A search means that searches for the recorded location information based on input from the user and provides a response, An object management system that includes this.
2. The object management system according to claim 1, characterized in that the analysis means has a function to recognize objects using a generative artificial intelligence model.
3. The object management system according to claim 1, characterized in that the search means performs object searches by analyzing the user's voice input using natural language processing.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A