System
A system with a traveling device, server, and user terminal provides real-time detailed property information, addressing the challenge of obtaining accurate environmental data for efficient property evaluation.
Patent Information
- Application Number
- JP2024133543
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Conventional property viewings face challenges in obtaining detailed information about properties, such as environmental data, which limits accurate evaluation and efficient selection by prospective buyers.
A system comprising a traveling device within a property that captures video and audio data, a server for analysis, and a user terminal for displaying answers, enabling real-time generation and display of detailed property information using a generation engine.
Enables efficient and accurate property evaluation by providing detailed environmental information, such as noise levels, vibrations, and sunlight, allowing users to make informed decisions.
Smart Images

Figure 2026030560000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In conventional property viewings, much of the information could only be obtained by actually visiting the property, and there were time and cost constraints. It was also difficult to obtain detailed information about the property's environmental data (noise, vibration, sunlight, etc.). As a result, prospective buyers were often unable to fully evaluate the property's true value and suitability. The present invention aims to solve these problems and enable efficient and accurate property selection. [Means for solving the problem]
[0005] To solve the above problems, the following means is provided: Video and audio data captured by a traveling device installed within a property is received and analyzed. Answers to specific questions are generated using a generation engine based on the analysis results, and the generated answers are displayed via a user interface. In addition to the video and audio data, the received data also includes environmental data such as noise levels, vibrations, and sunlight. This allows prospectors to obtain detailed environmental information about the property in real time, enabling them to more accurately evaluate the property.
[0006] A "traveling device" is a device that collects video and audio data while moving around a property autonomously or remotely.
[0007] "Video data" refers to data containing visual information of the inside of a property captured by the traveling device.
[0008] "Audio data" refers to data that includes acoustic information captured by the traveling device within the property.
[0009] "Analysis" is a process for extracting specific information or features based on received video and audio data.
[0010] A "generation engine" is software or algorithm that uses the results of the analysis to generate natural language answers to specific questions.
[0011] A "user interface" is an interface through which a user inputs information and displays analysis results and generated answers.
[0012] "Environmental Data" refers to data relating to the noise level, vibrations, and amount of sunlight at the property. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2]1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] System Configuration
[0035] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves around the property. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal through a user interface.
[0036] System Operation
[0037] server
[0038] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including the property's noise level, vibration, and sunlight. It then analyzes the received data and extracts specific information to respond to the user's question. For example, if the question is about sunlight, it evaluates the changes in light during the day based on the video data. If the question is about noise levels, it analyzes the audio data and evaluates the noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device.
[0039] Terminal
[0040] The terminal accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. By checking these, the user can efficiently obtain detailed property information.
[0041] User
[0042] Users access the system through a dedicated application. Using the application's interface, they can begin viewing the property and remotely control the traveling device. During the viewing, users can enter questions about information they are interested in (e.g., "How sunny is the bedroom?") and receive answers in real time. Based on the answers, users can proceed with their property evaluation.
[0043] Specific examples
[0044] For example, consider the case where a user asks about sunlight in the living room. Using a dedicated application, the user inputs the question, "How much sunlight does the living room get?" The device sends this question to the server, which analyzes the video data of the living room. The server evaluates the sunlight in the living room from the analysis results and generates an answer using a natural language generation engine, such as "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface. This allows the user to obtain detailed information about the sunlight in the living room without actually visiting the property.
[0045] As described above, the present invention provides specific and detailed property information to users by combining a traveling device, server, and terminal installed within a property, allowing users to efficiently evaluate properties and select the most suitable home.
[0046] The processing flow will be explained below.
[0047] Step 1:
[0048] server
[0049] The server receives the activation signal from the traveling device and activates it within the property. The traveling device activates the camera and microphone and begins collecting video and audio data in real time.
[0050] Step 2:
[0051] server
[0052] The server receives video and audio data transmitted from the traveling device in real time, as well as environmental data such as noise levels, vibrations, and sunlight. The received data is temporarily stored.
[0053] Step 3:
[0054] User
[0055] The user launches the dedicated application and begins viewing the property. The user inputs a specific question (e.g., "How sunny is the living room?") through the interface.
[0056] Step 4:
[0057] Terminal
[0058] The terminal receives the user's question and transmits the question to the server.
[0059] Step 5:
[0060] server
[0061] The server analyzes the received question, identifies the question, and selects and analyzes relevant data (e.g., video data on sunlight) based on the identified question.
[0062] Step 6:
[0063] server
[0064] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0065] Step 7:
[0066] server
[0067] The server generates a response and sends it to the terminal.
[0068] Step 8:
[0069] Terminal
[0070] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[0071] Step 9:
[0072] User
[0073] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[0074] In this way, the program proceeds step by step, providing the user with specific and detailed property information in real time.
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] With conventional property viewing systems, it was difficult for users to obtain detailed information about the property in real time, resulting in low accuracy in property evaluations. Furthermore, there was a lack of efficient means for collecting and analyzing environmental data such as noise levels and sunlight, making it impossible for users to obtain the detailed answers they desired. This made it difficult for users to efficiently view properties and select the most suitable home.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, means for receiving and analyzing environmental data including noise levels, vibrations, and sunlight levels within the property, means for transmitting questions input by users to the server, and means for displaying the analysis results received from the server and the generated answers on the user interface, thereby enabling users to obtain detailed information about the property in real time and efficiently evaluate the property.
[0080] A "traveling device" is a device that collects video and audio data while moving around the property.
[0081] The "server" is the core computer system that receives and analyzes data sent from the traveling device and generates answers to user questions.
[0082] A "user terminal" is a device that allows a user to access the system using a dedicated application, enter questions, and check answers.
[0083] "Video data" refers to data containing visual information about the inside of a property, and is captured by a traveling device.
[0084] "Audio data" refers to data containing auditory information within a property, and is recorded by the traveling device.
[0085] "Environmental data" refers to data that includes information about the environment within the property, such as noise levels, vibrations, and sunlight.
[0086] "Analysis" is the process of extracting information from the received data and identifying information that responds to the user's question.
[0087] The "generation engine" is an engine for generating specific answers to user questions based on the analysis results.
[0088] A "generative AI model" is an artificial intelligence model that generates natural language responses based on input prompts.
[0089] A "prompt sentence" is a sentence that is input into a generative AI model and serves as the basis for the answer to a question that is generated based on the analysis results.
[0090] "User interface" refers to an interface that allows a user to interact with a system, and includes the screen displayed on the terminal and the operation method.
[0091] MODE FOR CARRYING OUT THE INVENTION
[0092] System configuration
[0093] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves within the property and transmits it to the server. The server analyzes the received data and generates answers to users' questions. The generated answers are displayed on the user terminal through a user interface. Users can access the system via a dedicated application and input questions.
[0094] Hardware and software used
[0095] Mobile device: A device that moves autonomously or remotely around a property and collects video and audio data using cameras and microphones.
[0096] Server: A high-performance computer that provides an environment for receiving data, analyzing it, and running generative AI models. For example, it can use cloud-based computing services.
[0097] User device: A device such as a smartphone, tablet, or PC with a dedicated application installed.
[0098] Data processing and calculation
[0099] 1. Data reception:
[0100] The server receives real-time video and audio data transmitted from the traveling device, as well as environmental data such as noise levels, vibrations, and sunlight levels at the property.
[0101] 2. Data Analysis:
[0102] The server analyzes the received data and extracts specific information that responds to the user's question, for example by implementing algorithms that analyze video data to evaluate changes in light and shadow movement during the day.
[0103] 3. Answer generation:
[0104] Based on the analysis results, the generative AI model creates a prompt sentence and generates a natural language answer based on that prompt sentence. For example, the prompt sentence could be, "The user is asking about the sunlight in the living room. Please evaluate the sunlight in the living room based on the video data of the living room and generate a natural language answer explaining the results."
[0105] 4. Answer display:
[0106] The generated answer is sent to the user terminal through the user interface and displayed on the application screen.
[0107] Specific examples
[0108] For example, if a user asks about sunlight in the living room, the following process is performed.
[0109] 1. User enters a question: The user uses a dedicated application to enter a question such as, "How much sunlight does the living room get?"
[0110] 2. The device sends the question to the server: The device sends this question to the server.
[0111] 3. The server analyzes the data: The server analyzes the video data from the living room and evaluates the amount of sunlight.
[0112] 4. Answer generation: Based on the prompt, the server's generative AI model generates the answer, "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0113] 5. Displaying the answer: This answer is sent to the terminal and displayed to the user through the user interface.
[0114] Prompt Sentence Examples
[0115] "The user is asking about the sunlight in the living room. Based on the video data of the living room, please evaluate the sunlight in the living room and generate a natural language answer explaining the results."
[0116] As described above, this system efficiently provides users with specific and detailed property information and helps them evaluate properties, enabling them to obtain information to select the most suitable home.
[0117] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0118] Step 1: User enters question
[0119] The user launches the dedicated application and inputs a question through the interface, such as "How much sunlight does the living room get?" The application temporarily stores the question in memory.
[0120] Input: Questions entered into the user interface
[0121] Output: Question data stored in memory
[0122] Specific behavior:
[0123] The user starts the application and a question input screen is displayed.
[0124] Enter your question and tap the "Submit" button.
[0125] Step 2: Sending a question to the server from the device
[0126] The device uses a dedicated communication protocol to send the question data to the server, such as an HTTP POST request.
[0127] Input: Question data stored in memory
[0128] Output: Question data sent to the server
[0129] Specific behavior:
[0130] The device retrieves the question data stored in memory and sends it to the server as an HTTP POST request.
[0131] Step 3: Server receives data
[0132] The server receives the query data sent from the terminal. It then begins receiving video and audio data sent from the traveling device in real time. It also simultaneously receives environmental data such as noise levels, vibrations, and sunlight.
[0133] Input: Question data sent from the terminal, video and audio data from the driving device, environmental data
[0134] Output: Question data and various data stored on the server
[0135] Specific behavior:
[0136] The server receives the query data and stores it in a database.
[0137] The server receives data from the running gear and environmental sensors in real time and prepares it for analysis.
[0138] Step 4: Data analysis by the server
[0139] The server analyzes the data it receives and extracts information to answer questions, such as analyzing video data to assess changes in light during the day or audio data to assess noise levels.
[0140] Input: Video, audio, environmental data, and question data stored on the server
[0141] Output: Specific analysis results that address your questions
[0142] Specific behavior:
[0143] An algorithm is run to detect changes in light from the video data.
[0144] Audio data is analyzed to evaluate noise levels by time of day.
[0145] Step 5: Server Generates Answer
[0146] Based on the analysis results, the server creates a prompt to be applied to the generative AI model. The prompt is input into the generative AI model, which generates a natural language answer. For example, it generates an answer such as, "The living room has good sunlight in the morning and is partially shaded in the afternoon."
[0147] Input: Analysis results, generative AI model, prompt
[0148] Output: Natural language response
[0149] Specific behavior:
[0150] The server creates a prompt based on the analysis results.
[0151] The prompt sentence is input into a generative AI model to generate a natural language response.
[0152] Step 6: Sending the response from the server to the device
[0153] The server sends the generated response to the user's terminal using a communication protocol such as an HTTP POST request.
[0154] Input: Natural language response data
[0155] Output: Answer data sent to the user's device
[0156] Specific behavior:
[0157] The server sends the generated response to the device as an HTTP POST request.
[0158] Step 7: Viewing the Answers on Your Device
[0159] The terminal stores the response data received from the server in a memory and displays it on a user interface.
[0160] Input: Response data received from the server
[0161] Output: The answer displayed in the user interface
[0162] Specific behavior:
[0163] The terminal analyzes the received response data and displays it on the user interface.
[0164] (Application example 1)
[0165] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0166] Management and monitoring of logistics facilities are important for improving efficiency and ensuring safety. However, conventional systems make it difficult to physically inspect the site and integrate and analyze a wide range of environmental data, placing a heavy burden on managers. Furthermore, there are challenges in responding quickly to abnormalities and providing detailed information in real time.
[0167] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0168] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within the property, means for analyzing the received video and audio data, and a generation engine means for generating answers to specific questions based on the analysis results. This enables means for receiving and analyzing environmental data for monitoring and management within the logistics facility, such as noise levels, vibrations, and lighting conditions, means for issuing automatic alerts when an abnormality is detected, means for allowing a manager of the logistics facility to input a question and generating a specific answer to the question using a generative AI model, and means for displaying the answer to the question via a user interface.
[0169] A "traveling device installed within the property" is a device that moves autonomously or remotely within the logistics facility and collects video and audio data.
[0170] The "means for receiving video and audio data" refers to a means having a function for transmitting video and audio data collected by the traveling device to a server in real time and for the server to receive the data.
[0171] "Means for analyzing video and audio data" refers to technical means for analyzing received video and audio data and extracting specific information, and examples include video analysis software and audio analysis software.
[0172] The "generation engine means" refers to means including algorithms and programs for generating specific answers to user questions based on analyzed data.
[0173] The "means for displaying via a user interface" refers to a means including an interface and its functions for displaying the generated answer on a user terminal.
[0174] "Means for receiving and analyzing environmental data" refers to the technical means for collecting and analyzing environmental data such as noise levels, vibrations, and lighting conditions within a logistics facility.
[0175] "Means for issuing automatic alerts when an abnormality is detected" refers to a means that has the function of immediately notifying an administrator when an abnormality is detected based on collected and analyzed data.
[0176] "Generative AI models" are artificial intelligence algorithms and models used for natural language processing and information generation, examples of which include GPT-4.
[0177] The "means for displaying answers to questions via a user interface" refers to an interface and means including its functions for displaying answers generated by the generation engine to questions from users on the display screen of a user terminal.
[0178] System Configuration
[0179] This invention is a system consisting of a traveling device, a server, and a user terminal installed within a logistics facility. The traveling device collects video and audio data as it moves within the logistics facility and transmits it to the server. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal via a user interface.
[0180] Specific operation of the system
[0181] server
[0182] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including noise levels, vibrations, and lighting conditions within the logistics facility. It then analyzes the received data and extracts specific information to respond to user questions. For example, if a question is about noise levels, the server analyzes the audio data and evaluates noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device via the user interface.
[0183] The server uses the following specific technologies:
[0184] Hardware: Server machine
[0185] Software: Video analysis software (e.g., OpenCV), audio analysis software (e.g., DeepSpeech), natural language generation engines (e.g., GPT-4), real-time communication systems (e.g., WebSocket)
[0186] Specific examples
[0187] For example, consider the case where a manager asks about the noise level in a specific location in a logistics facility. Using a dedicated application, the manager inputs the question, "What is the current noise level?" The device sends this question to the server, which analyzes the collected noise data. The server evaluates the noise level based on the analysis results and generates an answer using a natural language generation engine, such as "The current noise level is an average of 85 decibels." This answer is sent to the device and displayed to the manager through a user interface. This allows the manager to obtain detailed information about the specific situation at the logistics facility.
[0188] User terminal
[0189] The terminal accepts questions from the user and sends them to the server. It receives the analysis results and generated answers from the server and displays them on the user interface. By checking these, the user can efficiently obtain detailed information about the logistics facility. The system also has a function that displays an automatic alert if an abnormality is detected.
[0190] User
[0191] Users access the system through a dedicated application. Using the application's interface, they can start monitoring the logistics facility and remotely control the traveling device. While monitoring, users can input questions about information of interest (e.g., "What is the lighting situation in this area?") and receive answers in real time. Based on the answers, users can evaluate and manage the logistics facility.
[0192] Prompt Sentence Examples
[0193] People: What is the noise level in your current distribution center?
[0194] AI: The current noise level averages 85 decibels, with peaks occurring between 2 and 3 p.m.
[0195] As described above, this system combines on-site traveling devices with a server and terminals to provide users with specific and detailed information about logistics facilities, allowing them to efficiently manage and evaluate the facilities.
[0196] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0197] Step 1:
[0198] The user launches the dedicated application and begins monitoring the logistics facility.
[0199] Input: The user launches the application
[0200] Processing: The user interface is launched and the running device is ready for remote operation.
[0201] Output: The control screen for the running gear is displayed.
[0202] Step 2:
[0203] The user remotely controls the traveling device, which patrols the logistics facility and collects video and audio data.
[0204] Input: User operates the running gear
[0205] Processing: As the traveling device moves, it collects video and audio using cameras and microphones.
[0206] Output: Real-time collected video and audio data
[0207] Step 3:
[0208] The traveling device transmits the collected video and audio data to a server.
[0209] Input: Video and audio data from the traveling device
[0210] Processing: Data is sent to the server via wireless communication
[0211] Output: Video and audio data received by the server
[0212] Step 4:
[0213] The server parses the received data.
[0214] Input: Video and audio data received by the server
[0215] Processing: The server analyzes the data using video analysis software (e.g., OpenCV) and audio analysis software (e.g., DeepSpeech).
[0216] Output: Analysis results (e.g. noise level, specific object recognition)
[0217] Step 5:
[0218] The user inputs a question via a dedicated application.
[0219] Input: User types a question into the device (e.g., "What is the current noise level?")
[0220] Process: The question is sent to the server
[0221] Output: The question sent to the server
[0222] Step 6:
[0223] Based on the analysis results, the server generates an answer to the question using a generative AI model (e.g., GPT-4).
[0224] Input: The question and analysis results sent to the server
[0225] Processing: The generative AI model generates a natural language answer to the question.
[0226] Output: The generated answer (e.g., "The current noise level is 85 decibels")
[0227] Step 7:
[0228] The generated answer is displayed on the user terminal via a user interface.
[0229] Input: Generated Answer
[0230] Processing: The answer is sent to the terminal and displayed in the user interface.
[0231] Output: The answer displayed in the user interface
[0232] Step 8:
[0233] The server will automatically send an alert if an abnormality is detected.
[0234] Input: Analysis results (e.g., noise level exceeds threshold)
[0235] Action: The server generates an automatic alert and notifies the user interface.
[0236] Output: The alert displayed in the user interface
[0237] In this way, the entire system can monitor detailed conditions within the logistics facility in real time, enabling efficient and safe operation and management.
[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0239] System Configuration
[0240] The present invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. The traveling device collects video, audio, and environmental data as it moves within the property. The server receives this data and analyzes it based on the user's questions. The emotion engine then recognizes emotions from the user's audio and facial data, and displays the analysis results and generated answers on the user terminal via a user interface.
[0241] System Operation
[0242] server
[0243] The server first receives video, audio, and environmental data transmitted from the traveling device in real time. This includes data on noise levels, vibrations, and sunlight. The server then temporarily stores the received data and analyzes it based on the user's questions. For example, video data is used for questions about sunlight, and audio data is analyzed for questions about noise. Specific information is extracted and a specific answer is generated using a generation engine. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and provides feedback according to the user's emotional state.
[0244] Terminal
[0245] The device accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. Furthermore, the device also displays the analysis results of the emotion engine, adjusting the information provided and questions asked during the viewing as necessary.
[0246] User
[0247] Users access the system through a dedicated application. Using the application's interface, they can begin viewing properties and remotely control the vehicle. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[0248] Specific examples
[0249] For example, consider the case where a user asks about sunlight in the living room. The user inputs the question, "How sunny is the living room?" The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer generated is, "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface.
[0250] If a user shows a surprised expression during a viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, and appropriate feedback will be provided according to the user's emotion.
[0251] As described above, the present invention combines a traveling device placed within a property with a server, a terminal, and an emotion engine to provide specific and detailed property information and personalized feedback to users, allowing them to efficiently evaluate properties and choose the most suitable home.
[0252] The processing flow will be explained below.
[0253] Step 1:
[0254] server
[0255] The server receives the activation signal from the vehicle and activates it within the property. The vehicle then activates its camera and microphone to begin collecting video and audio data in real time. It also collects environmental data such as noise levels, vibrations, and sunlight.
[0256] Step 2:
[0257] server
[0258] The server receives video, audio and environmental data transmitted from the traveling device in real time and temporarily stores this data.
[0259] Step 3:
[0260] User
[0261] The user launches the dedicated application and begins viewing the property, entering a specific question (e.g., "How sunny is the living room?") through the application's interface.
[0262] Step 4:
[0263] Terminal
[0264] The terminal receives the user's question and transmits the question to the server.
[0265] Step 5:
[0266] server
[0267] The server analyzes the received question, identifies the content of the question, and selects and analyzes data related to the question, such as video data and audio data related to sunlight.
[0268] Step 6:
[0269] server
[0270] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0271] Step 7:
[0272] server
[0273] The server generates a response and sends it to the terminal.
[0274] Step 8:
[0275] Terminal
[0276] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[0277] Step 9:
[0278] User
[0279] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[0280] Step 10:
[0281] server
[0282] The server collects the user's voice and facial data and analyzes it with an emotion engine to recognize the user's emotional state (e.g., joy, sadness, surprise, anger).
[0283] Step 11:
[0284] server
[0285] Based on the analysis results of the emotion engine, feedback and additional information are generated according to the user's emotional state. For example, if the user is surprised, a message such as "Would you like to know more about the part that surprised you?" is generated.
[0286] Step 12:
[0287] server
[0288] Feedback and additional information from the emotion engine is sent to the device.
[0289] Step 13:
[0290] Terminal
[0291] The device displays feedback and additional information in the user interface according to the user's emotions, allowing the user to view the property in more detail.
[0292] In this way, the program proceeds step by step, providing the user with specific and detailed property information and emotional feedback in real time.
[0293] Example 2
[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] Conventional property viewing systems are limited to collecting and analyzing video and audio data, and are unable to provide feedback tailored to the user's emotional state. Furthermore, they are limited in generating answers when users input specific questions, and lack the ability to provide personalized information in real time. This makes it difficult for users to efficiently obtain more accurate and detailed property information.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0297] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation AI model means for generating answers to specific questions based on the analysis results, means for recognizing emotions from the user's voice data and face data, means for generating feedback based on the emotion recognition results and the analysis results, and means for displaying the generated answers and feedback via a user interface. This allows the user to obtain specific and detailed property information in real time and receive personalized feedback according to their emotions.
[0298] A "traveling device" is a device that collects video data, audio data, and environmental data while moving around the property.
[0299] "Video data" refers to data of moving images and still images taken by the traveling device.
[0300] "Audio data" is data of sounds collected by the traveling device.
[0301] "Environmental data" refers to data that includes information such as noise levels, vibrations, and sunlight levels within the property.
[0302] A "generative AI model" is an artificial intelligence model that analyzes received data and generates answers to specific questions.
[0303] "Emotion recognition" is the process of analyzing a user's voice and facial data to identify the user's emotional state.
[0304] A "user interface" is a screen or operating means that allows a user to exchange information with a system.
[0305] "Feedback" refers to responses or information provided to the user based on the analysis results and emotion recognition results.
[0306] "Real-time" refers to immediate processing or response with little or no delay.
[0307] "Analysis" is the process of extracting information and deriving meaning from received data.
[0308] The present invention is a property viewing system that includes a traveling device, a server, a user terminal, and an emotion engine. The traveling device moves around the property and collects video, audio, and environmental data. The server receives this data in real time and analyzes it based on the user's questions. Furthermore, the emotion engine recognizes emotions from the user's voice data and facial data, and displays the analysis results and generated answers on the user terminal through a user interface.
[0309] Hardware and software used
[0310] Hardware:
[0311] Mobile devices (e.g., robotic cameras, mobile sensor devices)
[0312] Servers (e.g., high-performance servers in data centers)
[0313] User device (e.g. smartphone, tablet)
[0314] Hardware for emotion engine (e.g. high-resolution camera, directional microphone)
[0315] software:
[0316] Data collection software (e.g. camera control software, sensor interface)
[0317] Data analysis software (e.g., image analysis algorithms, audio analysis algorithms)
[0318] Emotion recognition software (e.g., voice emotion recognition model, facial expression analysis model)
[0319] User interfaces (e.g., mobile apps, web applications)
[0320] Specific operation of the system
[0321] The server first receives video and audio data and environmental data (such as noise levels, vibrations, and sunlight) sent from the driving device in real time. This data is temporarily stored and analyzed based on questions from the user. Specific information is extracted using various algorithms for analysis, and specific answers are generated using generative AI models. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and generates feedback according to the user's emotional state.
[0322] The device accepts questions from the user as input and sends them to the server. The analysis results and generated answers are received from the server in real time and displayed through the user interface. Furthermore, the emotion engine's analysis results are also displayed, and the information provided during the viewing and questions can be adjusted as needed.
[0323] Users access the system through a dedicated application. Using the application's interface, they can start viewing properties and remotely control the traveling device. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[0324] Specific use cases
[0325] For example, consider the case where a user asks a question about the sunlight in the living room. The user enters "How sunny is the living room?" into the device. The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer "The living room gets good sunlight in the morning and is partially shaded in the afternoon" is generated. This answer is sent to the device and displayed to the user through the user interface.
[0326] Furthermore, if the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, providing appropriate feedback according to the user's emotion.
[0327] Examples of prompt statements
[0328] An example of a prompt is as follows:
[0329] prompt:
[0330] How's the sunlight in the living room?
[0331] Expected answer:
[0332] "The living room gets good sunlight in the morning, but some shading is expected in the afternoon."
[0333] As described above, the present invention combines a traveling device, a server, a user terminal, and an emotion engine to provide users with specific and detailed property information and personalized feedback, allowing them to efficiently evaluate properties and select the most suitable home.
[0334] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0335] Program processing flow
[0336] Step 1: Data collection
[0337] Processing Description:
[0338] The vehicle travels around the property, collecting video and audio data using cameras and microphones, and environmental sensors to collect environmental data such as noise levels, vibrations, and sunlight.
[0339] input:
[0340] Video, audio, noise level, vibration, sunlight
[0341] output:
[0342] Real-time transmission of video data, audio data, and environmental data
[0343] Specific operation:
[0344] As the traveling device moves around the living room, it takes pictures of the entire room with a camera, collects sound with a microphone, and measures the temperature with a sensor.
[0345] Step 2: Data reception and storage
[0346] Processing Description:
[0347] The server receives video data, audio data, and environmental data transmitted from the traveling device in real time and temporarily stores it.
[0348] input:
[0349] Data transmitted from the traveling device (video, audio, environmental data)
[0350] output:
[0351] Data stored in the database
[0352] Specific operation:
[0353] The server stores the received data in a dedicated database and prepares it for analysis.
[0354] Step 3: Ask a question
[0355] Processing Description:
[0356] The terminal accepts a question from the user as input and sends it to the server.
[0357] input:
[0358] User questions (e.g., "How sunny is the living room?")
[0359] output:
[0360] Send the question to the server
[0361] Specific operation:
[0362] The user enters a question in the application interface and presses the submit button, which sends the question to the server via the API.
[0363] Step 4: Data analysis
[0364] Processing Description:
[0365] The server analyzes the question, filters and extracts relevant video, audio, and environmental data, and uses a generative AI model to generate a specific answer.
[0366] input:
[0367] User questions, saved data (video, audio, environmental data)
[0368] output:
[0369] Generated Answer
[0370] Specific operation:
[0371] The server analyzes the question "How much sunlight does the living room get?", filters and analyzes the video data of the living room, and uses a generative AI model to generate the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon."
[0372] Step 5: Emotion Recognition
[0373] Processing Description:
[0374] The emotion engine recognizes emotions from the user's voice data and facial data, and sends the recognition results to the server, which then analyzes the emotion data and generates feedback.
[0375] input:
[0376] User voice data, face data
[0377] output:
[0378] Emotion recognition results and generated feedback
[0379] Specific operation:
[0380] The emotion engine monitors the user's facial expressions to detect smiles and surprises. For example, if the user shows a surprised expression, the emotion engine recognizes the surprise and generates feedback such as, "Would you like to provide additional information about what is surprising?"
[0381] Step 6: Submit and view your answers and feedback
[0382] Processing Description:
[0383] The server transmits the generated answers and feedback to the user terminal, which displays the received information on a user interface.
[0384] input:
[0385] Generated answers, generated feedback
[0386] output:
[0387] Data sent to user terminals, display data
[0388] Specific operation:
[0389] The server sends the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon" and feedback "Do you like this property?" to the user's device. The device receives this information and displays it on the application screen.
[0390] In this way, the entire system works together to consistently collect and analyze data, provide real-time answers to user questions, and even provide emotional feedback.
[0391] (Application example 2)
[0392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0393] Conventional property viewing systems required users to personally visit the property, which was time-consuming and costly, and the information obtained from a single viewing was limited. Furthermore, it was difficult to provide personalized information based on the user's emotions and personal preferences, resulting in insufficient evaluations of properties depending on the user. Furthermore, due to a lack of analysis of environmental data and appropriate feedback, users were sometimes unaware of potential problems with the property. Thus, to improve user convenience and satisfaction, a more detailed and individually tailored property information system was needed.
[0394] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0395] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, and emotion recognition means for recognizing the user's emotions and providing feedback according to the user's emotional state. This allows the user to remotely view the property and receive answers in real time, and also allows the user to be provided with information according to their individual emotions, making it possible to efficiently obtain more detailed and personalized property information.
[0396] A "moving device" is a device that moves autonomously or remotely within a property to collect video and audio data.
[0397] The "generation engine means" is an engine that has the function of generating an answer to a specific question from a user based on the analysis results.
[0398] A "user interface" is an interface for displaying generated answers and analysis results to the user, and is a means for exchanging information between the user and the system.
[0399] The "emotion recognition means" is a means that has the function of recognizing emotions from data such as the user's voice and facial expression, and providing feedback according to those emotions.
[0400] "Environmental data" refers to data relating to environmental conditions such as noise levels, vibrations, and sunlight levels within the property.
[0401] A "visualization device" is a device worn by a user that collects video and audio data and provides the user with visual and audio information in real time.
[0402] A "biological signal" is a signal that contains information about the user's health condition and fatigue level, and includes heart rate, body temperature, skin potential, and the like.
[0403] "Personalized information" refers to information and feedback that is customized based on a user's individual emotional and health state.
[0404] This invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. Specific embodiments of this system are described below.
[0405] System Configuration
[0406] The server has a means for receiving video and audio data captured by the traveling device installed within the property. The received data is processed using an analysis engine to generate an optimal answer based on the user's question. The generated answer is displayed on the user's terminal via a user interface. In addition, an emotion recognition means is used to analyze the user's emotional state and provide feedback as necessary.
[0407] Hardware and Software
[0408] The system uses the following hardware and software:
[0409] Mobile device: Responsible for moving around the property and collecting video and audio data.
[0410] Server: Receives and analyzes data, generates answers via a generation engine, and recognizes emotions.
[0411] User terminal (smartphone, smart glasses, head-mounted display, etc.): displays information through a user interface.
[0412] Emotion engine: Recognizes emotions from the user's facial data and voice.
[0413] Software used: OpenCV (video data processing), EmotionRecognizer (emotion recognition), Server (communication with server).
[0414] Data processing and calculation
[0415] The server receives video and audio data from the vehicle in real time and temporarily stores it. The stored data is analyzed to extract specific information in response to user questions. A generative AI model is used to generate specific answers, which are then displayed to the user through a user interface.
[0416] The emotion engine analyzes the user's voice and facial data to recognize their emotions. The recognized emotions are sent to the server, which provides appropriate feedback along with the analysis results.
[0417] Specific examples
[0418] For example, if a user asks about the noise level in a room, the user inputs the question, "What is the noise level in this room?" The device sends this question to the server, which analyzes the voice data. As a result of the analysis, the answer "The average noise level in this room is 50 decibels" is generated and sent to the device. This answer is then displayed to the user through the user interface.
[0419] If the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and provide appropriate feedback according to the user's emotions, such as displaying a message like, "Would you like to provide additional information about the part that surprised you?"
[0420] Prompt Sentence Examples
[0421] "When you inquire about a new product, please send us the data of the moment that surprised you the most."
[0422] "Analyze user interests using facial recognition data and suggest appropriate products."
[0423] This system allows users to efficiently obtain more detailed and personalized information, and also provides feedback according to the user's emotions, thereby improving user satisfaction.
[0424] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0425] Step 1:
[0426] The device accepts questions from users as input. The user inputs the question through an interface such as smart glasses or a smartphone. The input data is sent to the server. The input content may include, for example, "What is the noise level in this room?"
[0427] Step 2:
[0428] The server receives the query data sent from the device. Based on the received query data, it acquires and temporarily stores the video and audio data collected from the traveling device. The temporarily stored data includes video of the room, environmental sounds, vibration data, and sunlight amount.
[0429] Step 3:
[0430] The server analyzes the temporarily stored data. Specifically, if a question is about noise levels, for example, the server analyzes the audio data. This analysis analyzes the noise level and frequency components, and outputs the noise level in decibels (dB). The analyzed data is then converted into a specific answer by a generation engine.
[0431] Step 4:
[0432] The generation engine uses the analysis results to create a generated answer, which may contain specific information such as "The average noise level in this room is 50 decibels," and sends the generated answer to the user interface.
[0433] Step 5:
[0434] The terminal receives the generated answer data sent from the server and displays it to the user via a user interface, and if the user is using smart glasses, the answer is presented visually.
[0435] Step 6:
[0436] The emotion recognition means analyzes the user's voice data and face data. The emotion engine analyzes the user's facial expression and tone of voice when they input a question, and recognizes the user's emotional state (for example, surprise, satisfaction, dissatisfaction, etc.).
[0437] Step 7:
[0438] The server analyzes the emotion data sent from the emotion engine and provides additional feedback as needed. For example, if the user shows a surprised expression, a message such as "Would you like to provide additional information about why you are surprised?" is generated and sent to the device.
[0439] Step 8:
[0440] The terminal receives the emotion-based feedback data transmitted from the server and displays it to the user via a user interface, allowing the user to receive information corresponding to their emotions in real time.
[0441] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0442] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0443] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0444] [Second embodiment]
[0445] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0446] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0447] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0448] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0449] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0451] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0452] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0453] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0454] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0455] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0456] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0457] System Configuration
[0458] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves around the property. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal through a user interface.
[0459] System Operation
[0460] server
[0461] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including the property's noise level, vibration, and sunlight. It then analyzes the received data and extracts specific information to respond to the user's question. For example, if the question is about sunlight, it evaluates the changes in light during the day based on the video data. If the question is about noise levels, it analyzes the audio data and evaluates the noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device.
[0462] Terminal
[0463] The terminal accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. By checking these, the user can efficiently obtain detailed property information.
[0464] User
[0465] Users access the system through a dedicated application. Using the application's interface, they can begin viewing the property and remotely control the traveling device. During the viewing, users can enter questions about information they are interested in (e.g., "How sunny is the bedroom?") and receive answers in real time. Based on the answers, users can proceed with their property evaluation.
[0466] Specific examples
[0467] For example, consider the case where a user asks about sunlight in the living room. Using a dedicated application, the user inputs the question, "How much sunlight does the living room get?" The device sends this question to the server, which analyzes the video data of the living room. The server evaluates the sunlight in the living room from the analysis results and generates an answer using a natural language generation engine, such as "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface. This allows the user to obtain detailed information about the sunlight in the living room without actually visiting the property.
[0468] As described above, the present invention provides specific and detailed property information to users by combining a traveling device, server, and terminal installed within a property, allowing users to efficiently evaluate properties and select the most suitable home.
[0469] The processing flow will be explained below.
[0470] Step 1:
[0471] server
[0472] The server receives the activation signal from the traveling device and activates it within the property. The traveling device activates the camera and microphone and begins collecting video and audio data in real time.
[0473] Step 2:
[0474] server
[0475] The server receives video and audio data transmitted from the traveling device in real time, as well as environmental data such as noise levels, vibrations, and sunlight. The received data is temporarily stored.
[0476] Step 3:
[0477] User
[0478] The user launches the dedicated application and begins viewing the property. The user inputs a specific question (e.g., "How sunny is the living room?") through the interface.
[0479] Step 4:
[0480] Terminal
[0481] The terminal receives the user's question and transmits the question to the server.
[0482] Step 5:
[0483] server
[0484] The server analyzes the received question, identifies the question, and selects and analyzes relevant data (e.g., video data on sunlight) based on the identified question.
[0485] Step 6:
[0486] server
[0487] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0488] Step 7:
[0489] server
[0490] The server generates a response and sends it to the terminal.
[0491] Step 8:
[0492] Terminal
[0493] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[0494] Step 9:
[0495] User
[0496] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[0497] In this way, the program proceeds step by step, providing the user with specific and detailed property information in real time.
[0498] Example 1
[0499] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0500] With conventional property viewing systems, it was difficult for users to obtain detailed information about the property in real time, resulting in low accuracy in property evaluations. Furthermore, there was a lack of efficient means for collecting and analyzing environmental data such as noise levels and sunlight, making it impossible for users to obtain the detailed answers they desired. This made it difficult for users to efficiently view properties and select the most suitable home.
[0501] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0502] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, means for receiving and analyzing environmental data including noise levels, vibrations, and sunlight levels within the property, means for transmitting questions input by users to the server, and means for displaying the analysis results received from the server and the generated answers on the user interface, thereby enabling users to obtain detailed information about the property in real time and efficiently evaluate the property.
[0503] A "traveling device" is a device that collects video and audio data while moving around the property.
[0504] The "server" is the core computer system that receives and analyzes data sent from the traveling device and generates answers to user questions.
[0505] A "user terminal" is a device that allows a user to access the system using a dedicated application, enter questions, and check answers.
[0506] "Video data" refers to data containing visual information about the inside of a property, and is captured by a traveling device.
[0507] "Audio data" refers to data containing auditory information within a property, and is recorded by the traveling device.
[0508] "Environmental data" refers to data that includes information about the environment within the property, such as noise levels, vibrations, and sunlight.
[0509] "Analysis" is the process of extracting information from the received data and identifying information that responds to the user's question.
[0510] The "generation engine" is an engine for generating specific answers to user questions based on the analysis results.
[0511] A "generative AI model" is an artificial intelligence model that generates natural language responses based on input prompts.
[0512] A "prompt sentence" is a sentence that is input into a generative AI model and serves as the basis for the answer to a question that is generated based on the analysis results.
[0513] "User interface" refers to an interface that allows a user to interact with a system, and includes the screen displayed on the terminal and the operation method.
[0514] MODE FOR CARRYING OUT THE INVENTION
[0515] System configuration
[0516] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves within the property and transmits it to the server. The server analyzes the received data and generates answers to users' questions. The generated answers are displayed on the user terminal through a user interface. Users can access the system via a dedicated application and input questions.
[0517] Hardware and software used
[0518] Mobile device: A device that moves autonomously or remotely around a property and collects video and audio data using cameras and microphones.
[0519] Server: A high-performance computer that provides an environment for receiving data, analyzing it, and running generative AI models. For example, it can use cloud-based computing services.
[0520] User device: A device such as a smartphone, tablet, or PC with a dedicated application installed.
[0521] Data processing and calculation
[0522] 1. Data reception:
[0523] The server receives real-time video and audio data transmitted from the traveling device, as well as environmental data such as noise levels, vibrations, and sunlight levels at the property.
[0524] 2. Data Analysis:
[0525] The server analyzes the received data and extracts specific information that responds to the user's question, for example by implementing algorithms that analyze video data to evaluate changes in light and shadow movement during the day.
[0526] 3. Answer generation:
[0527] Based on the analysis results, the generative AI model creates a prompt sentence and generates a natural language answer based on that prompt sentence. For example, the prompt sentence could be, "The user is asking about the sunlight in the living room. Please evaluate the sunlight in the living room based on the video data of the living room and generate a natural language answer explaining the results."
[0528] 4. Answer display:
[0529] The generated answer is sent to the user terminal through the user interface and displayed on the application screen.
[0530] Specific examples
[0531] For example, if a user asks about sunlight in the living room, the following process is performed.
[0532] 1. User enters a question: The user uses a dedicated application to enter a question such as, "How much sunlight does the living room get?"
[0533] 2. The device sends the question to the server: The device sends this question to the server.
[0534] 3. The server analyzes the data: The server analyzes the video data from the living room and evaluates the amount of sunlight.
[0535] 4. Answer generation: Based on the prompt, the server's generative AI model generates the answer, "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0536] 5. Displaying the answer: This answer is sent to the terminal and displayed to the user through the user interface.
[0537] Prompt Sentence Examples
[0538] "The user is asking about the sunlight in the living room. Based on the video data of the living room, please evaluate the sunlight in the living room and generate a natural language answer explaining the results."
[0539] As described above, this system efficiently provides users with specific and detailed property information and helps them evaluate properties, enabling them to obtain information to select the most suitable home.
[0540] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0541] Step 1: User enters question
[0542] The user launches the dedicated application and inputs a question through the interface, such as "How much sunlight does the living room get?" The application temporarily stores the question in memory.
[0543] Input: Questions entered into the user interface
[0544] Output: Question data stored in memory
[0545] Specific behavior:
[0546] The user starts the application and a question input screen is displayed.
[0547] Enter your question and tap the "Submit" button.
[0548] Step 2: Sending a question to the server from the device
[0549] The device uses a dedicated communication protocol to send the question data to the server, such as an HTTP POST request.
[0550] Input: Question data stored in memory
[0551] Output: Question data sent to the server
[0552] Specific behavior:
[0553] The device retrieves the question data stored in memory and sends it to the server as an HTTP POST request.
[0554] Step 3: Server receives data
[0555] The server receives the query data sent from the terminal. It then begins receiving video and audio data sent from the traveling device in real time. It also simultaneously receives environmental data such as noise levels, vibrations, and sunlight.
[0556] Input: Question data sent from the terminal, video and audio data from the driving device, environmental data
[0557] Output: Question data and various data stored on the server
[0558] Specific behavior:
[0559] The server receives the query data and stores it in a database.
[0560] The server receives data from the running gear and environmental sensors in real time and prepares it for analysis.
[0561] Step 4: Data analysis by the server
[0562] The server analyzes the data it receives and extracts information to answer questions, such as analyzing video data to assess changes in light during the day or audio data to assess noise levels.
[0563] Input: Video, audio, environmental data, and question data stored on the server
[0564] Output: Specific analysis results that address your questions
[0565] Specific behavior:
[0566] An algorithm is run to detect changes in light from the video data.
[0567] Audio data is analyzed to evaluate noise levels by time of day.
[0568] Step 5: Server Generates Answer
[0569] Based on the analysis results, the server creates a prompt to be applied to the generative AI model. The prompt is input into the generative AI model, which generates a natural language answer. For example, it generates an answer such as, "The living room has good sunlight in the morning and is partially shaded in the afternoon."
[0570] Input: Analysis results, generative AI model, prompt
[0571] Output: Natural language response
[0572] Specific behavior:
[0573] The server creates a prompt based on the analysis results.
[0574] The prompt sentence is input into a generative AI model to generate a natural language response.
[0575] Step 6: Sending the response from the server to the device
[0576] The server sends the generated response to the user's terminal using a communication protocol such as an HTTP POST request.
[0577] Input: Natural language response data
[0578] Output: Answer data sent to the user's device
[0579] Specific behavior:
[0580] The server sends the generated response to the device as an HTTP POST request.
[0581] Step 7: Viewing the Answers on Your Device
[0582] The terminal stores the response data received from the server in a memory and displays it on a user interface.
[0583] Input: Response data received from the server
[0584] Output: The answer displayed in the user interface
[0585] Specific behavior:
[0586] The terminal analyzes the received response data and displays it on the user interface.
[0587] (Application example 1)
[0588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] Management and monitoring of logistics facilities are important for improving efficiency and ensuring safety. However, conventional systems make it difficult to physically inspect the site and integrate and analyze a wide range of environmental data, placing a heavy burden on managers. Furthermore, there are challenges in responding quickly to abnormalities and providing detailed information in real time.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0591] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within the property, means for analyzing the received video and audio data, and a generation engine means for generating answers to specific questions based on the analysis results. This enables means for receiving and analyzing environmental data for monitoring and management within the logistics facility, such as noise levels, vibrations, and lighting conditions, means for issuing automatic alerts when an abnormality is detected, means for allowing a manager of the logistics facility to input a question and generating a specific answer to the question using a generative AI model, and means for displaying the answer to the question via a user interface.
[0592] A "traveling device installed within the property" is a device that moves autonomously or remotely within the logistics facility and collects video and audio data.
[0593] The "means for receiving video and audio data" refers to a means having a function for transmitting video and audio data collected by the traveling device to a server in real time and for the server to receive the data.
[0594] "Means for analyzing video and audio data" refers to technical means for analyzing received video and audio data and extracting specific information, and examples include video analysis software and audio analysis software.
[0595] The "generation engine means" refers to means including algorithms and programs for generating specific answers to user questions based on analyzed data.
[0596] The "means for displaying via a user interface" refers to a means including an interface and its functions for displaying the generated answer on a user terminal.
[0597] "Means for receiving and analyzing environmental data" refers to the technical means for collecting and analyzing environmental data such as noise levels, vibrations, and lighting conditions within a logistics facility.
[0598] "Means for issuing automatic alerts when an abnormality is detected" refers to a means that has the function of immediately notifying an administrator when an abnormality is detected based on collected and analyzed data.
[0599] "Generative AI models" are artificial intelligence algorithms and models used for natural language processing and information generation, examples of which include GPT-4.
[0600] The "means for displaying answers to questions via a user interface" refers to an interface and means including its functions for displaying answers generated by the generation engine to questions from users on the display screen of a user terminal.
[0601] System Configuration
[0602] This invention is a system consisting of a traveling device, a server, and a user terminal installed within a logistics facility. The traveling device collects video and audio data as it moves within the logistics facility and transmits it to the server. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal via a user interface.
[0603] Specific operation of the system
[0604] server
[0605] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including noise levels, vibrations, and lighting conditions within the logistics facility. It then analyzes the received data and extracts specific information to respond to user questions. For example, if a question is about noise levels, the server analyzes the audio data and evaluates noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device via the user interface.
[0606] The server uses the following specific technologies:
[0607] Hardware: Server machine
[0608] Software: Video analysis software (e.g., OpenCV), audio analysis software (e.g., DeepSpeech), natural language generation engines (e.g., GPT-4), real-time communication systems (e.g., WebSocket)
[0609] Specific examples
[0610] For example, consider the case where a manager asks about the noise level in a specific location in a logistics facility. Using a dedicated application, the manager inputs the question, "What is the current noise level?" The device sends this question to the server, which analyzes the collected noise data. The server evaluates the noise level based on the analysis results and generates an answer using a natural language generation engine, such as "The current noise level is an average of 85 decibels." This answer is sent to the device and displayed to the manager through a user interface. This allows the manager to obtain detailed information about the specific situation at the logistics facility.
[0611] User terminal
[0612] The terminal accepts questions from the user and sends them to the server. It receives the analysis results and generated answers from the server and displays them on the user interface. By checking these, the user can efficiently obtain detailed information about the logistics facility. The system also has a function that displays an automatic alert if an abnormality is detected.
[0613] User
[0614] Users access the system through a dedicated application. Using the application's interface, they can start monitoring the logistics facility and remotely control the traveling device. While monitoring, users can input questions about information of interest (e.g., "What is the lighting situation in this area?") and receive answers in real time. Based on the answers, users can evaluate and manage the logistics facility.
[0615] Prompt Sentence Examples
[0616] People: What is the noise level in your current distribution center?
[0617] AI: The current noise level averages 85 decibels, with peaks occurring between 2 and 3 p.m.
[0618] As described above, this system combines on-site traveling devices with a server and terminals to provide users with specific and detailed information about logistics facilities, allowing them to efficiently manage and evaluate the facilities.
[0619] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0620] Step 1:
[0621] The user launches the dedicated application and begins monitoring the logistics facility.
[0622] Input: The user launches the application
[0623] Processing: The user interface is launched and the running device is ready for remote operation.
[0624] Output: The control screen for the running gear is displayed.
[0625] Step 2:
[0626] The user remotely controls the traveling device, which patrols the logistics facility and collects video and audio data.
[0627] Input: User operates the running gear
[0628] Processing: As the traveling device moves, it collects video and audio using cameras and microphones.
[0629] Output: Real-time collected video and audio data
[0630] Step 3:
[0631] The traveling device transmits the collected video and audio data to a server.
[0632] Input: Video and audio data from the traveling device
[0633] Processing: Data is sent to the server via wireless communication
[0634] Output: Video and audio data received by the server
[0635] Step 4:
[0636] The server parses the received data.
[0637] Input: Video and audio data received by the server
[0638] Processing: The server analyzes the data using video analysis software (e.g., OpenCV) and audio analysis software (e.g., DeepSpeech).
[0639] Output: Analysis results (e.g. noise level, specific object recognition)
[0640] Step 5:
[0641] The user inputs a question via a dedicated application.
[0642] Input: User types a question into the device (e.g., "What is the current noise level?")
[0643] Process: The question is sent to the server
[0644] Output: The question sent to the server
[0645] Step 6:
[0646] Based on the analysis results, the server generates an answer to the question using a generative AI model (e.g., GPT-4).
[0647] Input: The question and analysis results sent to the server
[0648] Processing: The generative AI model generates a natural language answer to the question.
[0649] Output: The generated answer (e.g., "The current noise level is 85 decibels")
[0650] Step 7:
[0651] The generated answer is displayed on the user terminal via a user interface.
[0652] Input: Generated Answer
[0653] Processing: The answer is sent to the terminal and displayed in the user interface.
[0654] Output: The answer displayed in the user interface
[0655] Step 8:
[0656] The server will automatically send an alert if an abnormality is detected.
[0657] Input: Analysis results (e.g., noise level exceeds threshold)
[0658] Action: The server generates an automatic alert and notifies the user interface.
[0659] Output: The alert displayed in the user interface
[0660] In this way, the entire system can monitor detailed conditions within the logistics facility in real time, enabling efficient and safe operation and management.
[0661] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0662] System Configuration
[0663] The present invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. The traveling device collects video, audio, and environmental data as it moves within the property. The server receives this data and analyzes it based on the user's questions. The emotion engine then recognizes emotions from the user's audio and facial data, and displays the analysis results and generated answers on the user terminal via a user interface.
[0664] System Operation
[0665] server
[0666] The server first receives video, audio, and environmental data transmitted from the traveling device in real time. This includes data on noise levels, vibrations, and sunlight. The server then temporarily stores the received data and analyzes it based on the user's questions. For example, video data is used for questions about sunlight, and audio data is analyzed for questions about noise. Specific information is extracted and a specific answer is generated using a generation engine. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and provides feedback according to the user's emotional state.
[0667] Terminal
[0668] The device accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. Furthermore, the device also displays the analysis results of the emotion engine, adjusting the information provided and questions asked during the viewing as necessary.
[0669] User
[0670] Users access the system through a dedicated application. Using the application's interface, they can begin viewing properties and remotely control the vehicle. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[0671] Specific examples
[0672] For example, consider the case where a user asks about sunlight in the living room. The user inputs the question, "How sunny is the living room?" The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer generated is, "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface.
[0673] If a user shows a surprised expression during a viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, and appropriate feedback will be provided according to the user's emotion.
[0674] As described above, the present invention combines a traveling device placed within a property with a server, a terminal, and an emotion engine to provide specific and detailed property information and personalized feedback to users, allowing them to efficiently evaluate properties and choose the most suitable home.
[0675] The processing flow will be explained below.
[0676] Step 1:
[0677] server
[0678] The server receives the activation signal from the vehicle and activates it within the property. The vehicle then activates its camera and microphone to begin collecting video and audio data in real time. It also collects environmental data such as noise levels, vibrations, and sunlight.
[0679] Step 2:
[0680] server
[0681] The server receives video, audio and environmental data transmitted from the traveling device in real time and temporarily stores this data.
[0682] Step 3:
[0683] User
[0684] The user launches the dedicated application and begins viewing the property, entering a specific question (e.g., "How sunny is the living room?") through the application's interface.
[0685] Step 4:
[0686] Terminal
[0687] The terminal receives the user's question and transmits the question to the server.
[0688] Step 5:
[0689] server
[0690] The server analyzes the received question, identifies the content of the question, and selects and analyzes data related to the question, such as video data and audio data related to sunlight.
[0691] Step 6:
[0692] server
[0693] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0694] Step 7:
[0695] server
[0696] The server generates a response and sends it to the terminal.
[0697] Step 8:
[0698] Terminal
[0699] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[0700] Step 9:
[0701] User
[0702] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[0703] Step 10:
[0704] server
[0705] The server collects the user's voice and facial data and analyzes it with an emotion engine to recognize the user's emotional state (e.g., joy, sadness, surprise, anger).
[0706] Step 11:
[0707] server
[0708] Based on the analysis results of the emotion engine, feedback and additional information are generated according to the user's emotional state. For example, if the user is surprised, a message such as "Would you like to know more about the part that surprised you?" is generated.
[0709] Step 12:
[0710] server
[0711] Feedback and additional information from the emotion engine is sent to the device.
[0712] Step 13:
[0713] Terminal
[0714] The device displays feedback and additional information in the user interface according to the user's emotions, allowing the user to view the property in more detail.
[0715] In this way, the program proceeds step by step, providing the user with specific and detailed property information and emotional feedback in real time.
[0716] Example 2
[0717] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0718] Conventional property viewing systems are limited to collecting and analyzing video and audio data, and are unable to provide feedback tailored to the user's emotional state. Furthermore, they are limited in generating answers when users input specific questions, and lack the ability to provide personalized information in real time. This makes it difficult for users to efficiently obtain more accurate and detailed property information.
[0719] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0720] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation AI model means for generating answers to specific questions based on the analysis results, means for recognizing emotions from the user's voice data and face data, means for generating feedback based on the emotion recognition results and the analysis results, and means for displaying the generated answers and feedback via a user interface. This allows the user to obtain specific and detailed property information in real time and receive personalized feedback according to their emotions.
[0721] A "traveling device" is a device that collects video data, audio data, and environmental data while moving around the property.
[0722] "Video data" refers to data of moving images and still images taken by the traveling device.
[0723] "Audio data" is data of sounds collected by the traveling device.
[0724] "Environmental data" refers to data that includes information such as noise levels, vibrations, and sunlight levels within the property.
[0725] A "generative AI model" is an artificial intelligence model that analyzes received data and generates answers to specific questions.
[0726] "Emotion recognition" is the process of analyzing a user's voice and facial data to identify the user's emotional state.
[0727] A "user interface" is a screen or operating means that allows a user to exchange information with a system.
[0728] "Feedback" refers to responses or information provided to the user based on the analysis results and emotion recognition results.
[0729] "Real-time" refers to immediate processing or response with little or no delay.
[0730] "Analysis" is the process of extracting information and deriving meaning from received data.
[0731] The present invention is a property viewing system that includes a traveling device, a server, a user terminal, and an emotion engine. The traveling device moves around the property and collects video, audio, and environmental data. The server receives this data in real time and analyzes it based on the user's questions. Furthermore, the emotion engine recognizes emotions from the user's voice data and facial data, and displays the analysis results and generated answers on the user terminal through a user interface.
[0732] Hardware and software used
[0733] Hardware:
[0734] Mobile devices (e.g., robotic cameras, mobile sensor devices)
[0735] Servers (e.g., high-performance servers in data centers)
[0736] User device (e.g. smartphone, tablet)
[0737] Hardware for emotion engine (e.g. high-resolution camera, directional microphone)
[0738] software:
[0739] Data collection software (e.g. camera control software, sensor interface)
[0740] Data analysis software (e.g., image analysis algorithms, audio analysis algorithms)
[0741] Emotion recognition software (e.g., voice emotion recognition model, facial expression analysis model)
[0742] User interfaces (e.g., mobile apps, web applications)
[0743] Specific operation of the system
[0744] The server first receives video and audio data and environmental data (such as noise levels, vibrations, and sunlight) sent from the driving device in real time. This data is temporarily stored and analyzed based on questions from the user. Specific information is extracted using various algorithms for analysis, and specific answers are generated using generative AI models. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and generates feedback according to the user's emotional state.
[0745] The device accepts questions from the user as input and sends them to the server. The analysis results and generated answers are received from the server in real time and displayed through the user interface. Furthermore, the emotion engine's analysis results are also displayed, and the information provided during the viewing and questions can be adjusted as needed.
[0746] Users access the system through a dedicated application. Using the application's interface, they can start viewing properties and remotely control the traveling device. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[0747] Specific use cases
[0748] For example, consider the case where a user asks a question about the sunlight in the living room. The user enters "How sunny is the living room?" into the device. The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer "The living room gets good sunlight in the morning and is partially shaded in the afternoon" is generated. This answer is sent to the device and displayed to the user through the user interface.
[0749] Furthermore, if the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, providing appropriate feedback according to the user's emotion.
[0750] Examples of prompt statements
[0751] An example of a prompt is as follows:
[0752] prompt:
[0753] How's the sunlight in the living room?
[0754] Expected answer:
[0755] "The living room gets good sunlight in the morning, but some shading is expected in the afternoon."
[0756] As described above, the present invention combines a traveling device, a server, a user terminal, and an emotion engine to provide users with specific and detailed property information and personalized feedback, allowing them to efficiently evaluate properties and select the most suitable home.
[0757] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0758] Program processing flow
[0759] Step 1: Data collection
[0760] Processing Description:
[0761] The vehicle travels around the property, collecting video and audio data using cameras and microphones, and environmental sensors to collect environmental data such as noise levels, vibrations, and sunlight.
[0762] input:
[0763] Video, audio, noise level, vibration, sunlight
[0764] output:
[0765] Real-time transmission of video data, audio data, and environmental data
[0766] Specific operation:
[0767] As the traveling device moves around the living room, it takes pictures of the entire room with a camera, collects sound with a microphone, and measures the temperature with a sensor.
[0768] Step 2: Data reception and storage
[0769] Processing Description:
[0770] The server receives video data, audio data, and environmental data transmitted from the traveling device in real time and temporarily stores it.
[0771] input:
[0772] Data transmitted from the traveling device (video, audio, environmental data)
[0773] output:
[0774] Data stored in the database
[0775] Specific operation:
[0776] The server stores the received data in a dedicated database and prepares it for analysis.
[0777] Step 3: Ask a question
[0778] Processing Description:
[0779] The terminal accepts a question from the user as input and sends it to the server.
[0780] input:
[0781] User questions (e.g., "How sunny is the living room?")
[0782] output:
[0783] Send the question to the server
[0784] Specific operation:
[0785] The user enters a question in the application interface and presses the submit button, which sends the question to the server via the API.
[0786] Step 4: Data analysis
[0787] Processing Description:
[0788] The server analyzes the question, filters and extracts relevant video, audio, and environmental data, and uses a generative AI model to generate a specific answer.
[0789] input:
[0790] User questions, saved data (video, audio, environmental data)
[0791] output:
[0792] Generated Answer
[0793] Specific operation:
[0794] The server analyzes the question "How much sunlight does the living room get?", filters and analyzes the video data of the living room, and uses a generative AI model to generate the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon."
[0795] Step 5: Emotion Recognition
[0796] Processing Description:
[0797] The emotion engine recognizes emotions from the user's voice data and facial data, and sends the recognition results to the server, which then analyzes the emotion data and generates feedback.
[0798] input:
[0799] User voice data, face data
[0800] output:
[0801] Emotion recognition results and generated feedback
[0802] Specific operation:
[0803] The emotion engine monitors the user's facial expressions to detect smiles and surprises. For example, if the user shows a surprised expression, the emotion engine recognizes the surprise and generates feedback such as, "Would you like to provide additional information about what is surprising?"
[0804] Step 6: Submit and view your answers and feedback
[0805] Processing Description:
[0806] The server transmits the generated answers and feedback to the user terminal, which displays the received information on a user interface.
[0807] input:
[0808] Generated answers, generated feedback
[0809] output:
[0810] Data sent to user terminals, display data
[0811] Specific operation:
[0812] The server sends the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon" and feedback "Do you like this property?" to the user's device. The device receives this information and displays it on the application screen.
[0813] In this way, the entire system works together to consistently collect and analyze data, provide real-time answers to user questions, and even provide emotional feedback.
[0814] (Application example 2)
[0815] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0816] Conventional property viewing systems required users to personally visit the property, which was time-consuming and costly, and the information obtained from a single viewing was limited. Furthermore, it was difficult to provide personalized information based on the user's emotions and personal preferences, resulting in insufficient evaluations of properties depending on the user. Furthermore, due to a lack of analysis of environmental data and appropriate feedback, users were sometimes unaware of potential problems with the property. Thus, to improve user convenience and satisfaction, a more detailed and individually tailored property information system was needed.
[0817] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0818] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, and emotion recognition means for recognizing the user's emotions and providing feedback according to the user's emotional state. This allows the user to remotely view the property and receive answers in real time, and also allows the user to be provided with information according to their individual emotions, making it possible to efficiently obtain more detailed and personalized property information.
[0819] A "moving device" is a device that moves autonomously or remotely within a property to collect video and audio data.
[0820] The "generation engine means" is an engine that has the function of generating an answer to a specific question from a user based on the analysis results.
[0821] A "user interface" is an interface for displaying generated answers and analysis results to the user, and is a means for exchanging information between the user and the system.
[0822] The "emotion recognition means" is a means that has the function of recognizing emotions from data such as the user's voice and facial expression, and providing feedback according to those emotions.
[0823] "Environmental data" refers to data relating to environmental conditions such as noise levels, vibrations, and sunlight levels within the property.
[0824] A "visualization device" is a device worn by a user that collects video and audio data and provides the user with visual and audio information in real time.
[0825] A "biological signal" is a signal that contains information about the user's health condition and fatigue level, and includes heart rate, body temperature, skin potential, and the like.
[0826] "Personalized information" refers to information and feedback that is customized based on a user's individual emotional and health state.
[0827] This invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. Specific embodiments of this system are described below.
[0828] System Configuration
[0829] The server has a means for receiving video and audio data captured by the traveling device installed within the property. The received data is processed using an analysis engine to generate an optimal answer based on the user's question. The generated answer is displayed on the user's terminal via a user interface. In addition, an emotion recognition means is used to analyze the user's emotional state and provide feedback as necessary.
[0830] Hardware and Software
[0831] The system uses the following hardware and software:
[0832] Mobile device: Responsible for moving around the property and collecting video and audio data.
[0833] Server: Receives and analyzes data, generates answers via a generation engine, and recognizes emotions.
[0834] User terminal (smartphone, smart glasses, head-mounted display, etc.): displays information through a user interface.
[0835] Emotion engine: Recognizes emotions from the user's facial data and voice.
[0836] Software used: OpenCV (video data processing), EmotionRecognizer (emotion recognition), Server (communication with server).
[0837] Data processing and calculation
[0838] The server receives video and audio data from the vehicle in real time and temporarily stores it. The stored data is analyzed to extract specific information in response to user questions. A generative AI model is used to generate specific answers, which are then displayed to the user through a user interface.
[0839] The emotion engine analyzes the user's voice and facial data to recognize their emotions. The recognized emotions are sent to the server, which provides appropriate feedback along with the analysis results.
[0840] Specific examples
[0841] For example, if a user asks about the noise level in a room, the user inputs the question, "What is the noise level in this room?" The device sends this question to the server, which analyzes the voice data. As a result of the analysis, the answer "The average noise level in this room is 50 decibels" is generated and sent to the device. This answer is then displayed to the user through the user interface.
[0842] If the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and provide appropriate feedback according to the user's emotions, such as displaying a message like, "Would you like to provide additional information about the part that surprised you?"
[0843] Prompt Sentence Examples
[0844] "When you inquire about a new product, please send us the data of the moment that surprised you the most."
[0845] "Analyze user interests using facial recognition data and suggest appropriate products."
[0846] This system allows users to efficiently obtain more detailed and personalized information, and also provides feedback according to the user's emotions, thereby improving user satisfaction.
[0847] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0848] Step 1:
[0849] The device accepts questions from users as input. The user inputs the question through an interface such as smart glasses or a smartphone. The input data is sent to the server. The input content may include, for example, "What is the noise level in this room?"
[0850] Step 2:
[0851] The server receives the query data sent from the device. Based on the received query data, it acquires and temporarily stores the video and audio data collected from the traveling device. The temporarily stored data includes video of the room, environmental sounds, vibration data, and sunlight amount.
[0852] Step 3:
[0853] The server analyzes the temporarily stored data. Specifically, if a question is about noise levels, for example, the server analyzes the audio data. This analysis analyzes the noise level and frequency components, and outputs the noise level in decibels (dB). The analyzed data is then converted into a specific answer by a generation engine.
[0854] Step 4:
[0855] The generation engine uses the analysis results to create a generated answer, which may contain specific information such as "The average noise level in this room is 50 decibels," and sends the generated answer to the user interface.
[0856] Step 5:
[0857] The terminal receives the generated answer data sent from the server and displays it to the user via a user interface, and if the user is using smart glasses, the answer is presented visually.
[0858] Step 6:
[0859] The emotion recognition means analyzes the user's voice data and face data. The emotion engine analyzes the user's facial expression and tone of voice when they input a question, and recognizes the user's emotional state (for example, surprise, satisfaction, dissatisfaction, etc.).
[0860] Step 7:
[0861] The server analyzes the emotion data sent from the emotion engine and provides additional feedback as needed. For example, if the user shows a surprised expression, a message such as "Would you like to provide additional information about why you are surprised?" is generated and sent to the device.
[0862] Step 8:
[0863] The terminal receives the emotion-based feedback data transmitted from the server and displays it to the user via a user interface, allowing the user to receive information corresponding to their emotions in real time.
[0864] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0865] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0866] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0867] [Third embodiment]
[0868] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0869] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0870] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0871] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0872] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0873] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0874] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0875] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0876] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0877] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0878] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0879] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0880] System Configuration
[0881] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves around the property. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal through a user interface.
[0882] System Operation
[0883] server
[0884] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including the property's noise level, vibration, and sunlight. It then analyzes the received data and extracts specific information to respond to the user's question. For example, if the question is about sunlight, it evaluates the changes in light during the day based on the video data. If the question is about noise levels, it analyzes the audio data and evaluates the noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device.
[0885] Terminal
[0886] The terminal accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. By checking these, the user can efficiently obtain detailed property information.
[0887] User
[0888] Users access the system through a dedicated application. Using the application's interface, they can begin viewing the property and remotely control the traveling device. During the viewing, users can enter questions about information they are interested in (e.g., "How sunny is the bedroom?") and receive answers in real time. Based on the answers, users can proceed with their property evaluation.
[0889] Specific examples
[0890] For example, consider the case where a user asks about sunlight in the living room. Using a dedicated application, the user inputs the question, "How much sunlight does the living room get?" The device sends this question to the server, which analyzes the video data of the living room. The server evaluates the sunlight in the living room from the analysis results and generates an answer using a natural language generation engine, such as "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface. This allows the user to obtain detailed information about the sunlight in the living room without actually visiting the property.
[0891] As described above, the present invention provides specific and detailed property information to users by combining a traveling device, server, and terminal installed within a property, allowing users to efficiently evaluate properties and select the most suitable home.
[0892] The processing flow will be explained below.
[0893] Step 1:
[0894] server
[0895] The server receives the activation signal from the traveling device and activates it within the property. The traveling device activates the camera and microphone and begins collecting video and audio data in real time.
[0896] Step 2:
[0897] server
[0898] The server receives video and audio data transmitted from the traveling device in real time, as well as environmental data such as noise levels, vibrations, and sunlight. The received data is temporarily stored.
[0899] Step 3:
[0900] User
[0901] The user launches the dedicated application and begins viewing the property. The user inputs a specific question (e.g., "How sunny is the living room?") through the interface.
[0902] Step 4:
[0903] Terminal
[0904] The terminal receives the user's question and transmits the question to the server.
[0905] Step 5:
[0906] server
[0907] The server analyzes the received question, identifies the question, and selects and analyzes relevant data (e.g., video data on sunlight) based on the identified question.
[0908] Step 6:
[0909] server
[0910] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0911] Step 7:
[0912] server
[0913] The server generates a response and sends it to the terminal.
[0914] Step 8:
[0915] Terminal
[0916] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[0917] Step 9:
[0918] User
[0919] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[0920] In this way, the program proceeds step by step, providing the user with specific and detailed property information in real time.
[0921] Example 1
[0922] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0923] With conventional property viewing systems, it was difficult for users to obtain detailed information about the property in real time, resulting in low accuracy in property evaluations. Furthermore, there was a lack of efficient means for collecting and analyzing environmental data such as noise levels and sunlight, making it impossible for users to obtain the detailed answers they desired. This made it difficult for users to efficiently view properties and select the most suitable home.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0925] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, means for receiving and analyzing environmental data including noise levels, vibrations, and sunlight levels within the property, means for transmitting questions input by users to the server, and means for displaying the analysis results received from the server and the generated answers on the user interface, thereby enabling users to obtain detailed information about the property in real time and efficiently evaluate the property.
[0926] A "traveling device" is a device that collects video and audio data while moving around the property.
[0927] The "server" is the core computer system that receives and analyzes data sent from the traveling device and generates answers to user questions.
[0928] A "user terminal" is a device that allows a user to access the system using a dedicated application, enter questions, and check answers.
[0929] "Video data" refers to data containing visual information about the inside of a property, and is captured by a traveling device.
[0930] "Audio data" refers to data containing auditory information within a property, and is recorded by the traveling device.
[0931] "Environmental data" refers to data that includes information about the environment within the property, such as noise levels, vibrations, and sunlight.
[0932] "Analysis" is the process of extracting information from the received data and identifying information that responds to the user's question.
[0933] The "generation engine" is an engine for generating specific answers to user questions based on the analysis results.
[0934] A "generative AI model" is an artificial intelligence model that generates natural language responses based on input prompts.
[0935] A "prompt sentence" is a sentence that is input into a generative AI model and serves as the basis for the answer to a question that is generated based on the analysis results.
[0936] "User interface" refers to an interface that allows a user to interact with a system, and includes the screen displayed on the terminal and the operation method.
[0937] MODE FOR CARRYING OUT THE INVENTION
[0938] System configuration
[0939] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves within the property and transmits it to the server. The server analyzes the received data and generates answers to users' questions. The generated answers are displayed on the user terminal through a user interface. Users can access the system via a dedicated application and input questions.
[0940] Hardware and software used
[0941] Mobile device: A device that moves autonomously or remotely around a property and collects video and audio data using cameras and microphones.
[0942] Server: A high-performance computer that provides an environment for receiving data, analyzing it, and running generative AI models. For example, it can use cloud-based computing services.
[0943] User device: A device such as a smartphone, tablet, or PC with a dedicated application installed.
[0944] Data processing and calculation
[0945] 1. Data reception:
[0946] The server receives real-time video and audio data transmitted from the traveling device, as well as environmental data such as noise levels, vibrations, and sunlight levels at the property.
[0947] 2. Data Analysis:
[0948] The server analyzes the received data and extracts specific information that responds to the user's question, for example by implementing algorithms that analyze video data to evaluate changes in light and shadow movement during the day.
[0949] 3. Answer generation:
[0950] Based on the analysis results, the generative AI model creates a prompt sentence and generates a natural language answer based on that prompt sentence. For example, the prompt sentence could be, "The user is asking about the sunlight in the living room. Please evaluate the sunlight in the living room based on the video data of the living room and generate a natural language answer explaining the results."
[0951] 4. Answer display:
[0952] The generated answer is sent to the user terminal through the user interface and displayed on the application screen.
[0953] Specific examples
[0954] For example, if a user asks about sunlight in the living room, the following process is performed.
[0955] 1. User enters a question: The user uses a dedicated application to enter a question such as, "How much sunlight does the living room get?"
[0956] 2. The device sends the question to the server: The device sends this question to the server.
[0957] 3. The server analyzes the data: The server analyzes the video data from the living room and evaluates the amount of sunlight.
[0958] 4. Answer generation: Based on the prompt, the server's generative AI model generates the answer, "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[0959] 5. Displaying the answer: This answer is sent to the terminal and displayed to the user through the user interface.
[0960] Prompt Sentence Examples
[0961] "The user is asking about the sunlight in the living room. Based on the video data of the living room, please evaluate the sunlight in the living room and generate a natural language answer explaining the results."
[0962] As described above, this system efficiently provides users with specific and detailed property information and helps them evaluate properties, enabling them to obtain information to select the most suitable home.
[0963] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0964] Step 1: User enters question
[0965] The user launches the dedicated application and inputs a question through the interface, such as "How much sunlight does the living room get?" The application temporarily stores the question in memory.
[0966] Input: Questions entered into the user interface
[0967] Output: Question data stored in memory
[0968] Specific behavior:
[0969] The user starts the application and a question input screen is displayed.
[0970] Enter your question and tap the "Submit" button.
[0971] Step 2: Sending a question to the server from the device
[0972] The device uses a dedicated communication protocol to send the question data to the server, such as an HTTP POST request.
[0973] Input: Question data stored in memory
[0974] Output: Question data sent to the server
[0975] Specific behavior:
[0976] The device retrieves the question data stored in memory and sends it to the server as an HTTP POST request.
[0977] Step 3: Server receives data
[0978] The server receives the query data sent from the terminal. It then begins receiving video and audio data sent from the traveling device in real time. It also simultaneously receives environmental data such as noise levels, vibrations, and sunlight.
[0979] Input: Question data sent from the terminal, video and audio data from the driving device, environmental data
[0980] Output: Question data and various data stored on the server
[0981] Specific behavior:
[0982] The server receives the query data and stores it in a database.
[0983] The server receives data from the running gear and environmental sensors in real time and prepares it for analysis.
[0984] Step 4: Data analysis by the server
[0985] The server analyzes the data it receives and extracts information to answer questions, such as analyzing video data to assess changes in light during the day or audio data to assess noise levels.
[0986] Input: Video, audio, environmental data, and question data stored on the server
[0987] Output: Specific analysis results that address your questions
[0988] Specific behavior:
[0989] An algorithm is run to detect changes in light from the video data.
[0990] Audio data is analyzed to evaluate noise levels by time of day.
[0991] Step 5: Server Generates Answer
[0992] Based on the analysis results, the server creates a prompt to be applied to the generative AI model. The prompt is input into the generative AI model, which generates a natural language answer. For example, it generates an answer such as, "The living room has good sunlight in the morning and is partially shaded in the afternoon."
[0993] Input: Analysis results, generative AI model, prompt
[0994] Output: Natural language response
[0995] Specific behavior:
[0996] The server creates a prompt based on the analysis results.
[0997] The prompt sentence is input into a generative AI model to generate a natural language response.
[0998] Step 6: Sending the response from the server to the device
[0999] The server sends the generated response to the user's terminal using a communication protocol such as an HTTP POST request.
[1000] Input: Natural language response data
[1001] Output: Answer data sent to the user's device
[1002] Specific behavior:
[1003] The server sends the generated response to the device as an HTTP POST request.
[1004] Step 7: Viewing the Answers on Your Device
[1005] The terminal stores the response data received from the server in a memory and displays it on a user interface.
[1006] Input: Response data received from the server
[1007] Output: The answer displayed in the user interface
[1008] Specific behavior:
[1009] The terminal analyzes the received response data and displays it on the user interface.
[1010] (Application example 1)
[1011] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1012] Management and monitoring of logistics facilities are important for improving efficiency and ensuring safety. However, conventional systems make it difficult to physically inspect the site and integrate and analyze a wide range of environmental data, placing a heavy burden on managers. Furthermore, there are challenges in responding quickly to abnormalities and providing detailed information in real time.
[1013] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1014] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within the property, means for analyzing the received video and audio data, and a generation engine means for generating answers to specific questions based on the analysis results. This enables means for receiving and analyzing environmental data for monitoring and management within the logistics facility, such as noise levels, vibrations, and lighting conditions, means for issuing automatic alerts when an abnormality is detected, means for allowing a manager of the logistics facility to input a question and generating a specific answer to the question using a generative AI model, and means for displaying the answer to the question via a user interface.
[1015] A "traveling device installed within the property" is a device that moves autonomously or remotely within the logistics facility and collects video and audio data.
[1016] The "means for receiving video and audio data" refers to a means having a function for transmitting video and audio data collected by the traveling device to a server in real time and for the server to receive the data.
[1017] "Means for analyzing video and audio data" refers to technical means for analyzing received video and audio data and extracting specific information, and examples include video analysis software and audio analysis software.
[1018] The "generation engine means" refers to means including algorithms and programs for generating specific answers to user questions based on analyzed data.
[1019] The "means for displaying via a user interface" refers to a means including an interface and its functions for displaying the generated answer on a user terminal.
[1020] "Means for receiving and analyzing environmental data" refers to the technical means for collecting and analyzing environmental data such as noise levels, vibrations, and lighting conditions within a logistics facility.
[1021] "Means for issuing automatic alerts when an abnormality is detected" refers to a means that has the function of immediately notifying an administrator when an abnormality is detected based on collected and analyzed data.
[1022] "Generative AI models" are artificial intelligence algorithms and models used for natural language processing and information generation, examples of which include GPT-4.
[1023] The "means for displaying answers to questions via a user interface" refers to an interface and means including its functions for displaying answers generated by the generation engine to questions from users on the display screen of a user terminal.
[1024] System Configuration
[1025] This invention is a system consisting of a traveling device, a server, and a user terminal installed within a logistics facility. The traveling device collects video and audio data as it moves within the logistics facility and transmits it to the server. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal via a user interface.
[1026] Specific operation of the system
[1027] server
[1028] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including noise levels, vibrations, and lighting conditions within the logistics facility. It then analyzes the received data and extracts specific information to respond to user questions. For example, if a question is about noise levels, the server analyzes the audio data and evaluates noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device via the user interface.
[1029] The server uses the following specific technologies:
[1030] Hardware: Server machine
[1031] Software: Video analysis software (e.g., OpenCV), audio analysis software (e.g., DeepSpeech), natural language generation engines (e.g., GPT-4), real-time communication systems (e.g., WebSocket)
[1032] Specific examples
[1033] For example, consider the case where a manager asks about the noise level in a specific location in a logistics facility. Using a dedicated application, the manager inputs the question, "What is the current noise level?" The device sends this question to the server, which analyzes the collected noise data. The server evaluates the noise level based on the analysis results and generates an answer using a natural language generation engine, such as "The current noise level is an average of 85 decibels." This answer is sent to the device and displayed to the manager through a user interface. This allows the manager to obtain detailed information about the specific situation at the logistics facility.
[1034] User terminal
[1035] The terminal accepts questions from the user and sends them to the server. It receives the analysis results and generated answers from the server and displays them on the user interface. By checking these, the user can efficiently obtain detailed information about the logistics facility. The system also has a function that displays an automatic alert if an abnormality is detected.
[1036] User
[1037] Users access the system through a dedicated application. Using the application's interface, they can start monitoring the logistics facility and remotely control the traveling device. While monitoring, users can input questions about information of interest (e.g., "What is the lighting situation in this area?") and receive answers in real time. Based on the answers, users can evaluate and manage the logistics facility.
[1038] Prompt Sentence Examples
[1039] People: What is the noise level in your current distribution center?
[1040] AI: The current noise level averages 85 decibels, with peaks occurring between 2 and 3 p.m.
[1041] As described above, this system combines on-site traveling devices with a server and terminals to provide users with specific and detailed information about logistics facilities, allowing them to efficiently manage and evaluate the facilities.
[1042] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1043] Step 1:
[1044] The user launches the dedicated application and begins monitoring the logistics facility.
[1045] Input: The user launches the application
[1046] Processing: The user interface is launched and the running device is ready for remote operation.
[1047] Output: The control screen for the running gear is displayed.
[1048] Step 2:
[1049] The user remotely controls the traveling device, which patrols the logistics facility and collects video and audio data.
[1050] Input: User operates the running gear
[1051] Processing: As the traveling device moves, it collects video and audio using cameras and microphones.
[1052] Output: Real-time collected video and audio data
[1053] Step 3:
[1054] The traveling device transmits the collected video and audio data to a server.
[1055] Input: Video and audio data from the traveling device
[1056] Processing: Data is sent to the server via wireless communication
[1057] Output: Video and audio data received by the server
[1058] Step 4:
[1059] The server parses the received data.
[1060] Input: Video and audio data received by the server
[1061] Processing: The server analyzes the data using video analysis software (e.g., OpenCV) and audio analysis software (e.g., DeepSpeech).
[1062] Output: Analysis results (e.g. noise level, specific object recognition)
[1063] Step 5:
[1064] The user inputs a question via a dedicated application.
[1065] Input: User types a question into the device (e.g., "What is the current noise level?")
[1066] Process: The question is sent to the server
[1067] Output: The question sent to the server
[1068] Step 6:
[1069] Based on the analysis results, the server generates an answer to the question using a generative AI model (e.g., GPT-4).
[1070] Input: The question and analysis results sent to the server
[1071] Processing: The generative AI model generates a natural language answer to the question.
[1072] Output: The generated answer (e.g., "The current noise level is 85 decibels")
[1073] Step 7:
[1074] The generated answer is displayed on the user terminal via a user interface.
[1075] Input: Generated Answer
[1076] Processing: The answer is sent to the terminal and displayed in the user interface.
[1077] Output: The answer displayed in the user interface
[1078] Step 8:
[1079] The server will automatically send an alert if an abnormality is detected.
[1080] Input: Analysis results (e.g., noise level exceeds threshold)
[1081] Action: The server generates an automatic alert and notifies the user interface.
[1082] Output: The alert displayed in the user interface
[1083] In this way, the entire system can monitor detailed conditions within the logistics facility in real time, enabling efficient and safe operation and management.
[1084] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1085] System Configuration
[1086] The present invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. The traveling device collects video, audio, and environmental data as it moves within the property. The server receives this data and analyzes it based on the user's questions. The emotion engine then recognizes emotions from the user's audio and facial data, and displays the analysis results and generated answers on the user terminal via a user interface.
[1087] System Operation
[1088] server
[1089] The server first receives video, audio, and environmental data transmitted from the traveling device in real time. This includes data on noise levels, vibrations, and sunlight. The server then temporarily stores the received data and analyzes it based on the user's questions. For example, video data is used for questions about sunlight, and audio data is analyzed for questions about noise. Specific information is extracted and a specific answer is generated using a generation engine. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and provides feedback according to the user's emotional state.
[1090] Terminal
[1091] The device accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. Furthermore, the device also displays the analysis results of the emotion engine, adjusting the information provided and questions asked during the viewing as necessary.
[1092] User
[1093] Users access the system through a dedicated application. Using the application's interface, they can begin viewing properties and remotely control the vehicle. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[1094] Specific examples
[1095] For example, consider the case where a user asks about sunlight in the living room. The user inputs the question, "How sunny is the living room?" The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer generated is, "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface.
[1096] If a user shows a surprised expression during a viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, and appropriate feedback will be provided according to the user's emotion.
[1097] As described above, the present invention combines a traveling device placed within a property with a server, a terminal, and an emotion engine to provide specific and detailed property information and personalized feedback to users, allowing them to efficiently evaluate properties and choose the most suitable home.
[1098] The processing flow will be explained below.
[1099] Step 1:
[1100] server
[1101] The server receives the activation signal from the vehicle and activates it within the property. The vehicle then activates its camera and microphone to begin collecting video and audio data in real time. It also collects environmental data such as noise levels, vibrations, and sunlight.
[1102] Step 2:
[1103] server
[1104] The server receives video, audio and environmental data transmitted from the traveling device in real time and temporarily stores this data.
[1105] Step 3:
[1106] User
[1107] The user launches the dedicated application and begins viewing the property, entering a specific question (e.g., "How sunny is the living room?") through the application's interface.
[1108] Step 4:
[1109] Terminal
[1110] The terminal receives the user's question and transmits the question to the server.
[1111] Step 5:
[1112] server
[1113] The server analyzes the received question, identifies the content of the question, and selects and analyzes data related to the question, such as video data and audio data related to sunlight.
[1114] Step 6:
[1115] server
[1116] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[1117] Step 7:
[1118] server
[1119] The server generates a response and sends it to the terminal.
[1120] Step 8:
[1121] Terminal
[1122] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[1123] Step 9:
[1124] User
[1125] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[1126] Step 10:
[1127] server
[1128] The server collects the user's voice and facial data and analyzes it with an emotion engine to recognize the user's emotional state (e.g., joy, sadness, surprise, anger).
[1129] Step 11:
[1130] server
[1131] Based on the analysis results of the emotion engine, feedback and additional information are generated according to the user's emotional state. For example, if the user is surprised, a message such as "Would you like to know more about the part that surprised you?" is generated.
[1132] Step 12:
[1133] server
[1134] Feedback and additional information from the emotion engine is sent to the device.
[1135] Step 13:
[1136] Terminal
[1137] The device displays feedback and additional information in the user interface according to the user's emotions, allowing the user to view the property in more detail.
[1138] In this way, the program proceeds step by step, providing the user with specific and detailed property information and emotional feedback in real time.
[1139] Example 2
[1140] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1141] Conventional property viewing systems are limited to collecting and analyzing video and audio data, and are unable to provide feedback tailored to the user's emotional state. Furthermore, they are limited in generating answers when users input specific questions, and lack the ability to provide personalized information in real time. This makes it difficult for users to efficiently obtain more accurate and detailed property information.
[1142] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1143] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation AI model means for generating answers to specific questions based on the analysis results, means for recognizing emotions from the user's voice data and face data, means for generating feedback based on the emotion recognition results and the analysis results, and means for displaying the generated answers and feedback via a user interface. This allows the user to obtain specific and detailed property information in real time and receive personalized feedback according to their emotions.
[1144] A "traveling device" is a device that collects video data, audio data, and environmental data while moving around the property.
[1145] "Video data" refers to data of moving images and still images taken by the traveling device.
[1146] "Audio data" is data of sounds collected by the traveling device.
[1147] "Environmental data" refers to data that includes information such as noise levels, vibrations, and sunlight levels within the property.
[1148] A "generative AI model" is an artificial intelligence model that analyzes received data and generates answers to specific questions.
[1149] "Emotion recognition" is the process of analyzing a user's voice and facial data to identify the user's emotional state.
[1150] A "user interface" is a screen or operating means that allows a user to exchange information with a system.
[1151] "Feedback" refers to responses or information provided to the user based on the analysis results and emotion recognition results.
[1152] "Real-time" refers to immediate processing or response with little or no delay.
[1153] "Analysis" is the process of extracting information and deriving meaning from received data.
[1154] The present invention is a property viewing system that includes a traveling device, a server, a user terminal, and an emotion engine. The traveling device moves around the property and collects video, audio, and environmental data. The server receives this data in real time and analyzes it based on the user's questions. Furthermore, the emotion engine recognizes emotions from the user's voice data and facial data, and displays the analysis results and generated answers on the user terminal through a user interface.
[1155] Hardware and software used
[1156] Hardware:
[1157] Mobile devices (e.g., robotic cameras, mobile sensor devices)
[1158] Servers (e.g., high-performance servers in data centers)
[1159] User device (e.g. smartphone, tablet)
[1160] Hardware for emotion engine (e.g. high-resolution camera, directional microphone)
[1161] software:
[1162] Data collection software (e.g. camera control software, sensor interface)
[1163] Data analysis software (e.g., image analysis algorithms, audio analysis algorithms)
[1164] Emotion recognition software (e.g., voice emotion recognition model, facial expression analysis model)
[1165] User interfaces (e.g., mobile apps, web applications)
[1166] Specific operation of the system
[1167] The server first receives video and audio data and environmental data (such as noise levels, vibrations, and sunlight) sent from the driving device in real time. This data is temporarily stored and analyzed based on questions from the user. Specific information is extracted using various algorithms for analysis, and specific answers are generated using generative AI models. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and generates feedback according to the user's emotional state.
[1168] The device accepts questions from the user as input and sends them to the server. The analysis results and generated answers are received from the server in real time and displayed through the user interface. Furthermore, the emotion engine's analysis results are also displayed, and the information provided during the viewing and questions can be adjusted as needed.
[1169] Users access the system through a dedicated application. Using the application's interface, they can start viewing properties and remotely control the traveling device. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[1170] Specific use cases
[1171] For example, consider the case where a user asks a question about the sunlight in the living room. The user enters "How sunny is the living room?" into the device. The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer "The living room gets good sunlight in the morning and is partially shaded in the afternoon" is generated. This answer is sent to the device and displayed to the user through the user interface.
[1172] Furthermore, if the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, providing appropriate feedback according to the user's emotion.
[1173] Examples of prompt statements
[1174] An example of a prompt is as follows:
[1175] prompt:
[1176] How's the sunlight in the living room?
[1177] Expected answer:
[1178] "The living room gets good sunlight in the morning, but some shading is expected in the afternoon."
[1179] As described above, the present invention combines a traveling device, a server, a user terminal, and an emotion engine to provide users with specific and detailed property information and personalized feedback, allowing them to efficiently evaluate properties and select the most suitable home.
[1180] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1181] Program processing flow
[1182] Step 1: Data collection
[1183] Processing Description:
[1184] The vehicle travels around the property, collecting video and audio data using cameras and microphones, and environmental sensors to collect environmental data such as noise levels, vibrations, and sunlight.
[1185] input:
[1186] Video, audio, noise level, vibration, sunlight
[1187] output:
[1188] Real-time transmission of video data, audio data, and environmental data
[1189] Specific operation:
[1190] As the traveling device moves around the living room, it takes pictures of the entire room with a camera, collects sound with a microphone, and measures the temperature with a sensor.
[1191] Step 2: Data reception and storage
[1192] Processing Description:
[1193] The server receives video data, audio data, and environmental data transmitted from the traveling device in real time and temporarily stores it.
[1194] input:
[1195] Data transmitted from the traveling device (video, audio, environmental data)
[1196] output:
[1197] Data stored in the database
[1198] Specific operation:
[1199] The server stores the received data in a dedicated database and prepares it for analysis.
[1200] Step 3: Ask a question
[1201] Processing Description:
[1202] The terminal accepts a question from the user as input and sends it to the server.
[1203] input:
[1204] User questions (e.g., "How sunny is the living room?")
[1205] output:
[1206] Send the question to the server
[1207] Specific operation:
[1208] The user enters a question in the application interface and presses the submit button, which sends the question to the server via the API.
[1209] Step 4: Data analysis
[1210] Processing Description:
[1211] The server analyzes the question, filters and extracts relevant video, audio, and environmental data, and uses a generative AI model to generate a specific answer.
[1212] input:
[1213] User questions, saved data (video, audio, environmental data)
[1214] output:
[1215] Generated Answer
[1216] Specific operation:
[1217] The server analyzes the question "How much sunlight does the living room get?", filters and analyzes the video data of the living room, and uses a generative AI model to generate the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon."
[1218] Step 5: Emotion Recognition
[1219] Processing Description:
[1220] The emotion engine recognizes emotions from the user's voice data and facial data, and sends the recognition results to the server, which then analyzes the emotion data and generates feedback.
[1221] input:
[1222] User voice data, face data
[1223] output:
[1224] Emotion recognition results and generated feedback
[1225] Specific operation:
[1226] The emotion engine monitors the user's facial expressions to detect smiles and surprises. For example, if the user shows a surprised expression, the emotion engine recognizes the surprise and generates feedback such as, "Would you like to provide additional information about what is surprising?"
[1227] Step 6: Submit and view your answers and feedback
[1228] Processing Description:
[1229] The server transmits the generated answers and feedback to the user terminal, which displays the received information on a user interface.
[1230] input:
[1231] Generated answers, generated feedback
[1232] output:
[1233] Data sent to user terminals, display data
[1234] Specific operation:
[1235] The server sends the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon" and feedback "Do you like this property?" to the user's device. The device receives this information and displays it on the application screen.
[1236] In this way, the entire system works together to consistently collect and analyze data, provide real-time answers to user questions, and even provide emotional feedback.
[1237] (Application example 2)
[1238] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1239] Conventional property viewing systems required users to personally visit the property, which was time-consuming and costly, and the information obtained from a single viewing was limited. Furthermore, it was difficult to provide personalized information based on the user's emotions and personal preferences, resulting in insufficient evaluations of properties depending on the user. Furthermore, due to a lack of analysis of environmental data and appropriate feedback, users were sometimes unaware of potential problems with the property. Thus, to improve user convenience and satisfaction, a more detailed and individually tailored property information system was needed.
[1240] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1241] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, and emotion recognition means for recognizing the user's emotions and providing feedback according to the user's emotional state. This allows the user to remotely view the property and receive answers in real time, and also allows the user to be provided with information according to their individual emotions, making it possible to efficiently obtain more detailed and personalized property information.
[1242] A "moving device" is a device that moves autonomously or remotely within a property to collect video and audio data.
[1243] The "generation engine means" is an engine that has the function of generating an answer to a specific question from a user based on the analysis results.
[1244] A "user interface" is an interface for displaying generated answers and analysis results to the user, and is a means for exchanging information between the user and the system.
[1245] The "emotion recognition means" is a means that has the function of recognizing emotions from data such as the user's voice and facial expression, and providing feedback according to those emotions.
[1246] "Environmental data" refers to data relating to environmental conditions such as noise levels, vibrations, and sunlight levels within the property.
[1247] A "visualization device" is a device worn by a user that collects video and audio data and provides the user with visual and audio information in real time.
[1248] A "biological signal" is a signal that contains information about the user's health condition and fatigue level, and includes heart rate, body temperature, skin potential, and the like.
[1249] "Personalized information" refers to information and feedback that is customized based on a user's individual emotional and health state.
[1250] This invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. Specific embodiments of this system are described below.
[1251] System Configuration
[1252] The server has a means for receiving video and audio data captured by the traveling device installed within the property. The received data is processed using an analysis engine to generate an optimal answer based on the user's question. The generated answer is displayed on the user's terminal via a user interface. In addition, an emotion recognition means is used to analyze the user's emotional state and provide feedback as necessary.
[1253] Hardware and Software
[1254] The system uses the following hardware and software:
[1255] Mobile device: Responsible for moving around the property and collecting video and audio data.
[1256] Server: Receives and analyzes data, generates answers via a generation engine, and recognizes emotions.
[1257] User terminal (smartphone, smart glasses, head-mounted display, etc.): displays information through a user interface.
[1258] Emotion engine: Recognizes emotions from the user's facial data and voice.
[1259] Software used: OpenCV (video data processing), EmotionRecognizer (emotion recognition), Server (communication with server).
[1260] Data processing and calculation
[1261] The server receives video and audio data from the vehicle in real time and temporarily stores it. The stored data is analyzed to extract specific information in response to user questions. A generative AI model is used to generate specific answers, which are then displayed to the user through a user interface.
[1262] The emotion engine analyzes the user's voice and facial data to recognize their emotions. The recognized emotions are sent to the server, which provides appropriate feedback along with the analysis results.
[1263] Specific examples
[1264] For example, if a user asks about the noise level in a room, the user inputs the question, "What is the noise level in this room?" The device sends this question to the server, which analyzes the voice data. As a result of the analysis, the answer "The average noise level in this room is 50 decibels" is generated and sent to the device. This answer is then displayed to the user through the user interface.
[1265] If the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and provide appropriate feedback according to the user's emotions, such as displaying a message like, "Would you like to provide additional information about the part that surprised you?"
[1266] Prompt Sentence Examples
[1267] "When you inquire about a new product, please send us the data of the moment that surprised you the most."
[1268] "Analyze user interests using facial recognition data and suggest appropriate products."
[1269] This system allows users to efficiently obtain more detailed and personalized information, and also provides feedback according to the user's emotions, thereby improving user satisfaction.
[1270] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1271] Step 1:
[1272] The device accepts questions from users as input. The user inputs the question through an interface such as smart glasses or a smartphone. The input data is sent to the server. The input content may include, for example, "What is the noise level in this room?"
[1273] Step 2:
[1274] The server receives the query data sent from the device. Based on the received query data, it acquires and temporarily stores the video and audio data collected from the traveling device. The temporarily stored data includes video of the room, environmental sounds, vibration data, and sunlight amount.
[1275] Step 3:
[1276] The server analyzes the temporarily stored data. Specifically, if a question is about noise levels, for example, the server analyzes the audio data. This analysis analyzes the noise level and frequency components, and outputs the noise level in decibels (dB). The analyzed data is then converted into a specific answer by a generation engine.
[1277] Step 4:
[1278] The generation engine uses the analysis results to create a generated answer, which may contain specific information such as "The average noise level in this room is 50 decibels," and sends the generated answer to the user interface.
[1279] Step 5:
[1280] The terminal receives the generated answer data sent from the server and displays it to the user via a user interface, and if the user is using smart glasses, the answer is presented visually.
[1281] Step 6:
[1282] The emotion recognition means analyzes the user's voice data and face data. The emotion engine analyzes the user's facial expression and tone of voice when they input a question, and recognizes the user's emotional state (for example, surprise, satisfaction, dissatisfaction, etc.).
[1283] Step 7:
[1284] The server analyzes the emotion data sent from the emotion engine and provides additional feedback as needed. For example, if the user shows a surprised expression, a message such as "Would you like to provide additional information about why you are surprised?" is generated and sent to the device.
[1285] Step 8:
[1286] The terminal receives the emotion-based feedback data transmitted from the server and displays it to the user via a user interface, allowing the user to receive information corresponding to their emotions in real time.
[1287] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1288] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1289] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1290] [Fourth embodiment]
[1291] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1292] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1293] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1294] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1295] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1296] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1297] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1298] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1299] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1300] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1301] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1302] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1303] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1304] System Configuration
[1305] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves around the property. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal through a user interface.
[1306] System Operation
[1307] server
[1308] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including the property's noise level, vibration, and sunlight. It then analyzes the received data and extracts specific information to respond to the user's question. For example, if the question is about sunlight, it evaluates the changes in light during the day based on the video data. If the question is about noise levels, it analyzes the audio data and evaluates the noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device.
[1309] Terminal
[1310] The terminal accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. By checking these, the user can efficiently obtain detailed property information.
[1311] User
[1312] Users access the system through a dedicated application. Using the application's interface, they can begin viewing the property and remotely control the traveling device. During the viewing, users can enter questions about information they are interested in (e.g., "How sunny is the bedroom?") and receive answers in real time. Based on the answers, users can proceed with their property evaluation.
[1313] Specific examples
[1314] For example, consider the case where a user asks about sunlight in the living room. Using a dedicated application, the user inputs the question, "How much sunlight does the living room get?" The device sends this question to the server, which analyzes the video data of the living room. The server evaluates the sunlight in the living room from the analysis results and generates an answer using a natural language generation engine, such as "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface. This allows the user to obtain detailed information about the sunlight in the living room without actually visiting the property.
[1315] As described above, the present invention provides specific and detailed property information to users by combining a traveling device, server, and terminal installed within a property, allowing users to efficiently evaluate properties and select the most suitable home.
[1316] The processing flow will be explained below.
[1317] Step 1:
[1318] server
[1319] The server receives the activation signal from the traveling device and activates it within the property. The traveling device activates the camera and microphone and begins collecting video and audio data in real time.
[1320] Step 2:
[1321] server
[1322] The server receives video and audio data transmitted from the traveling device in real time, as well as environmental data such as noise levels, vibrations, and sunlight. The received data is temporarily stored.
[1323] Step 3:
[1324] User
[1325] The user launches the dedicated application and begins viewing the property. The user inputs a specific question (e.g., "How sunny is the living room?") through the interface.
[1326] Step 4:
[1327] Terminal
[1328] The terminal receives the user's question and transmits the question to the server.
[1329] Step 5:
[1330] server
[1331] The server analyzes the received question, identifies the question, and selects and analyzes relevant data (e.g., video data on sunlight) based on the identified question.
[1332] Step 6:
[1333] server
[1334] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[1335] Step 7:
[1336] server
[1337] The server generates a response and sends it to the terminal.
[1338] Step 8:
[1339] Terminal
[1340] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[1341] Step 9:
[1342] User
[1343] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[1344] In this way, the program proceeds step by step, providing the user with specific and detailed property information in real time.
[1345] Example 1
[1346] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1347] With conventional property viewing systems, it was difficult for users to obtain detailed information about the property in real time, resulting in low accuracy in property evaluations. Furthermore, there was a lack of efficient means for collecting and analyzing environmental data such as noise levels and sunlight, making it impossible for users to obtain the detailed answers they desired. This made it difficult for users to efficiently view properties and select the most suitable home.
[1348] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1349] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, means for receiving and analyzing environmental data including noise levels, vibrations, and sunlight levels within the property, means for transmitting questions input by users to the server, and means for displaying the analysis results received from the server and the generated answers on the user interface, thereby enabling users to obtain detailed information about the property in real time and efficiently evaluate the property.
[1350] A "traveling device" is a device that collects video and audio data while moving around the property.
[1351] The "server" is the core computer system that receives and analyzes data sent from the traveling device and generates answers to user questions.
[1352] A "user terminal" is a device that allows a user to access the system using a dedicated application, enter questions, and check answers.
[1353] "Video data" refers to data containing visual information about the inside of a property, and is captured by a traveling device.
[1354] "Audio data" refers to data containing auditory information within a property, and is recorded by the traveling device.
[1355] "Environmental data" refers to data that includes information about the environment within the property, such as noise levels, vibrations, and sunlight.
[1356] "Analysis" is the process of extracting information from the received data and identifying information that responds to the user's question.
[1357] The "generation engine" is an engine for generating specific answers to user questions based on the analysis results.
[1358] A "generative AI model" is an artificial intelligence model that generates natural language responses based on input prompts.
[1359] A "prompt sentence" is a sentence that is input into a generative AI model and serves as the basis for the answer to a question that is generated based on the analysis results.
[1360] "User interface" refers to an interface that allows a user to interact with a system, and includes the screen displayed on the terminal and the operation method.
[1361] MODE FOR CARRYING OUT THE INVENTION
[1362] System configuration
[1363] The present invention is a system consisting of a traveling device installed within a property, a server, and a user terminal. The traveling device collects video and audio data as it moves within the property and transmits it to the server. The server analyzes the received data and generates answers to users' questions. The generated answers are displayed on the user terminal through a user interface. Users can access the system via a dedicated application and input questions.
[1364] Hardware and software used
[1365] Mobile device: A device that moves autonomously or remotely around a property and collects video and audio data using cameras and microphones.
[1366] Server: A high-performance computer that provides an environment for receiving data, analyzing it, and running generative AI models. For example, it can use cloud-based computing services.
[1367] User device: A device such as a smartphone, tablet, or PC with a dedicated application installed.
[1368] Data processing and calculation
[1369] 1. Data reception:
[1370] The server receives real-time video and audio data transmitted from the traveling device, as well as environmental data such as noise levels, vibrations, and sunlight levels at the property.
[1371] 2. Data Analysis:
[1372] The server analyzes the received data and extracts specific information that responds to the user's question, for example by implementing algorithms that analyze video data to evaluate changes in light and shadow movement during the day.
[1373] 3. Answer generation:
[1374] Based on the analysis results, the generative AI model creates a prompt sentence and generates a natural language answer based on that prompt sentence. For example, the prompt sentence could be, "The user is asking about the sunlight in the living room. Please evaluate the sunlight in the living room based on the video data of the living room and generate a natural language answer explaining the results."
[1375] 4. Answer display:
[1376] The generated answer is sent to the user terminal through the user interface and displayed on the application screen.
[1377] Specific examples
[1378] For example, if a user asks about sunlight in the living room, the following process is performed.
[1379] 1. User enters a question: The user uses a dedicated application to enter a question such as, "How much sunlight does the living room get?"
[1380] 2. The device sends the question to the server: The device sends this question to the server.
[1381] 3. The server analyzes the data: The server analyzes the video data from the living room and evaluates the amount of sunlight.
[1382] 4. Answer generation: Based on the prompt, the server's generative AI model generates the answer, "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[1383] 5. Displaying the answer: This answer is sent to the terminal and displayed to the user through the user interface.
[1384] Prompt Sentence Examples
[1385] "The user is asking about the sunlight in the living room. Based on the video data of the living room, please evaluate the sunlight in the living room and generate a natural language answer explaining the results."
[1386] As described above, this system efficiently provides users with specific and detailed property information and helps them evaluate properties, enabling them to obtain information to select the most suitable home.
[1387] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1388] Step 1: User enters question
[1389] The user launches the dedicated application and inputs a question through the interface, such as "How much sunlight does the living room get?" The application temporarily stores the question in memory.
[1390] Input: Questions entered into the user interface
[1391] Output: Question data stored in memory
[1392] Specific behavior:
[1393] The user starts the application and a question input screen is displayed.
[1394] Enter your question and tap the "Submit" button.
[1395] Step 2: Sending a question to the server from the device
[1396] The device uses a dedicated communication protocol to send the question data to the server, such as an HTTP POST request.
[1397] Input: Question data stored in memory
[1398] Output: Question data sent to the server
[1399] Specific behavior:
[1400] The device retrieves the question data stored in memory and sends it to the server as an HTTP POST request.
[1401] Step 3: Server receives data
[1402] The server receives the query data sent from the terminal. It then begins receiving video and audio data sent from the traveling device in real time. It also simultaneously receives environmental data such as noise levels, vibrations, and sunlight.
[1403] Input: Question data sent from the terminal, video and audio data from the driving device, environmental data
[1404] Output: Question data and various data stored on the server
[1405] Specific behavior:
[1406] The server receives the query data and stores it in a database.
[1407] The server receives data from the running gear and environmental sensors in real time and prepares it for analysis.
[1408] Step 4: Data analysis by the server
[1409] The server analyzes the data it receives and extracts information to answer questions, such as analyzing video data to assess changes in light during the day or audio data to assess noise levels.
[1410] Input: Video, audio, environmental data, and question data stored on the server
[1411] Output: Specific analysis results that address your questions
[1412] Specific behavior:
[1413] An algorithm is run to detect changes in light from the video data.
[1414] Audio data is analyzed to evaluate noise levels by time of day.
[1415] Step 5: Server Generates Answer
[1416] Based on the analysis results, the server creates a prompt to be applied to the generative AI model. The prompt is input into the generative AI model, which generates a natural language answer. For example, it generates an answer such as, "The living room has good sunlight in the morning and is partially shaded in the afternoon."
[1417] Input: Analysis results, generative AI model, prompt
[1418] Output: Natural language response
[1419] Specific behavior:
[1420] The server creates a prompt based on the analysis results.
[1421] The prompt sentence is input into a generative AI model to generate a natural language response.
[1422] Step 6: Sending the response from the server to the device
[1423] The server sends the generated response to the user's terminal using a communication protocol such as an HTTP POST request.
[1424] Input: Natural language response data
[1425] Output: Answer data sent to the user's device
[1426] Specific behavior:
[1427] The server sends the generated response to the device as an HTTP POST request.
[1428] Step 7: Viewing the Answers on Your Device
[1429] The terminal stores the response data received from the server in a memory and displays it on a user interface.
[1430] Input: Response data received from the server
[1431] Output: The answer displayed in the user interface
[1432] Specific behavior:
[1433] The terminal analyzes the received response data and displays it on the user interface.
[1434] (Application example 1)
[1435] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1436] Management and monitoring of logistics facilities are important for improving efficiency and ensuring safety. However, conventional systems make it difficult to physically inspect the site and integrate and analyze a wide range of environmental data, placing a heavy burden on managers. Furthermore, there are challenges in responding quickly to abnormalities and providing detailed information in real time.
[1437] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1438] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within the property, means for analyzing the received video and audio data, and a generation engine means for generating answers to specific questions based on the analysis results. This enables means for receiving and analyzing environmental data for monitoring and management within the logistics facility, such as noise levels, vibrations, and lighting conditions, means for issuing automatic alerts when an abnormality is detected, means for allowing a manager of the logistics facility to input a question and generating a specific answer to the question using a generative AI model, and means for displaying the answer to the question via a user interface.
[1439] A "traveling device installed within the property" is a device that moves autonomously or remotely within the logistics facility and collects video and audio data.
[1440] The "means for receiving video and audio data" refers to a means having a function for transmitting video and audio data collected by the traveling device to a server in real time and for the server to receive the data.
[1441] "Means for analyzing video and audio data" refers to technical means for analyzing received video and audio data and extracting specific information, and examples include video analysis software and audio analysis software.
[1442] The "generation engine means" refers to means including algorithms and programs for generating specific answers to user questions based on analyzed data.
[1443] The "means for displaying via a user interface" refers to a means including an interface and its functions for displaying the generated answer on a user terminal.
[1444] "Means for receiving and analyzing environmental data" refers to the technical means for collecting and analyzing environmental data such as noise levels, vibrations, and lighting conditions within a logistics facility.
[1445] "Means for issuing automatic alerts when an abnormality is detected" refers to a means that has the function of immediately notifying an administrator when an abnormality is detected based on collected and analyzed data.
[1446] "Generative AI models" are artificial intelligence algorithms and models used for natural language processing and information generation, examples of which include GPT-4.
[1447] The "means for displaying answers to questions via a user interface" refers to an interface and means including its functions for displaying answers generated by the generation engine to questions from users on the display screen of a user terminal.
[1448] System Configuration
[1449] This invention is a system consisting of a traveling device, a server, and a user terminal installed within a logistics facility. The traveling device collects video and audio data as it moves within the logistics facility and transmits it to the server. The server receives and analyzes this data. Based on the analysis results, a generation engine generates specific answers to the user's questions and displays them on the user terminal via a user interface.
[1450] Specific operation of the system
[1451] server
[1452] The server first receives video and audio data transmitted from the traveling device in real time. It also simultaneously receives environmental data, including noise levels, vibrations, and lighting conditions within the logistics facility. It then analyzes the received data and extracts specific information to respond to user questions. For example, if a question is about noise levels, the server analyzes the audio data and evaluates noise levels during the day and night. Based on the analysis results, the generation engine generates a response in natural language and sends it to the user's device via the user interface.
[1453] The server uses the following specific technologies:
[1454] Hardware: Server machine
[1455] Software: Video analysis software (e.g., OpenCV), audio analysis software (e.g., DeepSpeech), natural language generation engines (e.g., GPT-4), real-time communication systems (e.g., WebSocket)
[1456] Specific examples
[1457] For example, consider the case where a manager asks about the noise level in a specific location in a logistics facility. Using a dedicated application, the manager inputs the question, "What is the current noise level?" The device sends this question to the server, which analyzes the collected noise data. The server evaluates the noise level based on the analysis results and generates an answer using a natural language generation engine, such as "The current noise level is an average of 85 decibels." This answer is sent to the device and displayed to the manager through a user interface. This allows the manager to obtain detailed information about the specific situation at the logistics facility.
[1458] User terminal
[1459] The terminal accepts questions from the user and sends them to the server. It receives the analysis results and generated answers from the server and displays them on the user interface. By checking these, the user can efficiently obtain detailed information about the logistics facility. The system also has a function that displays an automatic alert if an abnormality is detected.
[1460] User
[1461] Users access the system through a dedicated application. Using the application's interface, they can start monitoring the logistics facility and remotely control the traveling device. While monitoring, users can input questions about information of interest (e.g., "What is the lighting situation in this area?") and receive answers in real time. Based on the answers, users can evaluate and manage the logistics facility.
[1462] Prompt Sentence Examples
[1463] People: What is the noise level in your current distribution center?
[1464] AI: The current noise level averages 85 decibels, with peaks occurring between 2 and 3 p.m.
[1465] As described above, this system combines on-site traveling devices with a server and terminals to provide users with specific and detailed information about logistics facilities, allowing them to efficiently manage and evaluate the facilities.
[1466] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1467] Step 1:
[1468] The user launches the dedicated application and begins monitoring the logistics facility.
[1469] Input: The user launches the application
[1470] Processing: The user interface is launched and the running device is ready for remote operation.
[1471] Output: The control screen for the running gear is displayed.
[1472] Step 2:
[1473] The user remotely controls the traveling device, which patrols the logistics facility and collects video and audio data.
[1474] Input: User operates the running gear
[1475] Processing: As the traveling device moves, it collects video and audio using cameras and microphones.
[1476] Output: Real-time collected video and audio data
[1477] Step 3:
[1478] The traveling device transmits the collected video and audio data to a server.
[1479] Input: Video and audio data from the traveling device
[1480] Processing: Data is sent to the server via wireless communication
[1481] Output: Video and audio data received by the server
[1482] Step 4:
[1483] The server parses the received data.
[1484] Input: Video and audio data received by the server
[1485] Processing: The server analyzes the data using video analysis software (e.g., OpenCV) and audio analysis software (e.g., DeepSpeech).
[1486] Output: Analysis results (e.g. noise level, specific object recognition)
[1487] Step 5:
[1488] The user inputs a question via a dedicated application.
[1489] Input: User types a question into the device (e.g., "What is the current noise level?")
[1490] Process: The question is sent to the server
[1491] Output: The question sent to the server
[1492] Step 6:
[1493] Based on the analysis results, the server generates an answer to the question using a generative AI model (e.g., GPT-4).
[1494] Input: The question and analysis results sent to the server
[1495] Processing: The generative AI model generates a natural language answer to the question.
[1496] Output: The generated answer (e.g., "The current noise level is 85 decibels")
[1497] Step 7:
[1498] The generated answer is displayed on the user terminal via a user interface.
[1499] Input: Generated Answer
[1500] Processing: The answer is sent to the terminal and displayed in the user interface.
[1501] Output: The answer displayed in the user interface
[1502] Step 8:
[1503] The server will automatically send an alert if an abnormality is detected.
[1504] Input: Analysis results (e.g., noise level exceeds threshold)
[1505] Action: The server generates an automatic alert and notifies the user interface.
[1506] Output: The alert displayed in the user interface
[1507] In this way, the entire system can monitor detailed conditions within the logistics facility in real time, enabling efficient and safe operation and management.
[1508] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1509] System Configuration
[1510] The present invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. The traveling device collects video, audio, and environmental data as it moves within the property. The server receives this data and analyzes it based on the user's questions. The emotion engine then recognizes emotions from the user's audio and facial data, and displays the analysis results and generated answers on the user terminal via a user interface.
[1511] System Operation
[1512] server
[1513] The server first receives video, audio, and environmental data transmitted from the traveling device in real time. This includes data on noise levels, vibrations, and sunlight. The server then temporarily stores the received data and analyzes it based on the user's questions. For example, video data is used for questions about sunlight, and audio data is analyzed for questions about noise. Specific information is extracted and a specific answer is generated using a generation engine. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and provides feedback according to the user's emotional state.
[1514] Terminal
[1515] The device accepts questions from the user and sends them to the server. The server then receives the analysis results and generated answers, which are then displayed on the user interface. Furthermore, the device also displays the analysis results of the emotion engine, adjusting the information provided and questions asked during the viewing as necessary.
[1516] User
[1517] Users access the system through a dedicated application. Using the application's interface, they can begin viewing properties and remotely control the vehicle. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[1518] Specific examples
[1519] For example, consider the case where a user asks about sunlight in the living room. The user inputs the question, "How sunny is the living room?" The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer generated is, "The living room gets good sunlight in the morning, but is partially shaded in the afternoon." This answer is sent to the device and displayed to the user through the user interface.
[1520] If a user shows a surprised expression during a viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, and appropriate feedback will be provided according to the user's emotion.
[1521] As described above, the present invention combines a traveling device placed within a property with a server, a terminal, and an emotion engine to provide specific and detailed property information and personalized feedback to users, allowing them to efficiently evaluate properties and choose the most suitable home.
[1522] The processing flow will be explained below.
[1523] Step 1:
[1524] server
[1525] The server receives the activation signal from the vehicle and activates it within the property. The vehicle then activates its camera and microphone to begin collecting video and audio data in real time. It also collects environmental data such as noise levels, vibrations, and sunlight.
[1526] Step 2:
[1527] server
[1528] The server receives video, audio and environmental data transmitted from the traveling device in real time and temporarily stores this data.
[1529] Step 3:
[1530] User
[1531] The user launches the dedicated application and begins viewing the property, entering a specific question (e.g., "How sunny is the living room?") through the application's interface.
[1532] Step 4:
[1533] Terminal
[1534] The terminal receives the user's question and transmits the question to the server.
[1535] Step 5:
[1536] server
[1537] The server analyzes the received question, identifies the content of the question, and selects and analyzes data related to the question, such as video data and audio data related to sunlight.
[1538] Step 6:
[1539] server
[1540] Based on the analysis results, the server uses a generation engine to generate a specific answer to the user's question, such as "The living room gets good sunlight in the morning and is partially shaded in the afternoon."
[1541] Step 7:
[1542] server
[1543] The server generates a response and sends it to the terminal.
[1544] Step 8:
[1545] Terminal
[1546] The terminal displays the received answers on the user interface, and the user evaluates the property based on the displayed answers.
[1547] Step 9:
[1548] User
[1549] The user can then enter another question or end the preview. Clicking the End Preview button will end the session and stop the vehicle.
[1550] Step 10:
[1551] server
[1552] The server collects the user's voice and facial data and analyzes it with an emotion engine to recognize the user's emotional state (e.g., joy, sadness, surprise, anger).
[1553] Step 11:
[1554] server
[1555] Based on the analysis results of the emotion engine, feedback and additional information are generated according to the user's emotional state. For example, if the user is surprised, a message such as "Would you like to know more about the part that surprised you?" is generated.
[1556] Step 12:
[1557] server
[1558] Feedback and additional information from the emotion engine is sent to the device.
[1559] Step 13:
[1560] Terminal
[1561] The device displays feedback and additional information in the user interface according to the user's emotions, allowing the user to view the property in more detail.
[1562] In this way, the program proceeds step by step, providing the user with specific and detailed property information and emotional feedback in real time.
[1563] Example 2
[1564] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1565] Conventional property viewing systems are limited to collecting and analyzing video and audio data, and are unable to provide feedback tailored to the user's emotional state. Furthermore, they are limited in generating answers when users input specific questions, and lack the ability to provide personalized information in real time. This makes it difficult for users to efficiently obtain more accurate and detailed property information.
[1566] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1567] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation AI model means for generating answers to specific questions based on the analysis results, means for recognizing emotions from the user's voice data and face data, means for generating feedback based on the emotion recognition results and the analysis results, and means for displaying the generated answers and feedback via a user interface. This allows the user to obtain specific and detailed property information in real time and receive personalized feedback according to their emotions.
[1568] A "traveling device" is a device that collects video data, audio data, and environmental data while moving around the property.
[1569] "Video data" refers to data of moving images and still images taken by the traveling device.
[1570] "Audio data" is data of sounds collected by the traveling device.
[1571] "Environmental data" refers to data that includes information such as noise levels, vibrations, and sunlight levels within the property.
[1572] A "generative AI model" is an artificial intelligence model that analyzes received data and generates answers to specific questions.
[1573] "Emotion recognition" is the process of analyzing a user's voice and facial data to identify the user's emotional state.
[1574] A "user interface" is a screen or operating means that allows a user to exchange information with a system.
[1575] "Feedback" refers to responses or information provided to the user based on the analysis results and emotion recognition results.
[1576] "Real-time" refers to immediate processing or response with little or no delay.
[1577] "Analysis" is the process of extracting information and deriving meaning from received data.
[1578] The present invention is a property viewing system that includes a traveling device, a server, a user terminal, and an emotion engine. The traveling device moves around the property and collects video, audio, and environmental data. The server receives this data in real time and analyzes it based on the user's questions. Furthermore, the emotion engine recognizes emotions from the user's voice data and facial data, and displays the analysis results and generated answers on the user terminal through a user interface.
[1579] Hardware and software used
[1580] Hardware:
[1581] Mobile devices (e.g., robotic cameras, mobile sensor devices)
[1582] Servers (e.g., high-performance servers in data centers)
[1583] User device (e.g. smartphone, tablet)
[1584] Hardware for emotion engine (e.g. high-resolution camera, directional microphone)
[1585] software:
[1586] Data collection software (e.g. camera control software, sensor interface)
[1587] Data analysis software (e.g., image analysis algorithms, audio analysis algorithms)
[1588] Emotion recognition software (e.g., voice emotion recognition model, facial expression analysis model)
[1589] User interfaces (e.g., mobile apps, web applications)
[1590] Specific operation of the system
[1591] The server first receives video and audio data and environmental data (such as noise levels, vibrations, and sunlight) sent from the driving device in real time. This data is temporarily stored and analyzed based on questions from the user. Specific information is extracted using various algorithms for analysis, and specific answers are generated using generative AI models. Meanwhile, the emotion engine recognizes emotions from the user's voice and facial data and generates feedback according to the user's emotional state.
[1592] The device accepts questions from the user as input and sends them to the server. The analysis results and generated answers are received from the server in real time and displayed through the user interface. Furthermore, the emotion engine's analysis results are also displayed, and the information provided during the viewing and questions can be adjusted as needed.
[1593] Users access the system through a dedicated application. Using the application's interface, they can start viewing properties and remotely control the traveling device. During the viewing, they can enter questions about information they are interested in and receive answers in real time. Furthermore, the emotion engine recognizes the user's emotions, allowing them to receive more personalized information.
[1594] Specific use cases
[1595] For example, consider the case where a user asks a question about the sunlight in the living room. The user enters "How sunny is the living room?" into the device. The device sends this question to the server, which analyzes the video data of the living room. As a result of the analysis, the answer "The living room gets good sunlight in the morning and is partially shaded in the afternoon" is generated. This answer is sent to the device and displayed to the user through the user interface.
[1596] Furthermore, if the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and the server will analyze it. For example, a message such as "Would you like to provide additional information about the part that surprised you?" will be displayed, providing appropriate feedback according to the user's emotion.
[1597] Examples of prompt statements
[1598] An example of a prompt is as follows:
[1599] prompt:
[1600] How's the sunlight in the living room?
[1601] Expected answer:
[1602] "The living room gets good sunlight in the morning, but some shading is expected in the afternoon."
[1603] As described above, the present invention combines a traveling device, a server, a user terminal, and an emotion engine to provide users with specific and detailed property information and personalized feedback, allowing them to efficiently evaluate properties and select the most suitable home.
[1604] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1605] Program processing flow
[1606] Step 1: Data collection
[1607] Processing Description:
[1608] The vehicle travels around the property, collecting video and audio data using cameras and microphones, and environmental sensors to collect environmental data such as noise levels, vibrations, and sunlight.
[1609] input:
[1610] Video, audio, noise level, vibration, sunlight
[1611] output:
[1612] Real-time transmission of video data, audio data, and environmental data
[1613] Specific operation:
[1614] As the traveling device moves around the living room, it takes pictures of the entire room with a camera, collects sound with a microphone, and measures the temperature with a sensor.
[1615] Step 2: Data reception and storage
[1616] Processing Description:
[1617] The server receives video data, audio data, and environmental data transmitted from the traveling device in real time and temporarily stores it.
[1618] input:
[1619] Data transmitted from the traveling device (video, audio, environmental data)
[1620] output:
[1621] Data stored in the database
[1622] Specific operation:
[1623] The server stores the received data in a dedicated database and prepares it for analysis.
[1624] Step 3: Ask a question
[1625] Processing Description:
[1626] The terminal accepts a question from the user as input and sends it to the server.
[1627] input:
[1628] User questions (e.g., "How sunny is the living room?")
[1629] output:
[1630] Send the question to the server
[1631] Specific operation:
[1632] The user enters a question in the application interface and presses the submit button, which sends the question to the server via the API.
[1633] Step 4: Data analysis
[1634] Processing Description:
[1635] The server analyzes the question, filters and extracts relevant video, audio, and environmental data, and uses a generative AI model to generate a specific answer.
[1636] input:
[1637] User questions, saved data (video, audio, environmental data)
[1638] output:
[1639] Generated Answer
[1640] Specific operation:
[1641] The server analyzes the question "How much sunlight does the living room get?", filters and analyzes the video data of the living room, and uses a generative AI model to generate the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon."
[1642] Step 5: Emotion Recognition
[1643] Processing Description:
[1644] The emotion engine recognizes emotions from the user's voice data and facial data, and sends the recognition results to the server, which then analyzes the emotion data and generates feedback.
[1645] input:
[1646] User voice data, face data
[1647] output:
[1648] Emotion recognition results and generated feedback
[1649] Specific operation:
[1650] The emotion engine monitors the user's facial expressions to detect smiles and surprises. For example, if the user shows a surprised expression, the emotion engine recognizes the surprise and generates feedback such as, "Would you like to provide additional information about what is surprising?"
[1651] Step 6: Submit and view your answers and feedback
[1652] Processing Description:
[1653] The server transmits the generated answers and feedback to the user terminal, which displays the received information on a user interface.
[1654] input:
[1655] Generated answers, generated feedback
[1656] output:
[1657] Data sent to user terminals, display data
[1658] Specific operation:
[1659] The server sends the answer "The living room gets good sunlight in the morning, but is partially shaded in the afternoon" and feedback "Do you like this property?" to the user's device. The device receives this information and displays it on the application screen.
[1660] In this way, the entire system works together to consistently collect and analyze data, provide real-time answers to user questions, and even provide emotional feedback.
[1661] (Application example 2)
[1662] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1663] Conventional property viewing systems required users to personally visit the property, which was time-consuming and costly, and the information obtained from a single viewing was limited. Furthermore, it was difficult to provide personalized information based on the user's emotions and personal preferences, resulting in insufficient evaluations of properties depending on the user. Furthermore, due to a lack of analysis of environmental data and appropriate feedback, users were sometimes unaware of potential problems with the property. Thus, to improve user convenience and satisfaction, a more detailed and individually tailored property information system was needed.
[1664] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1665] In this invention, the server includes means for receiving video and audio data captured by a traveling device installed within a property, means for analyzing the received video and audio data, a generation engine means for generating answers to specific questions based on the analysis results, means for displaying the generated answers via a user interface, and emotion recognition means for recognizing the user's emotions and providing feedback according to the user's emotional state. This allows the user to remotely view the property and receive answers in real time, and also allows the user to be provided with information according to their individual emotions, making it possible to efficiently obtain more detailed and personalized property information.
[1666] A "moving device" is a device that moves autonomously or remotely within a property to collect video and audio data.
[1667] The "generation engine means" is an engine that has the function of generating an answer to a specific question from a user based on the analysis results.
[1668] A "user interface" is an interface for displaying generated answers and analysis results to the user, and is a means for exchanging information between the user and the system.
[1669] The "emotion recognition means" is a means that has the function of recognizing emotions from data such as the user's voice and facial expression, and providing feedback according to those emotions.
[1670] "Environmental data" refers to data relating to environmental conditions such as noise levels, vibrations, and sunlight levels within the property.
[1671] A "visualization device" is a device worn by a user that collects video and audio data and provides the user with visual and audio information in real time.
[1672] A "biological signal" is a signal that contains information about the user's health condition and fatigue level, and includes heart rate, body temperature, skin potential, and the like.
[1673] "Personalized information" refers to information and feedback that is customized based on a user's individual emotional and health state.
[1674] This invention is a system that includes a traveling device installed within a property, a server, a user terminal, and an emotion engine for recognizing the user's emotions. Specific embodiments of this system are described below.
[1675] System Configuration
[1676] The server has a means for receiving video and audio data captured by the traveling device installed within the property. The received data is processed using an analysis engine to generate an optimal answer based on the user's question. The generated answer is displayed on the user's terminal via a user interface. In addition, an emotion recognition means is used to analyze the user's emotional state and provide feedback as necessary.
[1677] Hardware and Software
[1678] The system uses the following hardware and software:
[1679] Mobile device: Responsible for moving around the property and collecting video and audio data.
[1680] Server: Receives and analyzes data, generates answers via a generation engine, and recognizes emotions.
[1681] User terminal (smartphone, smart glasses, head-mounted display, etc.): displays information through a user interface.
[1682] Emotion engine: Recognizes emotions from the user's facial data and voice.
[1683] Software used: OpenCV (video data processing), EmotionRecognizer (emotion recognition), Server (communication with server).
[1684] Data processing and calculation
[1685] The server receives video and audio data from the vehicle in real time and temporarily stores it. The stored data is analyzed to extract specific information in response to user questions. A generative AI model is used to generate specific answers, which are then displayed to the user through a user interface.
[1686] The emotion engine analyzes the user's voice and facial data to recognize their emotions. The recognized emotions are sent to the server, which provides appropriate feedback along with the analysis results.
[1687] Specific examples
[1688] For example, if a user asks about the noise level in a room, the user inputs the question, "What is the noise level in this room?" The device sends this question to the server, which analyzes the voice data. As a result of the analysis, the answer "The average noise level in this room is 50 decibels" is generated and sent to the device. This answer is then displayed to the user through the user interface.
[1689] If the user shows a surprised expression during the viewing, the emotion engine will recognize the emotion and provide appropriate feedback according to the user's emotions, such as displaying a message like, "Would you like to provide additional information about the part that surprised you?"
[1690] Prompt Sentence Examples
[1691] "When you inquire about a new product, please send us the data of the moment that surprised you the most."
[1692] "Analyze user interests using facial recognition data and suggest appropriate products."
[1693] This system allows users to efficiently obtain more detailed and personalized information, and also provides feedback according to the user's emotions, thereby improving user satisfaction.
[1694] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1695] Step 1:
[1696] The device accepts questions from users as input. The user inputs the question through an interface such as smart glasses or a smartphone. The input data is sent to the server. The input content may include, for example, "What is the noise level in this room?"
[1697] Step 2:
[1698] The server receives the query data sent from the device. Based on the received query data, it acquires and temporarily stores the video and audio data collected from the traveling device. The temporarily stored data includes video of the room, environmental sounds, vibration data, and sunlight amount.
[1699] Step 3:
[1700] The server analyzes the temporarily stored data. Specifically, if a question is about noise levels, for example, the server analyzes the audio data. This analysis analyzes the noise level and frequency components, and outputs the noise level in decibels (dB). The analyzed data is then converted into a specific answer by a generation engine.
[1701] Step 4:
[1702] The generation engine uses the analysis results to create a generated answer, which may contain specific information such as "The average noise level in this room is 50 decibels," and sends the generated answer to the user interface.
[1703] Step 5:
[1704] The terminal receives the generated answer data sent from the server and displays it to the user via a user interface, and if the user is using smart glasses, the answer is presented visually.
[1705] Step 6:
[1706] The emotion recognition means analyzes the user's voice data and face data. The emotion engine analyzes the user's facial expression and tone of voice when they input a question, and recognizes the user's emotional state (for example, surprise, satisfaction, dissatisfaction, etc.).
[1707] Step 7:
[1708] The server analyzes the emotion data sent from the emotion engine and provides additional feedback as needed. For example, if the user shows a surprised expression, a message such as "Would you like to provide additional information about why you are surprised?" is generated and sent to the device.
[1709] Step 8:
[1710] The terminal receives the emotion-based feedback data transmitted from the server and displays it to the user via a user interface, allowing the user to receive information corresponding to their emotions in real time.
[1711] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1712] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1713] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1714] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1715] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1716] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1717] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1718] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1719] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1720] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1721] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1722] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1723] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1724] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1725] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1726] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1727] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1728] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1729] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1730] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1731] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1732] The following is further disclosed regarding the above embodiment.
[1733] (Claim 1)
[1734] a means for receiving video and audio data captured by a traveling device installed within the property;
[1735] means for analyzing the received video and audio data;
[1736] a generation engine means for generating answers to specific questions based on the analysis results;
[1737] means for displaying the generated answers via a user interface;
[1738] A system including:
[1739] (Claim 2)
[1740] 10. The system of claim 1, further comprising means for receiving and analyzing environmental data in addition to the received video and audio data.
[1741] (Claim 3)
[1742] 3. The system of claim 2, wherein the environmental data includes noise levels, vibrations, and solar radiation.
[1743] "Example 1"
[1744] (Claim 1)
[1745] a means for receiving video and audio data captured by a traveling device installed within the property;
[1746] means for analyzing the received video and audio data;
[1747] a generation engine means for generating answers to specific questions based on the analysis results;
[1748] means for displaying the generated answers via a user interface;
[1749] means for receiving and analyzing environmental data including noise levels, vibrations, and sunlight levels within the property;
[1750] means for transmitting a question input by a user to a server;
[1751] means for displaying the analysis results received from the server and the generated answers on a user interface;
[1752] A system including:
[1753] (Claim 2)
[1754] The system according to claim 1, wherein the generation engine means creates a prompt sentence based on the analysis result and uses the prompt sentence to generate an answer using the generative AI model.
[1755] (Claim 3)
[1756] 2. The system according to claim 1, wherein a user inputs a question through a user interface and remotely controls the traveling device.
[1757] "Application Example 1"
[1758] (Claim 1)
[1759] a means for receiving video and audio data captured by a traveling device installed within the property;
[1760] means for analyzing the received video and audio data;
[1761] a generation engine means for generating answers to specific questions based on the analysis results;
[1762] means for displaying the generated answers via a user interface;
[1763] Furthermore, a means for receiving and analyzing environmental data for monitoring and management within the logistics facility, such as noise levels, vibrations, and lighting conditions;
[1764] A means of sending automatic alerts when an anomaly is detected; and
[1765] A means for a logistics facility manager to input a question and generate a specific answer to the question using a generative AI model;
[1766] means for displaying answers to questions via a user interface;
[1767] A system including:
[1768] (Claim 2)
[1769] 10. The system of claim 1, further comprising means for receiving and analyzing environmental data in addition to the received video and audio data.
[1770] (Claim 3)
[1771] 2. The system of claim 1, wherein the environmental data includes noise levels, vibrations, and lighting conditions.
[1772] "Example 2: Combining Emotion Engines"
[1773] (Claim 1)
[1774] a means for receiving video and audio data captured by a traveling device installed within the property;
[1775] means for analyzing the received video and audio data;
[1776] a generative AI model means for generating answers to specific questions based on the analysis results;
[1777] means for recognizing emotions from voice data and facial data of a user;
[1778] a means for generating feedback based on the emotion recognition and analysis results;
[1779] means for displaying the generated answers and feedback via a user interface;
[1780] A system including:
[1781] (Claim 2)
[1782] 10. The system of claim 1, further comprising means for receiving and analyzing environmental data in addition to the received video and audio data.
[1783] (Claim 3)
[1784] 3. The system of claim 2, wherein the environmental data includes noise levels, vibrations, and solar radiation.
[1785] "Application example 2 when combining emotion engines"
[1786] (Claim 1)
[1787] a means for receiving video and audio data captured by a traveling device installed within the property;
[1788] means for analyzing the received video and audio data;
[1789] a generation engine means for generating answers to specific questions based on the analysis results;
[1790] means for displaying the generated answers via a user interface;
[1791] emotion recognition means for recognizing the emotion of a user and providing feedback according to the emotional state;
[1792] A system including:
[1793] (Claim 2)
[1794] 10. The system of claim 1, further comprising means for receiving and analyzing environmental data in addition to the received video and audio data.
[1795] (Claim 3)
[1796] 3. The system of claim 2, wherein the environmental data includes noise levels, vibrations, and solar radiation.
[1797] (Claim 4)
[1798] The system according to claim 1, further comprising means for collecting video and audio data through a visualization device worn by the user, analyzing the user's emotions in real time, and providing information appropriate to the situation based on the analysis results.
[1799] (Claim 5)
[1800] 5. The system according to claim 4, further comprising means for receiving a biological signal and providing customized information based on the user's health condition and fatigue level. [Explanation of symbols]
[1801] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for receiving video and audio data captured by a traveling device installed within the property; means for analyzing the received video and audio data; a generation engine means for generating answers to specific questions based on the analysis results; means for displaying the generated answers via a user interface; A system including:
2. 10. The system of claim 1, further comprising means for receiving and analyzing environmental data in addition to the received video and audio data.
3. 3. The system of claim 2, wherein the environmental data includes noise levels, vibrations, and solar radiation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A