system
The system addresses the challenge of remote property viewing by providing real-time, detailed information and emotional analysis, enabling efficient and personalized virtual tours.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-30
AI Technical Summary
Conventional virtual interior viewing technologies for rental properties fail to convey specific usage feelings and details to users, especially for those living remotely, leading to inefficiencies and high costs in physical interior viewing.
A system that retrieves detailed property information from a database in real-time, converts user voice input to text, and provides accurate answers, using noise reduction and accent correction, while rendering 3D visuals and interactive maps for remote viewing.
Enables users to efficiently obtain detailed property information remotely, facilitating flexible responses and personalized interactions based on emotional analysis, enhancing the viewing experience.
Smart Images

Figure 2026071581000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional virtual interior viewing technologies for rental properties have the problem that they cannot fully convey the specific usage feelings and details of the properties to users, and for people living in remote areas, there are problems such as the trouble and cost of physical interior viewing. An effective method for users to obtain detailed information about the property under consideration without visiting the site is required.
Means for Solving the Problems
[0005] This invention provides a system that efficiently retrieves detailed information about a property from a database based on property selection information from a user's terminal, and provides it to the user in a compressed format in real time. In this way, users can view properties in detail through high-quality video and interactive maps. Furthermore, the system converts user voice input into text, quickly retrieves relevant information from the database, generates answers, and returns them to the user, enabling flexible responses to individual questions. In addition, noise reduction and accent correction functions using speech recognition technology enable the provision of accurate information to each user.
[0006] A "user terminal" refers to an electronic device used by a user to receive and input information.
[0007] "Property selection information" refers to data that includes identifying information about properties that the user is interested in.
[0008] A "database" refers to an electronic information storage system designed to efficiently store and retrieve information.
[0009] "Detailed information" refers to specific data related to the property, such as floor plan, facilities, and surrounding environment.
[0010] "Real-time compression and conversion to streaming format" refers to the process of instantly compressing data into a format that can be transmitted, and then continuously delivering it over the internet.
[0011] "Voice input" refers to voice information spoken by the user, which is used as instructions or questions for the system.
[0012] "Speech recognition technology" refers to the technology that analyzes input speech data and converts it into corresponding text strings.
[0013] "Noise reduction" refers to the process of removing unwanted noise from audio data.
[0014] "Accent correction" refers to a process that corrects pronunciation differences in voice data to perform accurate information recognition.
[0015] "Three-dimensional visual" refers to a visual representation generated based on a three-dimensional space, providing a realistic feeling to the user.
[0016] "Interactive map" refers to an interface that displays map information operable by a user, enabling the user to check navigation and information details.
Brief Explanation of Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of a data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Embodiments for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] This invention provides a system that allows users to understand rental properties in detail and resolve any questions they may have, even from a remote location. This system functions through the collaborative efforts of a server, a terminal, and the user.
[0039] Server-based data management and processing
[0040] The server manages a database containing detailed information about properties. When it receives property selection information from a user's terminal, the server immediately retrieves the corresponding property information from the database. The retrieved information includes 360-degree video, 3D models, basic property data, and surrounding environment information.
[0041] The server compresses this information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user's terminal over the internet. When the server receives voice input from the user, it uses speech recognition technology to convert the voice into text data and queries the database. Based on the results, the server generates and returns an answer to the user's question.
[0042] User interface via terminal
[0043] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail.
[0044] Furthermore, the device receives the user's voice input through its voice recording function and sends it directly to the server. The response sent back from the server is played back through the speaker using speech synthesis technology and simultaneously displayed as text on the device screen.
[0045] User actions
[0046] Users can easily search for their desired properties and begin virtual tours using their devices. Specifically, they can view property information and videos displayed on the interface and ask the system questions verbally about areas of interest or anything they are unsure of.
[0047] For example, if a user asks, "What is the ceiling height of this room?", the system will immediately provide accurate information to answer the user's question. This allows users to check property details remotely, enabling efficient and secure property selection.
[0048] Thus, this embodiment provides users with abundant information about properties and two-way communication, making the property viewing experience more convenient and effective.
[0049] The following describes the processing flow.
[0050] Step 1:
[0051] The user selects properties of interest using the terminal's interface. The terminal then prepares to send this selection information to the server.
[0052] Step 2:
[0053] The terminal sends the user's property selection information to the server. Upon receiving this information, the server identifies the corresponding property ID.
[0054] Step 3:
[0055] The server accesses the database based on the property ID and retrieves detailed information about the property, such as 360-degree video, 3D models, and basic information.
[0056] Step 4:
[0057] The server compresses the acquired detailed information in real time and converts it into a streamable format. This data is then prepared for transmission to the user's terminal.
[0058] Step 5:
[0059] The terminal receives compressed data from the server and renders it in real time as a 3D visual or interactive map.
[0060] Step 6:
[0061] While viewing videos of the property, users can ask questions using voice commands about anything that interests them. For example, they might ask, "Which way do the windows face?"
[0062] Step 7:
[0063] The terminal receives voice input from the user and sends the voice data to the server. The server receives the voice data.
[0064] Step 8:
[0065] The server uses speech recognition technology to convert the audio data into text, analyzes the question content, and retrieves relevant data using database queries.
[0066] Step 9:
[0067] The server generates a response based on the analysis results and prepares to send it back as audio and text data.
[0068] Step 10:
[0069] The device receives the response from the server, plays it back through the speaker using speech synthesis technology, and also displays the response text on the screen. The user then uses the received information to deepen their understanding of the property.
[0070] (Example 1)
[0071] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0072] Traditional methods for remotely gathering information on rental properties and resolving questions were inefficient, requiring users to visit properties in person. This resulted in a time-consuming and laborious property selection process. Furthermore, the fragmented nature of the information limited its effectiveness in supporting decision-making.
[0073] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0074] In this invention, the server includes means for receiving selection information transmitted from a user device, means for obtaining relevant detailed information from an information collection based on the received selection information, and means for compressing the information in real time and converting it into a continuous transmission format in order to transmit the obtained detailed information to the user device. This enables users to efficiently collect detailed property information and interactively review properties even from remote locations.
[0075] "User equipment" refers to terminal devices that users operate to send and receive information.
[0076] "Selected information" refers to identifying information about the subject selected by the user.
[0077] An "information collection" refers to a data storage structure in which multiple related data are aggregated.
[0078] "Real time" refers to the instantaneous processing and communication of information.
[0079] "Continuous transmission format" refers to a communication format for transmitting data in a continuous stream without interruption.
[0080] "Voice" refers to the sound signals that users use for data input.
[0081] "Textual information" refers to data in text format that has been converted from audio or other non-textual data.
[0082] "Information retrieval" refers to the process of efficiently searching for information within a database to find the desired information.
[0083] "Sound interference removal" refers to a technique for eliminating unwanted noise from an audio signal.
[0084] "Pronunciation correction" refers to a technology that compensates for the influence of different regions and accents during the speech recognition process to ensure accurate recognition.
[0085] This system allows users to gain a detailed understanding of rental properties even from a remote location, and it functions through the collaborative efforts of the server, terminal, and user.
[0086] The server manages a database containing detailed information about properties. Upon receiving selection information from a user terminal, the server immediately retrieves detailed information about the corresponding property from the database. This information includes 360-degree video, 3D models, basic property data, and surrounding environment information. The server compresses the retrieved information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user terminal via the internet.
[0087] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail. The terminal also has a function to receive user voice input via voice recording and send it directly to the server. The server uses speech recognition technology to convert the voice into text and queries the database. Based on the converted results, the server generates answers to the user's questions and sends them back to the terminal using speech synthesis technology. The terminal plays these answers through its speaker and simultaneously displays them as text on the screen.
[0088] Users can easily search for desired properties using their devices and begin viewings. Users can ask questions verbally about areas of interest or any concerns they may have. For example, if they ask, "What is the ceiling height of this room?", the system will immediately provide accurate information. This concrete example demonstrates how the system provides users with convenience and peace of mind, supporting efficient selection of rental properties.
[0089] An example of a prompt message is, "How should a user ask a question if they want to know the size of the windows in a property?" In this way, users can obtain detailed information efficiently and in real time.
[0090] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0091] Step 1:
[0092] The user selects their desired rental property via a terminal. The user's identification information for the property selected on the interface is used as input. This information is sent from the terminal to the server. The output is the state after the property identification information has been accurately transferred to the server.
[0093] Step 2:
[0094] The server searches the database based on the received property identification information and retrieves detailed information about the corresponding property. The input is property identification information, and the output is a data object containing detailed information. A database query is executed, retrieving 360-degree video, 3D models, basic data, and surrounding environment information for the property.
[0095] Step 3:
[0096] The server compresses the acquired property details in real time and converts them into a streaming format. The input here is a detailed information data object, and the output is compressed streaming data. This operation allows the data to be efficiently transmitted to the terminal.
[0097] Step 4:
[0098] The terminal renders streaming data received from the server in real time. The input is a compressed signal, and the output is a three-dimensional visual and interactive map presented to the user. The terminal decompresses the data, converts it to a format suitable for the display device, and visualizes it.
[0099] Step 5:
[0100] Users input their questions about properties via voice through a device. The input is the user's voice, and the output is the transfer of voice data to the server. The device uses a voice recording function to accurately capture the voice.
[0101] Step 6:
[0102] The server receives audio data and converts it to text using speech recognition technology. The input is audio data, and the output is the converted text data. The speech recognition algorithm performs noise reduction and accent correction to assist in query creation.
[0103] Step 7:
[0104] The server queries the database based on the converted text data to retrieve information relevant to the user's question. The input is a text query, and the output is the query result. Data filtering and searching are performed to efficiently aggregate the necessary information.
[0105] Step 8:
[0106] The server generates an answer based on the query results and sends it to the user's terminal. The input is the query result data, and the output is the answer in voice and text format. Natural language processing is performed using a generative AI model to provide contextually appropriate answers.
[0107] Step 9:
[0108] The terminal uses speech synthesis technology to play back the responses received from the server and also displays them as text on the screen. Input consists of audio and text data, while output is auditory and visual feedback to the user. This allows the user to confirm the answers and gain a deeper understanding of the property.
[0109] (Application Example 1)
[0110] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0111] In today's real estate market, a problem exists where users living remotely have difficulty understanding properties in detail. In particular, the limited access to detailed information about a property's internal structure and surrounding environment makes it difficult for users to make accurate decisions. There is a need to solve this problem and enable users to confidently select properties from the comfort of their homes.
[0112] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0113] In this invention, the server includes means for receiving object information transmitted from a user device, means for obtaining detailed data of the corresponding object from a data storage device based on the received object information, and means for immediately compressing the acquired detailed data and converting it into a continuous playback format in order to transmit it to the user device. As a result, users can visualize detailed information of properties in real time, even from a remote location, easily search for information via voice, and obtain immediate answers.
[0114] "User device" refers to a device used by a user to receive and operate information, and includes portable information terminals such as smartphones and smart glasses.
[0115] "Target information" refers to identifying information about properties or objects for which the user wishes to obtain detailed information.
[0116] "Detailed data" refers to comprehensive information including 360-degree video and 3D models of the object, basic property data, and surrounding environment information.
[0117] A "data storage device" refers to a database used to store and manage detailed data about properties or objects.
[0118] "Means of instantly compressing and converting to a continuous playback format" refers to technical methods that compress data and convert it into a streamable format so that detailed data can be transmitted efficiently.
[0119] "Voice input" refers to instructions or questions given by the user through the microphone on their device.
[0120] "Means of converting to text information" refers to processes and technologies that include converting voice input into text data.
[0121] "Speech understanding technology" refers to technology that analyzes speech data and converts speech into text information while removing noise.
[0122] A "three-dimensional virtual environment" refers to a visual environment created using computer graphics that allows users to observe properties and objects in three dimensions.
[0123] To implement this invention, it is necessary to construct a system in which various elements, such as a server, user equipment, and speech recognition technology, work together. The specific configuration and technology are described below.
[0124] The server first receives object information transmitted from the user's device. Based on this information, it retrieves the corresponding detailed data from the data storage device. This detailed data includes 360-degree video, a three-dimensional model, basic data, and surrounding environment information related to the property. The server immediately compresses this detailed data, converts it to a continuous playback format, and transmits it to the user's device. This allows the user to view the property smoothly.
[0125] The user's device renders data received from the server in real time, displaying it as a 3D visual and interactive map. This visual environment allows the user to experience a virtual tour of the property. The user's device also has a microphone to receive user voice input and transmit it to the server.
[0126] The server uses speech recognition technology to convert received audio data into text. Here, the server performs noise reduction and accent correction to analyze the data accurately. The converted text data is then searched in a data storage device to extract relevant information, and natural language is generated using automatic generation technology to respond to the user. This generated response is then converted into audio data using speech synthesis technology and transmitted to the user's device.
[0127] For example, if a user asks, "What is the ceiling height of this room?", the server can immediately search for the information corresponding to that question and provide a natural language response in both voice and text, such as, "The ceiling height of this room is 2.8 meters." This process allows users to quickly obtain accurate and detailed property information.
[0128] An example of a prompt message is: "When a user asks a question about the property they are viewing, retrieve the relevant property data from the database and generate a response. For example, if the user asks, 'What are the dimensions of this room?', tell me the dimensions based on the relevant data." This prompt allows the generating AI model to produce an appropriate response to the user's question.
[0129] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0130] Step 1:
[0131] The user's device sends object information to the server. The input is the user's specified property information, and the output is the transmission of information to the server. In this step, the user selects properties of interest through the application interface, and this selection information is sent to the server.
[0132] Step 2:
[0133] The server receives the transmitted object information and retrieves detailed data from the data storage device. The input is the object information received from the user's device, and the output is the detailed data of the property corresponding to that information. Here, the server collects 360-degree video, 3D models, basic data, and surrounding environment information related to the selected property.
[0134] Step 3:
[0135] The server immediately compresses the acquired detailed data and converts it into a continuous playback format. The input is detailed data, and the output is compressed, continuously playable data. In this step, a data compression algorithm is executed to efficiently transmit the detailed data and make it streamable, providing the user with a seamless experience.
[0136] Step 4:
[0137] The user's device renders streaming data received from the server in real time, displaying it as a 3D visual and interactive map. The input is compressed data received from the server, and the output is a user-interactive 3D visual. The user uses this view to conduct a virtual tour and check various parts of the property.
[0138] Step 5:
[0139] The user asks a question by voice, and the user's device receives this voice input. The input is the user's voice, and the output is sent to the server as voice data. The user gives voice instructions based on their interests and questions, which are recorded through the user's device and sent to the server.
[0140] Step 6:
[0141] The server uses speech understanding technology to convert received audio data into text information. The input is audio data sent from the user's device, and the output is text data. The server performs noise reduction and pronunciation correction, and uses the latest speech recognition algorithms to accurately convert speech into text.
[0142] Step 7:
[0143] The server searches the data storage device based on textual information and retrieves relevant information. The input is textual information, and the output is relevant information retrieved by the query. Here, the appropriate data for the user's question is quickly extracted from the database.
[0144] Step 8:
[0145] Based on the information acquired by the server, a generative AI model is used to generate natural language responses. The input is information from a database, and the output is a response in natural language. To achieve this, natural language processing techniques are used to convert machine-readable data into a format that users can understand.
[0146] Step 9:
[0147] The server generates speech, which is then converted into speech data using speech synthesis technology and sent to the user's device. The input is natural language response text, and the output is speech data playable on the user's device. The technology converts text to speech so that the user can hear the response.
[0148] Step 10:
[0149] The user's device plays back the received audio data and simultaneously displays it as text on the screen. The input is audio data received from the server, and the output is information provided to the user's eyes and ears. This allows the user to obtain answers to questions through both sight and hearing.
[0150] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0151] This invention provides a system that offers detailed information about a rental property selected by the user through a virtual viewing, and also recognizes the user's emotions and adjusts the interaction accordingly.
[0152] Server-based data management and processing
[0153] The server manages a database containing property information. Upon receiving property selection information from a user's terminal, the server immediately references the database based on that information and retrieves detailed information about the corresponding property (360-degree video, 3D model, basic information, etc.).
[0154] The acquired information is compressed in real time and converted into a streaming format. This data is then ready to be sent to the terminal. Furthermore, when voice input from the user is received via the terminal, the server combines speech recognition technology and an emotion engine to convert this data into text and analyze the emotional state.
[0155] Functions of the Emotion Engine
[0156] The emotion engine analyzes the user's voice data to capture emotional nuances. This utilizes information such as voice tone, pitch, and tempo to determine whether the user is in a positive, negative, or neutral emotional state.
[0157] Based on this emotional information, the server determines how to adjust the interface and present information according to the user's emotions. For example, if the user is in a positive state, it will provide more detailed information, while if they are negative, it will focus on encouraging messages and simpler information.
[0158] User interface via terminal
[0159] The terminal decompresses the compressed data received from the server and renders it in real time as a 3D visual or interactive map. This allows the user to visually experience the property.
[0160] If necessary, the device adjusts its display and audio output to improve the user's visual experience. Based on the results of the emotion engine, the device can offer the user feedback or recommended actions.
[0161] Examples
[0162] As a concrete example, consider a case where a user asks about the size of a property. If the user's voice indicates negative emotions, the system will focus on that and emphasize encouraging or positive information. For example, it might add a statement like, "This space allows for flexible furniture arrangement."
[0163] Thus, this embodiment realizes a system that allows users to remotely check the details of rental properties and dynamically adjusts the information provided and the interface to create an experience that matches their emotions.
[0164] The following describes the processing flow.
[0165] Step 1:
[0166] The user selects the properties they want to view through their device. This selection information is then sent from the device to the server.
[0167] Step 2:
[0168] Based on the property selection information received, the server consults the database and retrieves detailed information about the corresponding property.
[0169] Step 3:
[0170] The server compresses the acquired property information in real time, converts it into a streamable format, and then prepares it for transmission to the terminal.
[0171] Step 4:
[0172] The terminal decompresses the compressed data received from the server and displays the acquired information as a 3D visual and an interactive map. At the same time, it provides a navigation function to allow the user to freely browse properties.
[0173] Step 5:
[0174] As users view properties, they can ask questions about anything that concerns them using voice. This voice input is captured on the device and transmitted to the server.
[0175] Step 6:
[0176] The server receives voice input and converts the speech into text using speech recognition technology, while simultaneously analyzing the user's emotional state with an emotion engine.
[0177] Step 7:
[0178] Based on the transcribed questions, the server executes database queries to retrieve relevant information and determines the appropriate response style based on the analyzed sentiment data.
[0179] Step 8:
[0180] The server performs natural language processing in a tone and style that matches the user's emotions, generates a response optimized for the user, and sends it to the terminal.
[0181] Step 9:
[0182] The device plays back the response from the server using speech synthesis technology and simultaneously displays it on the screen as text. This allows the user to receive rich information and advice based on their responses.
[0183] (Example 2)
[0184] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0185] Traditional real estate information systems provide information without considering the user's emotional state, resulting in a uniform user experience that fails to adequately address individual needs. Furthermore, the lack of sufficient visual presentation of property information and limited opportunities for active interaction are also problematic.
[0186] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0187] In this invention, the server includes means for receiving property selection information transmitted from a user terminal, means for obtaining detailed information of the relevant property from a data storage medium based on the received property selection information, means for compressing the data in real time and converting it into a data streaming format in order to transmit the obtained detailed information to the user terminal, means for receiving voice input from the user and converting the voice data into text, means for performing sentiment analysis from the converted text, and means for dynamically adjusting the interface and changing the method of presenting information according to the user's emotional state. This makes it possible to provide information and adjust the interface to match the user's emotional state, thereby realizing a more personalized user experience.
[0188] A "user terminal" is an electronic device used by a user to input or output information.
[0189] "Property selection information" refers to data that represents the user's selections regarding real estate properties they are interested in.
[0190] A "data storage medium" is a physical or virtual device used to store digital information.
[0191] "Detailed information" refers to information about various attributes and characteristics of a real estate property.
[0192] A "data streaming format" is a data format for transmitting digital data continuously in real time.
[0193] "Voice input" refers to data of words and sounds spoken by the user.
[0194] "Text conversion" is the process of representing audio data as text.
[0195] "Emotional analysis" is an analytical technique used to identify emotional states from voice and other data.
[0196] An "interface" is a means or method for a user to exchange information with a system.
[0197] This invention is a system that allows users to remotely view details of real estate properties and provides interactions that take into account the user's emotional state. The system mainly consists of a server and a user terminal, each playing a specific role.
[0198] The server manages the data storage system and stores detailed information about properties. Based on the property selection information received from the user's terminal, the server quickly retrieves the necessary property details from the data storage, compresses the data in real time, and converts it into a data streaming format. Furthermore, the server uses speech recognition technology and an emotion analysis engine to convert the user's voice input into text and analyze their emotional state. Emotion analysis utilizes voice features such as tone, tempo, and pitch to determine whether the user is feeling positive, negative, or neutral. Based on these analysis results, the server dynamically adjusts the interface and presents information according to the user's emotional state.
[0199] The device decompresses compressed data sent from the server and renders it as a spatial digital representation in real time. This allows users to visually experience the property through 3D visuals and interactive maps. The device also adjusts the display and audio output to optimize the user's visual experience. Based on the results of sentiment analysis, the device provides feedback and recommended actions to the user, helping them to fully understand the value of the property.
[0200] As a concrete example, consider a scenario where a user asks a voice question about the size of a property, such as, "How big is the living room in this property?" If the device detects that the user's voice is expressing negative emotions, it will provide positive feedback such as, "This living room has large windows and gets plenty of light." This allows the user to focus on the positive features of the property.
[0201] An example of a prompt message might be: "What are the key points the user wants to know about this rental property? How should that information be presented, depending on the user's emotional state?"
[0202] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0203] Step 1:
[0204] Users select properties of interest via their device. Information about the selected properties is sent from the device to the server as "property selection information." This input information includes the property ID and the reason the user is interested in the property.
[0205] Step 2:
[0206] The server references the data storage system based on the received property selection information. The data processing performed here involves retrieving detailed information about the relevant property using queries based on the property ID. The output includes detailed information such as 360-degree video, 3D models, and basic property information.
[0207] Step 3:
[0208] The server compresses the acquired detailed information in real time and converts it into a streaming format. This data processing improves transmission efficiency. The output is a compressed data stream, ready to be sent to the user's terminal.
[0209] Step 4:
[0210] The user enters specific questions about the property via voice. This voice data is sent from the terminal to the server. The entered voice file is received on the server side.
[0211] Step 5:
[0212] The server analyzes the received audio data. Using speech recognition technology, it converts the audio to text, and then uses an emotion analysis engine to analyze the emotional state. Through data calculations, it determines whether the user is in a positive, negative, or neutral emotional state. The output consists of the converted text and the emotional state evaluation result.
[0213] Step 6:
[0214] The server adjusts the user interface based on the sentiment analysis results. If the sentiment is positive, it provides detailed information; if it's negative, it generates supplementary information and encouraging messages. This optimizes the way information is presented and its content. The output is the adjusted information presentation.
[0215] Step 7:
[0216] The terminal decompresses the compressed data sent from the server and renders it in real time as a spatial digital representation. Data processing is performed for display as 3D visuals or interactive maps. The output is a display that the user can visually observe.
[0217] Step 8:
[0218] The device presents the user with feedback and recommended actions based on the results of sentiment analysis. For example, positive feedback such as "This living room has large windows and gets plenty of light" might be displayed. The output consists of content and recommended actions presented to the user.
[0219] (Application Example 2)
[0220] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0221] Conventional virtual viewing systems simply provide property information, lacking dynamic interaction that responds to user emotions. Furthermore, because information is not provided in a way that considers user feelings, the user experience is uniform, making it difficult to enhance individual satisfaction.
[0222] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0223] In this invention, the server includes means for receiving real estate information transmitted from a user device, means for obtaining detailed information of the relevant real estate from a recording device based on the received real estate information, means for compressing the data in real time and converting it into a streaming format in order to transmit the obtained detailed information to the user device, means for analyzing the user's voice data using emotion recognition technology to determine the emotional state, and means for adjusting the displayed information based on the user's emotions. This enables the provision of dynamic information tailored to the user's emotions, resulting in a more satisfying virtual viewing experience.
[0224] A "user device" is a terminal device that a user operates to receive and transmit information.
[0225] "Real estate information" refers to detailed data about rental properties, including information such as location, floor plan, and amenities.
[0226] A "recording device" is a database system used to store and manage detailed property information.
[0227] "Streaming format" refers to a data transmission method for playing back data while transferring it in real time.
[0228] "Emotion recognition technology" is a technology that analyzes a user's emotions from voice data and image data to determine their emotional state, such as positive, negative, or neutral.
[0229] "Dynamic information provision" refers to a system that changes the data displayed and the answers provided in real time according to the user's current state and needs.
[0230] In order to implement this invention, the following system configuration and processing procedure are required.
[0231] First, the server receives property information from the user's device. Based on the received information, it retrieves detailed information about the relevant property from the recording device. This detailed information includes a 360-degree view and a 3D model of the property. The server compresses this information in real time and converts it into a streaming format. After that, it prepares it for transmission to the user's device.
[0232] The user device, acting as the terminal, decompresses received compressed data and renders it in real time to provide a three-dimensional visual output. It also manages user interaction through voice and eye-tracking input and uses emotion recognition technology to determine the user's emotional state. This technology includes software that analyzes voice tone and speed. Based on emotions, it dynamically adjusts the content and presentation of displayed information.
[0233] This system primarily operates using wearable devices such as smart glasses, and determines the user's emotional state based on their voice. For example, if a user expresses dissatisfaction by saying, "This room is small," it can then provide additional positive information such as, "Although there is limited space, you can use it more effectively depending on how you arrange the shelves."
[0234] As a concrete example, consider a scenario where a user wears smart glasses and requests to "show me the bathroom." This system displays a virtual view of the bathroom on the glasses and performs emotion recognition. Even if the user is disappointed, it provides encouraging information such as, "There is a relaxing bathtub installed."
[0235] A concrete example of a prompt message would be: "Design an application that virtually provides viewing information about a property selected by the user, analyzes the emotion in their voice, and provides feedback accordingly." In this way, the user can analyze and evaluate properties in real time in great detail.
[0236] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0237] Step 1:
[0238] The server retrieves real estate information received from the user's device. This real estate information includes identification information for the property selected by the user. Based on this input data, the server queries the recording device to retrieve detailed information about the relevant property. The retrieved detailed information includes 360-degree images and a three-dimensional model of the property.
[0239] Step 2:
[0240] The server compresses the acquired property details and converts them into a real-time streaming format. This process encodes the data to efficiently utilize communication bandwidth and prepares it for transmission to the user's device. The streaming data is then sent to the transmission queue as output.
[0241] Step 3:
[0242] The user device receives compressed streaming data sent from the server. This input data is decompressed and rendered in real time. Specifically, it provides the user with a visual output in the form of a three-dimensional image and an interactive map. Through the device, the user can virtually tour the property.
[0243] Step 4:
[0244] Users ask questions and make comments about properties via voice input. Once the voice data is input to the user's device, the terminal uses a speech recognition engine to convert it into text, and then uses emotion recognition technology to analyze the user's emotions. This analysis is performed by analyzing acoustic features, including voice tone and speed. The analysis results in the output of emotional state data.
[0245] Step 5:
[0246] The server dynamically adjusts the information displayed based on the user's emotional state. For positive emotions, it provides detailed information; for negative emotions, it generates feedback such as supplementary information or encouraging messages. This adjusted information is then sent to the user's device.
[0247] Step 6:
[0248] The user device receives pre-configured information from the server and presents it to the user according to a predetermined display method. Audio and visual output facilitates a deeper understanding of the property and a more positive viewing experience. Further interactions proceed based on the user's subsequent actions.
[0249] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0250] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0251] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0252] [Second Embodiment]
[0253] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0254] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0255] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0256] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0257] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0258] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0259] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0260] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0261] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0262] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0263] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0264] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0265] This invention provides a system that allows users to understand rental properties in detail and resolve any questions they may have, even from a remote location. This system functions through the collaborative efforts of a server, a terminal, and the user.
[0266] Server-based data management and processing
[0267] The server manages a database containing detailed information about properties. When it receives property selection information from a user's terminal, the server immediately retrieves the corresponding property information from the database. The retrieved information includes 360-degree video, 3D models, basic property data, and surrounding environment information.
[0268] The server compresses this information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user's terminal over the internet. When the server receives voice input from the user, it uses speech recognition technology to convert the voice into text data and queries the database. Based on the results, the server generates and returns an answer to the user's question.
[0269] User interface via terminal
[0270] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail.
[0271] Furthermore, the device receives the user's voice input through its voice recording function and sends it directly to the server. The response sent back from the server is played back through the speaker using speech synthesis technology and simultaneously displayed as text on the device screen.
[0272] User actions
[0273] Users can easily search for their desired properties and begin virtual tours using their devices. Specifically, they can view property information and videos displayed on the interface and ask the system questions verbally about areas of interest or anything they are unsure of.
[0274] For example, if a user asks, "What is the ceiling height of this room?", the system will immediately provide accurate information to answer the user's question. This allows users to check property details remotely, enabling efficient and secure property selection.
[0275] In this way, this embodiment provides rich information about the property and two-way communication to the user, making the in-house viewing experience more convenient and effective.
[0276] The following describes the processing flow.
[0277] Step 1:
[0278] The user selects a property of interest through the interface of the terminal and prepares to send the selection information to the server.
[0279] Step 2:
[0280] The terminal sends the user's property selection information to the server. When the server receives this information, it identifies the corresponding property ID.
[0281] Step 3:
[0282] Based on the property ID, the server accesses the database to obtain detailed information about the corresponding property, such as 360-degree videos, 3D models, and basic information.
[0283] Step 4:
[0284] The server compresses the obtained detailed information in real time and converts it into a streaming-compatible format. This data is prepared to be sent to the user terminal.
[0285] Step 5:
[0286] The terminal receives the compressed data from the server and renders it in real time as a three-dimensional visual or an interactive map.
[0287] Step 6:
[0288] While viewing the video of the property, the user asks questions about the points of concern in voice. For example, set "What is the direction of the window?"
[0289] Step 7:
[0290] The terminal receives voice input from the user and sends the voice data to the server. The server receives the voice data.
[0291] Step 8:
[0292] The server uses speech recognition technology to convert the audio data into text, analyzes the question content, and retrieves relevant data using database queries.
[0293] Step 9:
[0294] The server generates a response based on the analysis results and prepares to send it back as audio and text data.
[0295] Step 10:
[0296] The device receives the response from the server, plays it back through the speaker using speech synthesis technology, and also displays the response text on the screen. The user then uses the received information to deepen their understanding of the property.
[0297] (Example 1)
[0298] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0299] Traditional methods for remotely gathering information on rental properties and resolving questions were inefficient, requiring users to visit properties in person. This resulted in a time-consuming and laborious property selection process. Furthermore, the fragmented nature of the information limited its effectiveness in supporting decision-making.
[0300] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0301] In this invention, the server includes means for receiving selection information transmitted from a user device, means for obtaining corresponding detailed information from an information collection based on the received selection information, and means for compressing the information in real time and converting it into a continuous transmission format in order to transmit the obtained detailed information to the user device. As a result, the user can efficiently collect detailed property information even from a remote location and interactively consider properties.
[0302] The "user device" refers to a terminal device used by a user to operate and transmit / receive information.
[0303] The "selection information" refers to identification information regarding the object selected by the user.
[0304] The "information collection" refers to a data storage structure in which a plurality of related data are aggregated.
[0305] "Real time" means that information processing and communication are performed immediately.
[0306] The "continuous transmission format" refers to a communication format for transmitting data in a continuous stream without interruption.
[0307] "Voice" refers to a sound signal used by a user for data input.
[0308] "Character information" refers to data in text format obtained by converting voice and other non-character data.
[0309] "Information collection search" refers to the process of efficiently searching for information in a database to find the target information.
[0310] "Noise mixing removal" refers to a technology for eliminating unwanted noise from a voice signal.
[0311] "Pronunciation correction" refers to a technology for accurately recognizing speech by compensating for the effects of different regions and accents in the speech recognition process.
[0312] This system allows users to gain a detailed understanding of rental properties even from a remote location, and it functions through the collaborative efforts of the server, terminal, and user.
[0313] The server manages a database containing detailed information about properties. Upon receiving selection information from a user terminal, the server immediately retrieves detailed information about the corresponding property from the database. This information includes 360-degree video, 3D models, basic property data, and surrounding environment information. The server compresses the retrieved information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user terminal via the internet.
[0314] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail. The terminal also has a function to receive user voice input via voice recording and send it directly to the server. The server uses speech recognition technology to convert the voice into text and queries the database. Based on the converted results, the server generates answers to the user's questions and sends them back to the terminal using speech synthesis technology. The terminal plays these answers through its speaker and simultaneously displays them as text on the screen.
[0315] Users can easily search for desired properties using their devices and begin viewings. Users can ask questions verbally about areas of interest or any concerns they may have. For example, if they ask, "What is the ceiling height of this room?", the system will immediately provide accurate information. This concrete example demonstrates how the system provides users with convenience and peace of mind, supporting efficient selection of rental properties.
[0316] An example of a prompt message is, "How should a user ask a question if they want to know the size of the windows in a property?" In this way, users can obtain detailed information efficiently and in real time.
[0317] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0318] Step 1:
[0319] The user selects their desired rental property via a terminal. The user's identification information for the property selected on the interface is used as input. This information is sent from the terminal to the server. The output is the state after the property identification information has been accurately transferred to the server.
[0320] Step 2:
[0321] The server searches the database based on the received property identification information and retrieves detailed information about the corresponding property. The input is property identification information, and the output is a data object containing detailed information. A database query is executed, retrieving 360-degree video, 3D models, basic data, and surrounding environment information for the property.
[0322] Step 3:
[0323] The server compresses the acquired property details in real time and converts them into a streaming format. The input here is a detailed information data object, and the output is compressed streaming data. This operation allows the data to be efficiently transmitted to the terminal.
[0324] Step 4:
[0325] The terminal renders streaming data received from the server in real time. The input is a compressed signal, and the output is a three-dimensional visual and interactive map presented to the user. The terminal decompresses the data, converts it to a format suitable for the display device, and visualizes it.
[0326] Step 5:
[0327] Users input their questions about properties via voice through a device. The input is the user's voice, and the output is the transfer of voice data to the server. The device uses a voice recording function to accurately capture the voice.
[0328] Step 6:
[0329] The server receives audio data and converts it to text using speech recognition technology. The input is audio data, and the output is the converted text data. The speech recognition algorithm performs noise reduction and accent correction to assist in query creation.
[0330] Step 7:
[0331] The server queries the database based on the converted text data to retrieve information relevant to the user's question. The input is a text query, and the output is the query result. Data filtering and searching are performed to efficiently aggregate the necessary information.
[0332] Step 8:
[0333] The server generates an answer based on the query results and sends it to the user's terminal. The input is the query result data, and the output is the answer in voice and text format. Natural language processing is performed using a generative AI model to provide contextually appropriate answers.
[0334] Step 9:
[0335] The terminal uses speech synthesis technology to play back the responses received from the server and also displays them as text on the screen. Input consists of audio and text data, while output is auditory and visual feedback to the user. This allows the user to confirm the answers and gain a deeper understanding of the property.
[0336] (Application Example 1)
[0337] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0338] In today's real estate market, a problem exists where users living remotely have difficulty understanding properties in detail. In particular, the limited access to detailed information about a property's internal structure and surrounding environment makes it difficult for users to make accurate decisions. There is a need to solve this problem and enable users to confidently select properties from the comfort of their homes.
[0339] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0340] In this invention, the server includes means for receiving object information transmitted from a user device, means for obtaining detailed data of the corresponding object from a data storage device based on the received object information, and means for immediately compressing the acquired detailed data and converting it into a continuous playback format in order to transmit it to the user device. As a result, users can visualize detailed information of properties in real time, even from a remote location, easily search for information via voice, and obtain immediate answers.
[0341] "User device" refers to a device used by a user to receive and operate information, and includes portable information terminals such as smartphones and smart glasses.
[0342] "Target information" refers to identifying information about properties or objects for which the user wishes to obtain detailed information.
[0343] "Detailed data" refers to comprehensive information including 360-degree video and 3D models of the object, basic property data, and surrounding environment information.
[0344] A "data storage device" refers to a database used to store and manage detailed data about properties or objects.
[0345] "Means of instantly compressing and converting to a continuous playback format" refers to technical methods that compress data and convert it into a streamable format so that detailed data can be transmitted efficiently.
[0346] "Voice input" refers to instructions or questions given by the user through the microphone on their device.
[0347] "Means of converting to text information" refers to processes and technologies that include converting voice input into text data.
[0348] "Speech understanding technology" refers to technology that analyzes speech data and converts speech into text information while removing noise.
[0349] A "three-dimensional virtual environment" refers to a visual environment created using computer graphics that allows users to observe properties and objects in three dimensions.
[0350] To implement this invention, it is necessary to construct a system in which various elements, such as a server, user equipment, and speech recognition technology, work together. The specific configuration and technology are described below.
[0351] The server first receives object information transmitted from the user's device. Based on this information, it retrieves the corresponding detailed data from the data storage device. This detailed data includes 360-degree video, a three-dimensional model, basic data, and surrounding environment information related to the property. The server immediately compresses this detailed data, converts it to a continuous playback format, and transmits it to the user's device. This allows the user to view the property smoothly.
[0352] The user's device renders data received from the server in real time, displaying it as a 3D visual and interactive map. This visual environment allows the user to experience a virtual tour of the property. The user's device also has a microphone to receive user voice input and transmit it to the server.
[0353] The server uses speech recognition technology to convert received audio data into text. Here, the server performs noise reduction and accent correction to analyze the data accurately. The converted text data is then searched in a data storage device to extract relevant information, and natural language is generated using automatic generation technology to respond to the user. This generated response is then converted into audio data using speech synthesis technology and transmitted to the user's device.
[0354] For example, if a user asks, "What is the ceiling height of this room?", the server can immediately search for the information corresponding to that question and provide a natural language response in both voice and text, such as, "The ceiling height of this room is 2.8 meters." This process allows users to quickly obtain accurate and detailed property information.
[0355] An example of a prompt message is: "When a user asks a question about the property they are viewing, retrieve the relevant property data from the database and generate a response. For example, if the user asks, 'What are the dimensions of this room?', tell me the dimensions based on the relevant data." This prompt allows the generating AI model to produce an appropriate response to the user's question.
[0356] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0357] Step 1:
[0358] The user's device sends object information to the server. The input is the user's specified property information, and the output is the transmission of information to the server. In this step, the user selects properties of interest through the application interface, and this selection information is sent to the server.
[0359] Step 2:
[0360] The server receives the transmitted object information and retrieves detailed data from the data storage device. The input is the object information received from the user's device, and the output is the detailed data of the property corresponding to that information. Here, the server collects 360-degree video, 3D models, basic data, and surrounding environment information related to the selected property.
[0361] Step 3:
[0362] The server immediately compresses the acquired detailed data and converts it into a continuous playback format. The input is detailed data, and the output is compressed, continuously playable data. In this step, a data compression algorithm is executed to efficiently transmit the detailed data and make it streamable, providing the user with a seamless experience.
[0363] Step 4:
[0364] The user's device renders streaming data received from the server in real time, displaying it as a 3D visual and interactive map. The input is compressed data received from the server, and the output is a user-interactive 3D visual. The user uses this view to conduct a virtual tour and check various parts of the property.
[0365] Step 5:
[0366] The user asks a question by voice, and the user's device receives this voice input. The input is the user's voice, and the output is sent to the server as voice data. The user gives voice instructions based on their interests and questions, which are recorded through the user's device and sent to the server.
[0367] Step 6:
[0368] The server uses speech understanding technology to convert received audio data into text information. The input is audio data sent from the user's device, and the output is text data. The server performs noise reduction and pronunciation correction, and uses the latest speech recognition algorithms to accurately convert speech into text.
[0369] Step 7:
[0370] The server searches the data storage device based on textual information and retrieves relevant information. The input is textual information, and the output is relevant information retrieved by the query. Here, the appropriate data for the user's question is quickly extracted from the database.
[0371] Step 8:
[0372] Based on the information acquired by the server, a generative AI model is used to generate natural language responses. The input is information from a database, and the output is a response in natural language. To achieve this, natural language processing techniques are used to convert machine-readable data into a format that users can understand.
[0373] Step 9:
[0374] The server generates speech, which is then converted into speech data using speech synthesis technology and sent to the user's device. The input is natural language response text, and the output is speech data playable on the user's device. The technology converts text to speech so that the user can hear the response.
[0375] Step 10:
[0376] The user's device plays back the received audio data and simultaneously displays it as text on the screen. The input is audio data received from the server, and the output is information provided to the user's eyes and ears. This allows the user to obtain answers to questions through both sight and hearing.
[0377] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0378] This invention provides a system that offers detailed information about a rental property selected by the user through a virtual viewing, and also recognizes the user's emotions and adjusts the interaction accordingly.
[0379] Server-based data management and processing
[0380] The server manages a database containing property information. Upon receiving property selection information from a user's terminal, the server immediately references the database based on that information and retrieves detailed information about the corresponding property (360-degree video, 3D model, basic information, etc.).
[0381] The acquired information is compressed in real time and converted into a streaming format. This data is then ready to be sent to the terminal. Furthermore, when voice input from the user is received via the terminal, the server combines speech recognition technology and an emotion engine to convert this data into text and analyze the emotional state.
[0382] Functions of the Emotion Engine
[0383] The emotion engine analyzes the user's voice data to capture emotional nuances. This utilizes information such as voice tone, pitch, and tempo to determine whether the user is in a positive, negative, or neutral emotional state.
[0384] Based on this emotional information, the server determines how to adjust the interface and present information according to the user's emotions. For example, if the user is in a positive state, it will provide more detailed information, while if they are negative, it will focus on encouraging messages and simpler information.
[0385] User interface via terminal
[0386] The terminal decompresses the compressed data received from the server and renders it in real time as a 3D visual or interactive map. This allows the user to visually experience the property.
[0387] If necessary, the device adjusts its display and audio output to improve the user's visual experience. Based on the results of the emotion engine, the device can offer the user feedback or recommended actions.
[0388] Examples
[0389] As a concrete example, consider a case where a user asks about the size of a property. If the user's voice indicates negative emotions, the system will focus on that and emphasize encouraging or positive information. For example, it might add a statement like, "This space allows for flexible furniture arrangement."
[0390] Thus, this embodiment realizes a system that allows users to remotely check the details of rental properties and dynamically adjusts the information provided and the interface to create an experience that matches their emotions.
[0391] The following describes the processing flow.
[0392] Step 1:
[0393] The user selects the properties they want to view through their device. This selection information is then sent from the device to the server.
[0394] Step 2:
[0395] Based on the property selection information received, the server consults the database and retrieves detailed information about the corresponding property.
[0396] Step 3:
[0397] The server compresses the acquired property information in real time, converts it into a streamable format, and then prepares it for transmission to the terminal.
[0398] Step 4:
[0399] The terminal decompresses the compressed data received from the server and displays the acquired information as a 3D visual and an interactive map. At the same time, it provides a navigation function to allow the user to freely browse properties.
[0400] Step 5:
[0401] As users view properties, they can ask questions about anything that concerns them using voice. This voice input is captured on the device and transmitted to the server.
[0402] Step 6:
[0403] The server receives voice input and converts the speech into text using speech recognition technology, while simultaneously analyzing the user's emotional state with an emotion engine.
[0404] Step 7:
[0405] Based on the transcribed questions, the server executes database queries to retrieve relevant information and determines the appropriate response style based on the analyzed sentiment data.
[0406] Step 8:
[0407] The server performs natural language processing in a tone and style that matches the user's emotions, generates a response optimized for the user, and sends it to the terminal.
[0408] Step 9:
[0409] The device plays back the response from the server using speech synthesis technology and simultaneously displays it on the screen as text. This allows the user to receive rich information and advice based on their responses.
[0410] (Example 2)
[0411] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0412] Traditional real estate information systems provide information without considering the user's emotional state, resulting in a uniform user experience that fails to adequately address individual needs. Furthermore, the lack of sufficient visual presentation of property information and limited opportunities for active interaction are also problematic.
[0413] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0414] In this invention, the server includes means for receiving property selection information transmitted from a user terminal, means for obtaining detailed information of the relevant property from a data storage medium based on the received property selection information, means for compressing the data in real time and converting it into a data streaming format in order to transmit the obtained detailed information to the user terminal, means for receiving voice input from the user and converting the voice data into text, means for performing sentiment analysis from the converted text, and means for dynamically adjusting the interface and changing the method of presenting information according to the user's emotional state. This makes it possible to provide information and adjust the interface to match the user's emotional state, thereby realizing a more personalized user experience.
[0415] A "user terminal" is an electronic device used by a user to input or output information.
[0416] "Property selection information" refers to data that represents the user's selections regarding real estate properties they are interested in.
[0417] A "data storage medium" is a physical or virtual device used to store digital information.
[0418] "Detailed information" refers to information about various attributes and characteristics of a real estate property.
[0419] A "data streaming format" is a data format for transmitting digital data continuously in real time.
[0420] "Voice input" refers to data of words and sounds spoken by the user.
[0421] "Text conversion" is the process of representing audio data as text.
[0422] "Emotional analysis" is an analytical technique used to identify emotional states from voice and other data.
[0423] An "interface" is a means or method for a user to exchange information with a system.
[0424] This invention is a system that allows users to remotely view details of real estate properties and provides interactions that take into account the user's emotional state. The system mainly consists of a server and a user terminal, each playing a specific role.
[0425] The server manages the data storage system and stores detailed information about properties. Based on the property selection information received from the user's terminal, the server quickly retrieves the necessary property details from the data storage, compresses the data in real time, and converts it into a data streaming format. Furthermore, the server uses speech recognition technology and an emotion analysis engine to convert the user's voice input into text and analyze their emotional state. Emotion analysis utilizes voice features such as tone, tempo, and pitch to determine whether the user is feeling positive, negative, or neutral. Based on these analysis results, the server dynamically adjusts the interface and presents information according to the user's emotional state.
[0426] The device decompresses compressed data sent from the server and renders it as a spatial digital representation in real time. This allows users to visually experience the property through 3D visuals and interactive maps. The device also adjusts the display and audio output to optimize the user's visual experience. Based on the results of sentiment analysis, the device provides feedback and recommended actions to the user, helping them to fully understand the value of the property.
[0427] As a concrete example, consider a scenario where a user asks a voice question about the size of a property, such as, "How big is the living room in this property?" If the device detects that the user's voice is expressing negative emotions, it will provide positive feedback such as, "This living room has large windows and gets plenty of light." This allows the user to focus on the positive features of the property.
[0428] An example of a prompt message might be: "What are the key points the user wants to know about this rental property? How should that information be presented, depending on the user's emotional state?"
[0429] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0430] Step 1:
[0431] Users select properties of interest via their device. Information about the selected properties is sent from the device to the server as "property selection information." This input information includes the property ID and the reason the user is interested in the property.
[0432] Step 2:
[0433] The server references the data storage system based on the received property selection information. The data processing performed here involves retrieving detailed information about the relevant property using queries based on the property ID. The output includes detailed information such as 360-degree video, 3D models, and basic property information.
[0434] Step 3:
[0435] The server compresses the acquired detailed information in real time and converts it into a streaming format. This data processing improves transmission efficiency. The output is a compressed data stream, ready to be sent to the user's terminal.
[0436] Step 4:
[0437] The user enters specific questions about the property via voice. This voice data is sent from the terminal to the server. The entered voice file is received on the server side.
[0438] Step 5:
[0439] The server analyzes the received audio data. Using speech recognition technology, it converts the audio to text, and then uses an emotion analysis engine to analyze the emotional state. Through data calculations, it determines whether the user is in a positive, negative, or neutral emotional state. The output consists of the converted text and the emotional state evaluation result.
[0440] Step 6:
[0441] The server adjusts the user interface based on the sentiment analysis results. If the sentiment is positive, it provides detailed information; if it's negative, it generates supplementary information and encouraging messages. This optimizes the way information is presented and its content. The output is the adjusted information presentation.
[0442] Step 7:
[0443] The terminal decompresses the compressed data sent from the server and renders it in real time as a spatial digital representation. Data processing is performed for display as 3D visuals or interactive maps. The output is a display that the user can visually observe.
[0444] Step 8:
[0445] The device presents the user with feedback and recommended actions based on the results of sentiment analysis. For example, positive feedback such as "This living room has large windows and gets plenty of light" might be displayed. The output consists of content and recommended actions presented to the user.
[0446] (Application Example 2)
[0447] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0448] Conventional virtual viewing systems simply provide property information, lacking dynamic interaction that responds to user emotions. Furthermore, because information is not provided in a way that considers user feelings, the user experience is uniform, making it difficult to enhance individual satisfaction.
[0449] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0450] In this invention, the server includes means for receiving real estate information transmitted from a user device, means for obtaining detailed information of the relevant real estate from a recording device based on the received real estate information, means for compressing the data in real time and converting it into a streaming format in order to transmit the obtained detailed information to the user device, means for analyzing the user's voice data using emotion recognition technology to determine the emotional state, and means for adjusting the displayed information based on the user's emotions. This enables the provision of dynamic information tailored to the user's emotions, resulting in a more satisfying virtual viewing experience.
[0451] A "user device" is a terminal device that a user operates to receive and transmit information.
[0452] "Real estate information" refers to detailed data about rental properties, including information such as location, floor plan, and amenities.
[0453] A "recording device" is a database system used to store and manage detailed property information.
[0454] "Streaming format" refers to a data transmission method for playing back data while transferring it in real time.
[0455] "Emotion recognition technology" is a technology that analyzes a user's emotions from voice data and image data to determine their emotional state, such as positive, negative, or neutral.
[0456] "Dynamic information provision" refers to a system that changes the data displayed and the answers provided in real time according to the user's current state and needs.
[0457] In order to implement this invention, the following system configuration and processing procedure are required.
[0458] First, the server receives property information from the user's device. Based on the received information, it retrieves detailed information about the relevant property from the recording device. This detailed information includes a 360-degree view and a 3D model of the property. The server compresses this information in real time and converts it into a streaming format. After that, it prepares it for transmission to the user's device.
[0459] The user device, acting as the terminal, decompresses received compressed data and renders it in real time to provide a three-dimensional visual output. It also manages user interaction through voice and eye-tracking input and uses emotion recognition technology to determine the user's emotional state. This technology includes software that analyzes voice tone and speed. Based on emotions, it dynamically adjusts the content and presentation of displayed information.
[0460] This system primarily operates using wearable devices such as smart glasses, and determines the user's emotional state based on their voice. For example, if a user expresses dissatisfaction by saying, "This room is small," it can then provide additional positive information such as, "Although there is limited space, you can use it more effectively depending on how you arrange the shelves."
[0461] As a concrete example, consider a scenario where a user wears smart glasses and requests to "show me the bathroom." This system displays a virtual view of the bathroom on the glasses and performs emotion recognition. Even if the user is disappointed, it provides encouraging information such as, "There is a relaxing bathtub installed."
[0462] A concrete example of a prompt message would be: "Design an application that virtually provides viewing information about a property selected by the user, analyzes the emotion in their voice, and provides feedback accordingly." In this way, the user can analyze and evaluate properties in real time in great detail.
[0463] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0464] Step 1:
[0465] The server retrieves real estate information received from the user's device. This real estate information includes identification information for the property selected by the user. Based on this input data, the server queries the recording device to retrieve detailed information about the relevant property. The retrieved detailed information includes 360-degree images and a three-dimensional model of the property.
[0466] Step 2:
[0467] The server compresses the acquired property details and converts them into a real-time streaming format. This process encodes the data to efficiently utilize communication bandwidth and prepares it for transmission to the user's device. The streaming data is then sent to the transmission queue as output.
[0468] Step 3:
[0469] The user device receives compressed streaming data sent from the server. This input data is decompressed and rendered in real time. Specifically, it provides the user with a visual output in the form of a three-dimensional image and an interactive map. Through the device, the user can virtually tour the property.
[0470] Step 4:
[0471] Users ask questions and make comments about properties via voice input. Once the voice data is input to the user's device, the terminal uses a speech recognition engine to convert it into text, and then uses emotion recognition technology to analyze the user's emotions. This analysis is performed by analyzing acoustic features, including voice tone and speed. The analysis results in the output of emotional state data.
[0472] Step 5:
[0473] The server dynamically adjusts the information displayed based on the user's emotional state. For positive emotions, it provides detailed information; for negative emotions, it generates feedback such as supplementary information or encouraging messages. This adjusted information is then sent to the user's device.
[0474] Step 6:
[0475] The user device receives pre-configured information from the server and presents it to the user according to a predetermined display method. Audio and visual output facilitates a deeper understanding of the property and a more positive viewing experience. Further interactions proceed based on the user's subsequent actions.
[0476] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0477] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0478] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0479] [Third Embodiment]
[0480] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0481] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0482] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0483] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0484] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0485] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0486] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0487] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0488] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0489] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0490] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0491] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0492] This invention provides a system that allows users to understand rental properties in detail and resolve any questions they may have, even from a remote location. This system functions through the collaborative efforts of a server, a terminal, and the user.
[0493] Server-based data management and processing
[0494] The server manages a database containing detailed information about properties. When it receives property selection information from a user's terminal, the server immediately retrieves the corresponding property information from the database. The retrieved information includes 360-degree video, 3D models, basic property data, and surrounding environment information.
[0495] The server compresses this information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user's terminal over the internet. When the server receives voice input from the user, it uses speech recognition technology to convert the voice into text data and queries the database. Based on the results, the server generates and returns an answer to the user's question.
[0496] User interface via terminal
[0497] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail.
[0498] Furthermore, the device receives the user's voice input through its voice recording function and sends it directly to the server. The response sent back from the server is played back through the speaker using speech synthesis technology and simultaneously displayed as text on the device screen.
[0499] User actions
[0500] Users can easily search for their desired properties and begin virtual tours using their devices. Specifically, they can view property information and videos displayed on the interface and ask the system questions verbally about areas of interest or anything they are unsure of.
[0501] For example, if a user asks, "What is the ceiling height of this room?", the system will immediately provide accurate information to answer the user's question. This allows users to check property details remotely, enabling efficient and secure property selection.
[0502] Thus, this embodiment provides users with abundant information about properties and two-way communication, making the property viewing experience more convenient and effective.
[0503] The following describes the processing flow.
[0504] Step 1:
[0505] The user selects properties of interest using the terminal's interface. The terminal then prepares to send this selection information to the server.
[0506] Step 2:
[0507] The terminal sends the user's property selection information to the server. Upon receiving this information, the server identifies the corresponding property ID.
[0508] Step 3:
[0509] The server accesses the database based on the property ID and retrieves detailed information about the property, such as 360-degree video, 3D models, and basic information.
[0510] Step 4:
[0511] The server compresses the acquired detailed information in real time and converts it into a streamable format. This data is then prepared for transmission to the user's terminal.
[0512] Step 5:
[0513] The terminal receives compressed data from the server and renders it in real time as a 3D visual or interactive map.
[0514] Step 6:
[0515] While viewing videos of the property, users can ask questions using voice commands about anything that interests them. For example, they might ask, "Which way do the windows face?"
[0516] Step 7:
[0517] The terminal receives voice input from the user and sends the voice data to the server. The server receives the voice data.
[0518] Step 8:
[0519] The server uses speech recognition technology to convert the audio data into text, analyzes the question content, and retrieves relevant data using database queries.
[0520] Step 9:
[0521] The server generates a response based on the analysis results and prepares to send it back as audio and text data.
[0522] Step 10:
[0523] The device receives the response from the server, plays it back through the speaker using speech synthesis technology, and also displays the response text on the screen. The user then uses the received information to deepen their understanding of the property.
[0524] (Example 1)
[0525] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0526] Traditional methods for remotely gathering information on rental properties and resolving questions were inefficient, requiring users to visit properties in person. This resulted in a time-consuming and laborious property selection process. Furthermore, the fragmented nature of the information limited its effectiveness in supporting decision-making.
[0527] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0528] In this invention, the server includes means for receiving selection information transmitted from a user device, means for obtaining relevant detailed information from an information collection based on the received selection information, and means for compressing the information in real time and converting it into a continuous transmission format in order to transmit the obtained detailed information to the user device. This enables users to efficiently collect detailed property information and interactively review properties even from remote locations.
[0529] "User equipment" refers to terminal devices that users operate to send and receive information.
[0530] "Selected information" refers to identifying information about the subject selected by the user.
[0531] An "information collection" refers to a data storage structure in which multiple related data are aggregated.
[0532] "Real time" refers to the instantaneous processing and communication of information.
[0533] "Continuous transmission format" refers to a communication format for transmitting data in a continuous stream without interruption.
[0534] "Voice" refers to the sound signals that users use for data input.
[0535] "Textual information" refers to data in text format that has been converted from audio or other non-textual data.
[0536] "Information retrieval" refers to the process of efficiently searching for information within a database to find the desired information.
[0537] "Sound interference removal" refers to a technique for eliminating unwanted noise from an audio signal.
[0538] "Pronunciation correction" refers to a technology that compensates for the influence of different regions and accents during the speech recognition process to ensure accurate recognition.
[0539] This system allows users to gain a detailed understanding of rental properties even from a remote location, and it functions through the collaborative efforts of the server, terminal, and user.
[0540] The server manages a database containing detailed information about properties. Upon receiving selection information from a user terminal, the server immediately retrieves detailed information about the corresponding property from the database. This information includes 360-degree video, 3D models, basic property data, and surrounding environment information. The server compresses the retrieved information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user terminal via the internet.
[0541] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail. The terminal also has a function to receive user voice input via voice recording and send it directly to the server. The server uses speech recognition technology to convert the voice into text and queries the database. Based on the converted results, the server generates answers to the user's questions and sends them back to the terminal using speech synthesis technology. The terminal plays these answers through its speaker and simultaneously displays them as text on the screen.
[0542] Users can easily search for desired properties using their devices and begin viewings. Users can ask questions verbally about areas of interest or any concerns they may have. For example, if they ask, "What is the ceiling height of this room?", the system will immediately provide accurate information. This concrete example demonstrates how the system provides users with convenience and peace of mind, supporting efficient selection of rental properties.
[0543] An example of a prompt message is, "How should a user ask a question if they want to know the size of the windows in a property?" In this way, users can obtain detailed information efficiently and in real time.
[0544] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0545] Step 1:
[0546] The user selects their desired rental property via a terminal. The user's identification information for the property selected on the interface is used as input. This information is sent from the terminal to the server. The output is the state after the property identification information has been accurately transferred to the server.
[0547] Step 2:
[0548] The server searches the database based on the received property identification information and retrieves detailed information about the corresponding property. The input is property identification information, and the output is a data object containing detailed information. A database query is executed, retrieving 360-degree video, 3D models, basic data, and surrounding environment information for the property.
[0549] Step 3:
[0550] The server compresses the acquired property details in real time and converts them into a streaming format. The input here is a detailed information data object, and the output is compressed streaming data. This operation allows the data to be efficiently transmitted to the terminal.
[0551] Step 4:
[0552] The terminal renders streaming data received from the server in real time. The input is a compressed signal, and the output is a three-dimensional visual and interactive map presented to the user. The terminal decompresses the data, converts it to a format suitable for the display device, and visualizes it.
[0553] Step 5:
[0554] Users input their questions about properties via voice through a device. The input is the user's voice, and the output is the transfer of voice data to the server. The device uses a voice recording function to accurately capture the voice.
[0555] Step 6:
[0556] The server receives audio data and converts it to text using speech recognition technology. The input is audio data, and the output is the converted text data. The speech recognition algorithm performs noise reduction and accent correction to assist in query creation.
[0557] Step 7:
[0558] The server queries the database based on the converted text data to retrieve information relevant to the user's question. The input is a text query, and the output is the query result. Data filtering and searching are performed to efficiently aggregate the necessary information.
[0559] Step 8:
[0560] The server generates an answer based on the query results and sends it to the user's terminal. The input is the query result data, and the output is the answer in voice and text format. Natural language processing is performed using a generative AI model to provide contextually appropriate answers.
[0561] Step 9:
[0562] The terminal uses speech synthesis technology to play back the responses received from the server and also displays them as text on the screen. Input consists of audio and text data, while output is auditory and visual feedback to the user. This allows the user to confirm the answers and gain a deeper understanding of the property.
[0563] (Application Example 1)
[0564] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0565] In today's real estate market, a problem exists where users living remotely have difficulty understanding properties in detail. In particular, the limited access to detailed information about a property's internal structure and surrounding environment makes it difficult for users to make accurate decisions. There is a need to solve this problem and enable users to confidently select properties from the comfort of their homes.
[0566] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0567] In this invention, the server includes means for receiving object information transmitted from a user device, means for obtaining detailed data of the corresponding object from a data storage device based on the received object information, and means for immediately compressing the acquired detailed data and converting it into a continuous playback format in order to transmit it to the user device. As a result, users can visualize detailed information of properties in real time, even from a remote location, easily search for information via voice, and obtain immediate answers.
[0568] "User device" refers to a device used by a user to receive and operate information, and includes portable information terminals such as smartphones and smart glasses.
[0569] "Target information" refers to identifying information about properties or objects for which the user wishes to obtain detailed information.
[0570] "Detailed data" refers to comprehensive information including 360-degree video and 3D models of the object, basic property data, and surrounding environment information.
[0571] A "data storage device" refers to a database used to store and manage detailed data about properties or objects.
[0572] "Means of instantly compressing and converting to a continuous playback format" refers to technical methods that compress data and convert it into a streamable format so that detailed data can be transmitted efficiently.
[0573] "Voice input" refers to instructions or questions given by the user through the microphone on their device.
[0574] "Means of converting to text information" refers to processes and technologies that include converting voice input into text data.
[0575] "Speech understanding technology" refers to technology that analyzes speech data and converts speech into text information while removing noise.
[0576] A "three-dimensional virtual environment" refers to a visual environment created using computer graphics that allows users to observe properties and objects in three dimensions.
[0577] To implement this invention, it is necessary to construct a system in which various elements, such as a server, user equipment, and speech recognition technology, work together. The specific configuration and technology are described below.
[0578] The server first receives object information transmitted from the user's device. Based on this information, it retrieves the corresponding detailed data from the data storage device. This detailed data includes 360-degree video, a three-dimensional model, basic data, and surrounding environment information related to the property. The server immediately compresses this detailed data, converts it to a continuous playback format, and transmits it to the user's device. This allows the user to view the property smoothly.
[0579] The user's device renders data received from the server in real time, displaying it as a 3D visual and interactive map. This visual environment allows the user to experience a virtual tour of the property. The user's device also has a microphone to receive user voice input and transmit it to the server.
[0580] The server uses speech recognition technology to convert received audio data into text. Here, the server performs noise reduction and accent correction to analyze the data accurately. The converted text data is then searched in a data storage device to extract relevant information, and natural language is generated using automatic generation technology to respond to the user. This generated response is then converted into audio data using speech synthesis technology and transmitted to the user's device.
[0581] For example, if a user asks, "What is the ceiling height of this room?", the server can immediately search for the information corresponding to that question and provide a natural language response in both voice and text, such as, "The ceiling height of this room is 2.8 meters." This process allows users to quickly obtain accurate and detailed property information.
[0582] An example of a prompt message is: "When a user asks a question about the property they are viewing, retrieve the relevant property data from the database and generate a response. For example, if the user asks, 'What are the dimensions of this room?', tell me the dimensions based on the relevant data." This prompt allows the generating AI model to produce an appropriate response to the user's question.
[0583] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0584] Step 1:
[0585] The user's device sends object information to the server. The input is the user's specified property information, and the output is the transmission of information to the server. In this step, the user selects properties of interest through the application interface, and this selection information is sent to the server.
[0586] Step 2:
[0587] The server receives the transmitted object information and retrieves detailed data from the data storage device. The input is the object information received from the user's device, and the output is the detailed data of the property corresponding to that information. Here, the server collects 360-degree video, 3D models, basic data, and surrounding environment information related to the selected property.
[0588] Step 3:
[0589] The server immediately compresses the acquired detailed data and converts it into a continuous playback format. The input is detailed data, and the output is compressed, continuously playable data. In this step, a data compression algorithm is executed to efficiently transmit the detailed data and make it streamable, providing the user with a seamless experience.
[0590] Step 4:
[0591] The user's device renders streaming data received from the server in real time, displaying it as a 3D visual and interactive map. The input is compressed data received from the server, and the output is a user-interactive 3D visual. The user uses this view to conduct a virtual tour and check various parts of the property.
[0592] Step 5:
[0593] The user asks a question by voice, and the user's device receives this voice input. The input is the user's voice, and the output is sent to the server as voice data. The user gives voice instructions based on their interests and questions, which are recorded through the user's device and sent to the server.
[0594] Step 6:
[0595] The server uses speech understanding technology to convert received audio data into text information. The input is audio data sent from the user's device, and the output is text data. The server performs noise reduction and pronunciation correction, and uses the latest speech recognition algorithms to accurately convert speech into text.
[0596] Step 7:
[0597] The server searches the data storage device based on textual information and retrieves relevant information. The input is textual information, and the output is relevant information retrieved by the query. Here, the appropriate data for the user's question is quickly extracted from the database.
[0598] Step 8:
[0599] Based on the information acquired by the server, a generative AI model is used to generate natural language responses. The input is information from a database, and the output is a response in natural language. To achieve this, natural language processing techniques are used to convert machine-readable data into a format that users can understand.
[0600] Step 9:
[0601] The server generates speech, which is then converted into speech data using speech synthesis technology and sent to the user's device. The input is natural language response text, and the output is speech data playable on the user's device. The technology converts text to speech so that the user can hear the response.
[0602] Step 10:
[0603] The user's device plays back the received audio data and simultaneously displays it as text on the screen. The input is audio data received from the server, and the output is information provided to the user's eyes and ears. This allows the user to obtain answers to questions through both sight and hearing.
[0604] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0605] This invention provides a system that offers detailed information about a rental property selected by the user through a virtual viewing, and also recognizes the user's emotions and adjusts the interaction accordingly.
[0606] Server-based data management and processing
[0607] The server manages a database containing property information. Upon receiving property selection information from a user's terminal, the server immediately references the database based on that information and retrieves detailed information about the corresponding property (360-degree video, 3D model, basic information, etc.).
[0608] The acquired information is compressed in real time and converted into a streaming format. This data is then ready to be sent to the terminal. Furthermore, when voice input from the user is received via the terminal, the server combines speech recognition technology and an emotion engine to convert this data into text and analyze the emotional state.
[0609] Functions of the Emotion Engine
[0610] The emotion engine analyzes the user's voice data to capture emotional nuances. This utilizes information such as voice tone, pitch, and tempo to determine whether the user is in a positive, negative, or neutral emotional state.
[0611] Based on this emotional information, the server determines how to adjust the interface and present information according to the user's emotions. For example, if the user is in a positive state, it will provide more detailed information, while if they are negative, it will focus on encouraging messages and simpler information.
[0612] User interface via terminal
[0613] The terminal decompresses the compressed data received from the server and renders it in real time as a 3D visual or interactive map. This allows the user to visually experience the property.
[0614] If necessary, the device adjusts its display and audio output to improve the user's visual experience. Based on the results of the emotion engine, the device can offer the user feedback or recommended actions.
[0615] Examples
[0616] As a concrete example, consider a case where a user asks about the size of a property. If the user's voice indicates negative emotions, the system will focus on that and emphasize encouraging or positive information. For example, it might add a statement like, "This space allows for flexible furniture arrangement."
[0617] Thus, this embodiment realizes a system that allows users to remotely check the details of rental properties and dynamically adjusts the information provided and the interface to create an experience that matches their emotions.
[0618] The following describes the processing flow.
[0619] Step 1:
[0620] The user selects the properties they want to view through their device. This selection information is then sent from the device to the server.
[0621] Step 2:
[0622] Based on the property selection information received, the server consults the database and retrieves detailed information about the corresponding property.
[0623] Step 3:
[0624] The server compresses the acquired property information in real time, converts it into a streamable format, and then prepares it for transmission to the terminal.
[0625] Step 4:
[0626] The terminal decompresses the compressed data received from the server and displays the acquired information as a 3D visual and an interactive map. At the same time, it provides a navigation function to allow the user to freely browse properties.
[0627] Step 5:
[0628] As users view properties, they can ask questions about anything that concerns them using voice. This voice input is captured on the device and transmitted to the server.
[0629] Step 6:
[0630] The server receives voice input and converts the speech into text using speech recognition technology, while simultaneously analyzing the user's emotional state with an emotion engine.
[0631] Step 7:
[0632] Based on the transcribed questions, the server executes database queries to retrieve relevant information and determines the appropriate response style based on the analyzed sentiment data.
[0633] Step 8:
[0634] The server performs natural language processing in a tone and style that matches the user's emotions, generates a response optimized for the user, and sends it to the terminal.
[0635] Step 9:
[0636] The device plays back the response from the server using speech synthesis technology and simultaneously displays it on the screen as text. This allows the user to receive rich information and advice based on their responses.
[0637] (Example 2)
[0638] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0639] Traditional real estate information systems provide information without considering the user's emotional state, resulting in a uniform user experience that fails to adequately address individual needs. Furthermore, the lack of sufficient visual presentation of property information and limited opportunities for active interaction are also problematic.
[0640] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0641] In this invention, the server includes means for receiving property selection information transmitted from a user terminal, means for obtaining detailed information of the relevant property from a data storage medium based on the received property selection information, means for compressing the data in real time and converting it into a data streaming format in order to transmit the obtained detailed information to the user terminal, means for receiving voice input from the user and converting the voice data into text, means for performing sentiment analysis from the converted text, and means for dynamically adjusting the interface and changing the method of presenting information according to the user's emotional state. This makes it possible to provide information and adjust the interface to match the user's emotional state, thereby realizing a more personalized user experience.
[0642] A "user terminal" is an electronic device used by a user to input or output information.
[0643] "Property selection information" refers to data that represents the user's selections regarding real estate properties they are interested in.
[0644] A "data storage medium" is a physical or virtual device used to store digital information.
[0645] "Detailed information" refers to information about various attributes and characteristics of a real estate property.
[0646] A "data streaming format" is a data format for transmitting digital data continuously in real time.
[0647] "Voice input" refers to data of words and sounds spoken by the user.
[0648] "Text conversion" is the process of representing audio data as text.
[0649] "Emotional analysis" is an analytical technique used to identify emotional states from voice and other data.
[0650] An "interface" is a means or method for a user to exchange information with a system.
[0651] This invention is a system that allows users to remotely view details of real estate properties and provides interactions that take into account the user's emotional state. The system mainly consists of a server and a user terminal, each playing a specific role.
[0652] The server manages the data storage system and stores detailed information about properties. Based on the property selection information received from the user's terminal, the server quickly retrieves the necessary property details from the data storage, compresses the data in real time, and converts it into a data streaming format. Furthermore, the server uses speech recognition technology and an emotion analysis engine to convert the user's voice input into text and analyze their emotional state. Emotion analysis utilizes voice features such as tone, tempo, and pitch to determine whether the user is feeling positive, negative, or neutral. Based on these analysis results, the server dynamically adjusts the interface and presents information according to the user's emotional state.
[0653] The device decompresses compressed data sent from the server and renders it as a spatial digital representation in real time. This allows users to visually experience the property through 3D visuals and interactive maps. The device also adjusts the display and audio output to optimize the user's visual experience. Based on the results of sentiment analysis, the device provides feedback and recommended actions to the user, helping them to fully understand the value of the property.
[0654] As a concrete example, consider a scenario where a user asks a voice question about the size of a property, such as, "How big is the living room in this property?" If the device detects that the user's voice is expressing negative emotions, it will provide positive feedback such as, "This living room has large windows and gets plenty of light." This allows the user to focus on the positive features of the property.
[0655] An example of a prompt message might be: "What are the key points the user wants to know about this rental property? How should that information be presented, depending on the user's emotional state?"
[0656] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0657] Step 1:
[0658] Users select properties of interest via their device. Information about the selected properties is sent from the device to the server as "property selection information." This input information includes the property ID and the reason the user is interested in the property.
[0659] Step 2:
[0660] The server references the data storage system based on the received property selection information. The data processing performed here involves retrieving detailed information about the relevant property using queries based on the property ID. The output includes detailed information such as 360-degree video, 3D models, and basic property information.
[0661] Step 3:
[0662] The server compresses the acquired detailed information in real time and converts it into a streaming format. This data processing improves transmission efficiency. The output is a compressed data stream, ready to be sent to the user's terminal.
[0663] Step 4:
[0664] The user enters specific questions about the property via voice. This voice data is sent from the terminal to the server. The entered voice file is received on the server side.
[0665] Step 5:
[0666] The server analyzes the received audio data. Using speech recognition technology, it converts the audio to text, and then uses an emotion analysis engine to analyze the emotional state. Through data calculations, it determines whether the user is in a positive, negative, or neutral emotional state. The output consists of the converted text and the emotional state evaluation result.
[0667] Step 6:
[0668] The server adjusts the user interface based on the sentiment analysis results. If the sentiment is positive, it provides detailed information; if it's negative, it generates supplementary information and encouraging messages. This optimizes the way information is presented and its content. The output is the adjusted information presentation.
[0669] Step 7:
[0670] The terminal decompresses the compressed data sent from the server and renders it in real time as a spatial digital representation. Data processing is performed for display as 3D visuals or interactive maps. The output is a display that the user can visually observe.
[0671] Step 8:
[0672] The device presents the user with feedback and recommended actions based on the results of sentiment analysis. For example, positive feedback such as "This living room has large windows and gets plenty of light" might be displayed. The output consists of content and recommended actions presented to the user.
[0673] (Application Example 2)
[0674] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0675] Conventional virtual viewing systems simply provide property information, lacking dynamic interaction that responds to user emotions. Furthermore, because information is not provided in a way that considers user feelings, the user experience is uniform, making it difficult to enhance individual satisfaction.
[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0677] In this invention, the server includes means for receiving real estate information transmitted from a user device, means for obtaining detailed information of the relevant real estate from a recording device based on the received real estate information, means for compressing the data in real time and converting it into a streaming format in order to transmit the obtained detailed information to the user device, means for analyzing the user's voice data using emotion recognition technology to determine the emotional state, and means for adjusting the displayed information based on the user's emotions. This enables the provision of dynamic information tailored to the user's emotions, resulting in a more satisfying virtual viewing experience.
[0678] A "user device" is a terminal device that a user operates to receive and transmit information.
[0679] "Real estate information" refers to detailed data about rental properties, including information such as location, floor plan, and amenities.
[0680] A "recording device" is a database system used to store and manage detailed property information.
[0681] "Streaming format" refers to a data transmission method for playing back data while transferring it in real time.
[0682] "Emotion recognition technology" is a technology that analyzes a user's emotions from voice data and image data to determine their emotional state, such as positive, negative, or neutral.
[0683] "Dynamic information provision" refers to a system that changes the data displayed and the answers provided in real time according to the user's current state and needs.
[0684] In order to implement this invention, the following system configuration and processing procedure are required.
[0685] First, the server receives property information from the user's device. Based on the received information, it retrieves detailed information about the relevant property from the recording device. This detailed information includes a 360-degree view and a 3D model of the property. The server compresses this information in real time and converts it into a streaming format. After that, it prepares it for transmission to the user's device.
[0686] The user device, acting as the terminal, decompresses received compressed data and renders it in real time to provide a three-dimensional visual output. It also manages user interaction through voice and eye-tracking input and uses emotion recognition technology to determine the user's emotional state. This technology includes software that analyzes voice tone and speed. Based on emotions, it dynamically adjusts the content and presentation of displayed information.
[0687] This system primarily operates using wearable devices such as smart glasses, and determines the user's emotional state based on their voice. For example, if a user expresses dissatisfaction by saying, "This room is small," it can then provide additional positive information such as, "Although there is limited space, you can use it more effectively depending on how you arrange the shelves."
[0688] As a concrete example, consider a scenario where a user wears smart glasses and requests to "show me the bathroom." This system displays a virtual view of the bathroom on the glasses and performs emotion recognition. Even if the user is disappointed, it provides encouraging information such as, "There is a relaxing bathtub installed."
[0689] A concrete example of a prompt message would be: "Design an application that virtually provides viewing information about a property selected by the user, analyzes the emotion in their voice, and provides feedback accordingly." In this way, the user can analyze and evaluate properties in real time in great detail.
[0690] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0691] Step 1:
[0692] The server retrieves real estate information received from the user's device. This real estate information includes identification information for the property selected by the user. Based on this input data, the server queries the recording device to retrieve detailed information about the relevant property. The retrieved detailed information includes 360-degree images and a three-dimensional model of the property.
[0693] Step 2:
[0694] The server compresses the acquired property details and converts them into a real-time streaming format. This process encodes the data to efficiently utilize communication bandwidth and prepares it for transmission to the user's device. The streaming data is then sent to the transmission queue as output.
[0695] Step 3:
[0696] The user device receives compressed streaming data sent from the server. This input data is decompressed and rendered in real time. Specifically, it provides the user with a visual output in the form of a three-dimensional image and an interactive map. Through the device, the user can virtually tour the property.
[0697] Step 4:
[0698] Users ask questions and make comments about properties via voice input. Once the voice data is input to the user's device, the terminal uses a speech recognition engine to convert it into text, and then uses emotion recognition technology to analyze the user's emotions. This analysis is performed by analyzing acoustic features, including voice tone and speed. The analysis results in the output of emotional state data.
[0699] Step 5:
[0700] The server dynamically adjusts the information displayed based on the user's emotional state. For positive emotions, it provides detailed information; for negative emotions, it generates feedback such as supplementary information or encouraging messages. This adjusted information is then sent to the user's device.
[0701] Step 6:
[0702] The user device receives pre-configured information from the server and presents it to the user according to a predetermined display method. Audio and visual output facilitates a deeper understanding of the property and a more positive viewing experience. Further interactions proceed based on the user's subsequent actions.
[0703] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0704] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0705] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0706] [Fourth Embodiment]
[0707] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0708] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0709] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0710] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0711] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0712] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0713] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0714] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0715] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0716] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0717] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0718] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0719] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0720] This invention provides a system that allows users to understand rental properties in detail and resolve any questions they may have, even from a remote location. This system functions through the collaborative efforts of a server, a terminal, and the user.
[0721] Server-based data management and processing
[0722] The server manages a database containing detailed information about properties. When it receives property selection information from a user's terminal, the server immediately retrieves the corresponding property information from the database. The retrieved information includes 360-degree video, 3D models, basic property data, and surrounding environment information.
[0723] The server compresses this information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user's terminal over the internet. When the server receives voice input from the user, it uses speech recognition technology to convert the voice into text data and queries the database. Based on the results, the server generates and returns an answer to the user's question.
[0724] User interface via terminal
[0725] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail.
[0726] Furthermore, the device receives the user's voice input through its voice recording function and sends it directly to the server. The response sent back from the server is played back through the speaker using speech synthesis technology and simultaneously displayed as text on the device screen.
[0727] User actions
[0728] Users can easily search for their desired properties and begin virtual tours using their devices. Specifically, they can view property information and videos displayed on the interface and ask the system questions verbally about areas of interest or anything they are unsure of.
[0729] For example, if a user asks, "What is the ceiling height of this room?", the system will immediately provide accurate information to answer the user's question. This allows users to check property details remotely, enabling efficient and secure property selection.
[0730] Thus, this embodiment provides users with abundant information about properties and two-way communication, making the property viewing experience more convenient and effective.
[0731] The following describes the processing flow.
[0732] Step 1:
[0733] The user selects properties of interest using the terminal's interface. The terminal then prepares to send this selection information to the server.
[0734] Step 2:
[0735] The terminal sends the user's property selection information to the server. Upon receiving this information, the server identifies the corresponding property ID.
[0736] Step 3:
[0737] The server accesses the database based on the property ID and retrieves detailed information about the property, such as 360-degree video, 3D models, and basic information.
[0738] Step 4:
[0739] The server compresses the acquired detailed information in real time and converts it into a streamable format. This data is then prepared for transmission to the user's terminal.
[0740] Step 5:
[0741] The terminal receives compressed data from the server and renders it in real time as a 3D visual or interactive map.
[0742] Step 6:
[0743] While viewing videos of the property, users can ask questions using voice commands about anything that interests them. For example, they might ask, "Which way do the windows face?"
[0744] Step 7:
[0745] The terminal receives voice input from the user and sends the voice data to the server. The server receives the voice data.
[0746] Step 8:
[0747] The server uses speech recognition technology to convert the audio data into text, analyzes the question content, and retrieves relevant data using database queries.
[0748] Step 9:
[0749] The server generates a response based on the analysis results and prepares to send it back as audio and text data.
[0750] Step 10:
[0751] The device receives the response from the server, plays it back through the speaker using speech synthesis technology, and also displays the response text on the screen. The user then uses the received information to deepen their understanding of the property.
[0752] (Example 1)
[0753] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0754] Traditional methods for remotely gathering information on rental properties and resolving questions were inefficient, requiring users to visit properties in person. This resulted in a time-consuming and laborious property selection process. Furthermore, the fragmented nature of the information limited its effectiveness in supporting decision-making.
[0755] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0756] In this invention, the server includes means for receiving selection information transmitted from a user device, means for obtaining relevant detailed information from an information collection based on the received selection information, and means for compressing the information in real time and converting it into a continuous transmission format in order to transmit the obtained detailed information to the user device. This enables users to efficiently collect detailed property information and interactively review properties even from remote locations.
[0757] "User equipment" refers to terminal devices that users operate to send and receive information.
[0758] "Selected information" refers to identifying information about the subject selected by the user.
[0759] An "information collection" refers to a data storage structure in which multiple related data are aggregated.
[0760] "Real time" refers to the instantaneous processing and communication of information.
[0761] "Continuous transmission format" refers to a communication format for transmitting data in a continuous stream without interruption.
[0762] "Voice" refers to the sound signals that users use for data input.
[0763] "Textual information" refers to data in text format that has been converted from audio or other non-textual data.
[0764] "Information retrieval" refers to the process of efficiently searching for information within a database to find the desired information.
[0765] "Sound interference removal" refers to a technique for eliminating unwanted noise from an audio signal.
[0766] "Pronunciation correction" refers to a technology that compensates for the influence of different regions and accents during the speech recognition process to ensure accurate recognition.
[0767] This system allows users to gain a detailed understanding of rental properties even from a remote location, and it functions through the collaborative efforts of the server, terminal, and user.
[0768] The server manages a database containing detailed information about properties. Upon receiving selection information from a user terminal, the server immediately retrieves detailed information about the corresponding property from the database. This information includes 360-degree video, 3D models, basic property data, and surrounding environment information. The server compresses the retrieved information in real time and converts it into a streaming format. This format conversion allows the data to be quickly transmitted to the user terminal via the internet.
[0769] The terminal renders video data received from the server in real time, displaying a 3D visual and interactive map to the user. Through this interface, the user can view the property in detail. The terminal also has a function to receive user voice input via voice recording and send it directly to the server. The server uses speech recognition technology to convert the voice into text and queries the database. Based on the converted results, the server generates answers to the user's questions and sends them back to the terminal using speech synthesis technology. The terminal plays these answers through its speaker and simultaneously displays them as text on the screen.
[0770] Users can easily search for desired properties using their devices and begin viewings. Users can ask questions verbally about areas of interest or any concerns they may have. For example, if they ask, "What is the ceiling height of this room?", the system will immediately provide accurate information. This concrete example demonstrates how the system provides users with convenience and peace of mind, supporting efficient selection of rental properties.
[0771] An example of a prompt message is, "How should a user ask a question if they want to know the size of the windows in a property?" In this way, users can obtain detailed information efficiently and in real time.
[0772] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0773] Step 1:
[0774] The user selects their desired rental property via a terminal. The user's identification information for the property selected on the interface is used as input. This information is sent from the terminal to the server. The output is the state after the property identification information has been accurately transferred to the server.
[0775] Step 2:
[0776] The server searches the database based on the received property identification information and retrieves detailed information about the corresponding property. The input is property identification information, and the output is a data object containing detailed information. A database query is executed, retrieving 360-degree video, 3D models, basic data, and surrounding environment information for the property.
[0777] Step 3:
[0778] The server compresses the acquired property details in real time and converts them into a streaming format. The input here is a detailed information data object, and the output is compressed streaming data. This operation allows the data to be efficiently transmitted to the terminal.
[0779] Step 4:
[0780] The terminal renders streaming data received from the server in real time. The input is a compressed signal, and the output is a three-dimensional visual and interactive map presented to the user. The terminal decompresses the data, converts it to a format suitable for the display device, and visualizes it.
[0781] Step 5:
[0782] Users input their questions about properties via voice through a device. The input is the user's voice, and the output is the transfer of voice data to the server. The device uses a voice recording function to accurately capture the voice.
[0783] Step 6:
[0784] The server receives audio data and converts it to text using speech recognition technology. The input is audio data, and the output is the converted text data. The speech recognition algorithm performs noise reduction and accent correction to assist in query creation.
[0785] Step 7:
[0786] The server queries the database based on the converted text data to retrieve information relevant to the user's question. The input is a text query, and the output is the query result. Data filtering and searching are performed to efficiently aggregate the necessary information.
[0787] Step 8:
[0788] The server generates an answer based on the query results and sends it to the user's terminal. The input is the query result data, and the output is the answer in voice and text format. Natural language processing is performed using a generative AI model to provide contextually appropriate answers.
[0789] Step 9:
[0790] The terminal uses speech synthesis technology to play back the responses received from the server and also displays them as text on the screen. Input consists of audio and text data, while output is auditory and visual feedback to the user. This allows the user to confirm the answers and gain a deeper understanding of the property.
[0791] (Application Example 1)
[0792] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0793] In today's real estate market, a problem exists where users living remotely have difficulty understanding properties in detail. In particular, the limited access to detailed information about a property's internal structure and surrounding environment makes it difficult for users to make accurate decisions. There is a need to solve this problem and enable users to confidently select properties from the comfort of their homes.
[0794] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0795] In this invention, the server includes means for receiving object information transmitted from a user device, means for obtaining detailed data of the corresponding object from a data storage device based on the received object information, and means for immediately compressing the acquired detailed data and converting it into a continuous playback format in order to transmit it to the user device. As a result, users can visualize detailed information of properties in real time, even from a remote location, easily search for information via voice, and obtain immediate answers.
[0796] "User device" refers to a device used by a user to receive and operate information, and includes portable information terminals such as smartphones and smart glasses.
[0797] "Target information" refers to identifying information about properties or objects for which the user wishes to obtain detailed information.
[0798] "Detailed data" refers to comprehensive information including 360-degree video and 3D models of the object, basic property data, and surrounding environment information.
[0799] A "data storage device" refers to a database used to store and manage detailed data about properties or objects.
[0800] "Means of instantly compressing and converting to a continuous playback format" refers to technical methods that compress data and convert it into a streamable format so that detailed data can be transmitted efficiently.
[0801] "Voice input" refers to instructions or questions given by the user through the microphone on their device.
[0802] "Means of converting to text information" refers to processes and technologies that include converting voice input into text data.
[0803] "Speech understanding technology" refers to technology that analyzes speech data and converts speech into text information while removing noise.
[0804] A "three-dimensional virtual environment" refers to a visual environment created using computer graphics that allows users to observe properties and objects in three dimensions.
[0805] To implement this invention, it is necessary to construct a system in which various elements, such as a server, user equipment, and speech recognition technology, work together. The specific configuration and technology are described below.
[0806] The server first receives object information transmitted from the user's device. Based on this information, it retrieves the corresponding detailed data from the data storage device. This detailed data includes 360-degree video, a three-dimensional model, basic data, and surrounding environment information related to the property. The server immediately compresses this detailed data, converts it to a continuous playback format, and transmits it to the user's device. This allows the user to view the property smoothly.
[0807] The user's device renders data received from the server in real time, displaying it as a 3D visual and interactive map. This visual environment allows the user to experience a virtual tour of the property. The user's device also has a microphone to receive user voice input and transmit it to the server.
[0808] The server uses speech recognition technology to convert received audio data into text. Here, the server performs noise reduction and accent correction to analyze the data accurately. The converted text data is then searched in a data storage device to extract relevant information, and natural language is generated using automatic generation technology to respond to the user. This generated response is then converted into audio data using speech synthesis technology and transmitted to the user's device.
[0809] For example, if a user asks, "What is the ceiling height of this room?", the server can immediately search for the information corresponding to that question and provide a natural language response in both voice and text, such as, "The ceiling height of this room is 2.8 meters." This process allows users to quickly obtain accurate and detailed property information.
[0810] An example of a prompt message is: "When a user asks a question about the property they are viewing, retrieve the relevant property data from the database and generate a response. For example, if the user asks, 'What are the dimensions of this room?', tell me the dimensions based on the relevant data." This prompt allows the generating AI model to produce an appropriate response to the user's question.
[0811] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0812] Step 1:
[0813] The user's device sends object information to the server. The input is the user's specified property information, and the output is the transmission of information to the server. In this step, the user selects properties of interest through the application interface, and this selection information is sent to the server.
[0814] Step 2:
[0815] The server receives the transmitted object information and retrieves detailed data from the data storage device. The input is the object information received from the user's device, and the output is the detailed data of the property corresponding to that information. Here, the server collects 360-degree video, 3D models, basic data, and surrounding environment information related to the selected property.
[0816] Step 3:
[0817] The server immediately compresses the acquired detailed data and converts it into a continuous playback format. The input is detailed data, and the output is compressed, continuously playable data. In this step, a data compression algorithm is executed to efficiently transmit the detailed data and make it streamable, providing the user with a seamless experience.
[0818] Step 4:
[0819] The user's device renders streaming data received from the server in real time, displaying it as a 3D visual and interactive map. The input is compressed data received from the server, and the output is a user-interactive 3D visual. The user uses this view to conduct a virtual tour and check various parts of the property.
[0820] Step 5:
[0821] The user asks a question by voice, and the user's device receives this voice input. The input is the user's voice, and the output is sent to the server as voice data. The user gives voice instructions based on their interests and questions, which are recorded through the user's device and sent to the server.
[0822] Step 6:
[0823] The server uses speech understanding technology to convert received audio data into text information. The input is audio data sent from the user's device, and the output is text data. The server performs noise reduction and pronunciation correction, and uses the latest speech recognition algorithms to accurately convert speech into text.
[0824] Step 7:
[0825] The server searches the data storage device based on textual information and retrieves relevant information. The input is textual information, and the output is relevant information retrieved by the query. Here, the appropriate data for the user's question is quickly extracted from the database.
[0826] Step 8:
[0827] Based on the information acquired by the server, a generative AI model is used to generate natural language responses. The input is information from a database, and the output is a response in natural language. To achieve this, natural language processing techniques are used to convert machine-readable data into a format that users can understand.
[0828] Step 9:
[0829] The server generates speech, which is then converted into speech data using speech synthesis technology and sent to the user's device. The input is natural language response text, and the output is speech data playable on the user's device. The technology converts text to speech so that the user can hear the response.
[0830] Step 10:
[0831] The user's device plays back the received audio data and simultaneously displays it as text on the screen. The input is audio data received from the server, and the output is information provided to the user's eyes and ears. This allows the user to obtain answers to questions through both sight and hearing.
[0832] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0833] This invention provides a system that offers detailed information about a rental property selected by the user through a virtual viewing, and also recognizes the user's emotions and adjusts the interaction accordingly.
[0834] Server-based data management and processing
[0835] The server manages a database containing property information. Upon receiving property selection information from a user's terminal, the server immediately references the database based on that information and retrieves detailed information about the corresponding property (360-degree video, 3D model, basic information, etc.).
[0836] The acquired information is compressed in real time and converted into a streaming format. This data is then ready to be sent to the terminal. Furthermore, when voice input from the user is received via the terminal, the server combines speech recognition technology and an emotion engine to convert this data into text and analyze the emotional state.
[0837] Functions of the Emotion Engine
[0838] The emotion engine analyzes the user's voice data to capture emotional nuances. This utilizes information such as voice tone, pitch, and tempo to determine whether the user is in a positive, negative, or neutral emotional state.
[0839] Based on this emotional information, the server determines how to adjust the interface and present information according to the user's emotions. For example, if the user is in a positive state, it will provide more detailed information, while if they are negative, it will focus on encouraging messages and simpler information.
[0840] User interface via terminal
[0841] The terminal decompresses the compressed data received from the server and renders it in real time as a 3D visual or interactive map. This allows the user to visually experience the property.
[0842] If necessary, the device adjusts its display and audio output to improve the user's visual experience. Based on the results of the emotion engine, the device can offer the user feedback or recommended actions.
[0843] Examples
[0844] As a concrete example, consider a case where a user asks about the size of a property. If the user's voice indicates negative emotions, the system will focus on that and emphasize encouraging or positive information. For example, it might add a statement like, "This space allows for flexible furniture arrangement."
[0845] Thus, this embodiment realizes a system that allows users to remotely check the details of rental properties and dynamically adjusts the information provided and the interface to create an experience that matches their emotions.
[0846] The following describes the processing flow.
[0847] Step 1:
[0848] The user selects the properties they want to view through their device. This selection information is then sent from the device to the server.
[0849] Step 2:
[0850] Based on the property selection information received, the server consults the database and retrieves detailed information about the corresponding property.
[0851] Step 3:
[0852] The server compresses the acquired property information in real time, converts it into a streamable format, and then prepares it for transmission to the terminal.
[0853] Step 4:
[0854] The terminal decompresses the compressed data received from the server and displays the acquired information as a 3D visual and an interactive map. At the same time, it provides a navigation function to allow the user to freely browse properties.
[0855] Step 5:
[0856] As users view properties, they can ask questions about anything that concerns them using voice. This voice input is captured on the device and transmitted to the server.
[0857] Step 6:
[0858] The server receives voice input and converts the speech into text using speech recognition technology, while simultaneously analyzing the user's emotional state with an emotion engine.
[0859] Step 7:
[0860] Based on the transcribed questions, the server executes database queries to retrieve relevant information and determines the appropriate response style based on the analyzed sentiment data.
[0861] Step 8:
[0862] The server performs natural language processing in a tone and style that matches the user's emotions, generates a response optimized for the user, and sends it to the terminal.
[0863] Step 9:
[0864] The device plays back the response from the server using speech synthesis technology and simultaneously displays it on the screen as text. This allows the user to receive rich information and advice based on their responses.
[0865] (Example 2)
[0866] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0867] Traditional real estate information systems provide information without considering the user's emotional state, resulting in a uniform user experience that fails to adequately address individual needs. Furthermore, the lack of sufficient visual presentation of property information and limited opportunities for active interaction are also problematic.
[0868] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0869] In this invention, the server includes means for receiving property selection information transmitted from a user terminal, means for obtaining detailed information of the relevant property from a data storage medium based on the received property selection information, means for compressing the data in real time and converting it into a data streaming format in order to transmit the obtained detailed information to the user terminal, means for receiving voice input from the user and converting the voice data into text, means for performing sentiment analysis from the converted text, and means for dynamically adjusting the interface and changing the method of presenting information according to the user's emotional state. This makes it possible to provide information and adjust the interface to match the user's emotional state, thereby realizing a more personalized user experience.
[0870] A "user terminal" is an electronic device used by a user to input or output information.
[0871] "Property selection information" refers to data that represents the user's selections regarding real estate properties they are interested in.
[0872] A "data storage medium" is a physical or virtual device used to store digital information.
[0873] "Detailed information" refers to information about various attributes and characteristics of a real estate property.
[0874] A "data streaming format" is a data format for transmitting digital data continuously in real time.
[0875] "Voice input" refers to data of words and sounds spoken by the user.
[0876] "Text conversion" is the process of representing audio data as text.
[0877] "Emotional analysis" is an analytical technique used to identify emotional states from voice and other data.
[0878] An "interface" is a means or method for a user to exchange information with a system.
[0879] This invention is a system that allows users to remotely view details of real estate properties and provides interactions that take into account the user's emotional state. The system mainly consists of a server and a user terminal, each playing a specific role.
[0880] The server manages the data storage system and stores detailed information about properties. Based on the property selection information received from the user's terminal, the server quickly retrieves the necessary property details from the data storage, compresses the data in real time, and converts it into a data streaming format. Furthermore, the server uses speech recognition technology and an emotion analysis engine to convert the user's voice input into text and analyze their emotional state. Emotion analysis utilizes voice features such as tone, tempo, and pitch to determine whether the user is feeling positive, negative, or neutral. Based on these analysis results, the server dynamically adjusts the interface and presents information according to the user's emotional state.
[0881] The device decompresses compressed data sent from the server and renders it as a spatial digital representation in real time. This allows users to visually experience the property through 3D visuals and interactive maps. The device also adjusts the display and audio output to optimize the user's visual experience. Based on the results of sentiment analysis, the device provides feedback and recommended actions to the user, helping them to fully understand the value of the property.
[0882] As a concrete example, consider a scenario where a user asks a voice question about the size of a property, such as, "How big is the living room in this property?" If the device detects that the user's voice is expressing negative emotions, it will provide positive feedback such as, "This living room has large windows and gets plenty of light." This allows the user to focus on the positive features of the property.
[0883] An example of a prompt message might be: "What are the key points the user wants to know about this rental property? How should that information be presented, depending on the user's emotional state?"
[0884] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0885] Step 1:
[0886] Users select properties of interest via their device. Information about the selected properties is sent from the device to the server as "property selection information." This input information includes the property ID and the reason the user is interested in the property.
[0887] Step 2:
[0888] The server references the data storage system based on the received property selection information. The data processing performed here involves retrieving detailed information about the relevant property using queries based on the property ID. The output includes detailed information such as 360-degree video, 3D models, and basic property information.
[0889] Step 3:
[0890] The server compresses the acquired detailed information in real time and converts it into a streaming format. This data processing improves transmission efficiency. The output is a compressed data stream, ready to be sent to the user's terminal.
[0891] Step 4:
[0892] The user enters specific questions about the property via voice. This voice data is sent from the terminal to the server. The entered voice file is received on the server side.
[0893] Step 5:
[0894] The server analyzes the received audio data. Using speech recognition technology, it converts the audio to text, and then uses an emotion analysis engine to analyze the emotional state. Through data calculations, it determines whether the user is in a positive, negative, or neutral emotional state. The output consists of the converted text and the emotional state evaluation result.
[0895] Step 6:
[0896] The server adjusts the user interface based on the sentiment analysis results. If the sentiment is positive, it provides detailed information; if it's negative, it generates supplementary information and encouraging messages. This optimizes the way information is presented and its content. The output is the adjusted information presentation.
[0897] Step 7:
[0898] The terminal decompresses the compressed data sent from the server and renders it in real time as a spatial digital representation. Data processing is performed for display as 3D visuals or interactive maps. The output is a display that the user can visually observe.
[0899] Step 8:
[0900] The device presents the user with feedback and recommended actions based on the results of sentiment analysis. For example, positive feedback such as "This living room has large windows and gets plenty of light" might be displayed. The output consists of content and recommended actions presented to the user.
[0901] (Application Example 2)
[0902] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0903] Conventional virtual viewing systems simply provide property information, lacking dynamic interaction that responds to user emotions. Furthermore, because information is not provided in a way that considers user feelings, the user experience is uniform, making it difficult to enhance individual satisfaction.
[0904] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0905] In this invention, the server includes means for receiving real estate information transmitted from a user device, means for obtaining detailed information of the relevant real estate from a recording device based on the received real estate information, means for compressing the data in real time and converting it into a streaming format in order to transmit the obtained detailed information to the user device, means for analyzing the user's voice data using emotion recognition technology to determine the emotional state, and means for adjusting the displayed information based on the user's emotions. This enables the provision of dynamic information tailored to the user's emotions, resulting in a more satisfying virtual viewing experience.
[0906] A "user device" is a terminal device that a user operates to receive and transmit information.
[0907] "Real estate information" refers to detailed data about rental properties, including information such as location, floor plan, and amenities.
[0908] A "recording device" is a database system used to store and manage detailed property information.
[0909] "Streaming format" refers to a data transmission method for playing back data while transferring it in real time.
[0910] "Emotion recognition technology" is a technology that analyzes a user's emotions from voice data and image data to determine their emotional state, such as positive, negative, or neutral.
[0911] "Dynamic information provision" refers to a system that changes the data displayed and the answers provided in real time according to the user's current state and needs.
[0912] In order to implement this invention, the following system configuration and processing procedure are required.
[0913] First, the server receives property information from the user's device. Based on the received information, it retrieves detailed information about the relevant property from the recording device. This detailed information includes a 360-degree view and a 3D model of the property. The server compresses this information in real time and converts it into a streaming format. After that, it prepares it for transmission to the user's device.
[0914] The user device, acting as the terminal, decompresses received compressed data and renders it in real time to provide a three-dimensional visual output. It also manages user interaction through voice and eye-tracking input and uses emotion recognition technology to determine the user's emotional state. This technology includes software that analyzes voice tone and speed. Based on emotions, it dynamically adjusts the content and presentation of displayed information.
[0915] This system primarily operates using wearable devices such as smart glasses, and determines the user's emotional state based on their voice. For example, if a user expresses dissatisfaction by saying, "This room is small," it can then provide additional positive information such as, "Although there is limited space, you can use it more effectively depending on how you arrange the shelves."
[0916] As a concrete example, consider a scenario where a user wears smart glasses and requests to "show me the bathroom." This system displays a virtual view of the bathroom on the glasses and performs emotion recognition. Even if the user is disappointed, it provides encouraging information such as, "There is a relaxing bathtub installed."
[0917] A concrete example of a prompt message would be: "Design an application that virtually provides viewing information about a property selected by the user, analyzes the emotion in their voice, and provides feedback accordingly." In this way, the user can analyze and evaluate properties in real time in great detail.
[0918] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0919] Step 1:
[0920] The server retrieves real estate information received from the user's device. This real estate information includes identification information for the property selected by the user. Based on this input data, the server queries the recording device to retrieve detailed information about the relevant property. The retrieved detailed information includes 360-degree images and a three-dimensional model of the property.
[0921] Step 2:
[0922] The server compresses the acquired property details and converts them into a real-time streaming format. This process encodes the data to efficiently utilize communication bandwidth and prepares it for transmission to the user's device. The streaming data is then sent to the transmission queue as output.
[0923] Step 3:
[0924] The user device receives compressed streaming data sent from the server. This input data is decompressed and rendered in real time. Specifically, it provides the user with a visual output in the form of a three-dimensional image and an interactive map. Through the device, the user can virtually tour the property.
[0925] Step 4:
[0926] Users ask questions and make comments about properties via voice input. Once the voice data is input to the user's device, the terminal uses a speech recognition engine to convert it into text, and then uses emotion recognition technology to analyze the user's emotions. This analysis is performed by analyzing acoustic features, including voice tone and speed. The analysis results in the output of emotional state data.
[0927] Step 5:
[0928] The server dynamically adjusts the information displayed based on the user's emotional state. For positive emotions, it provides detailed information; for negative emotions, it generates feedback such as supplementary information or encouraging messages. This adjusted information is then sent to the user's device.
[0929] Step 6:
[0930] The user device receives pre-configured information from the server and presents it to the user according to a predetermined display method. Audio and visual output facilitates a deeper understanding of the property and a more positive viewing experience. Further interactions proceed based on the user's subsequent actions.
[0931] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0932] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0933] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0934] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0935] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0936] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0937] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0938] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0939] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0940] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0941] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0942] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0943] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0944] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0945] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0946] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0947] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0948] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0949] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0950] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0951] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0952] The following is further disclosed regarding the embodiments described above.
[0953] (Claim 1)
[0954] A means for receiving property selection information transmitted from a user terminal,
[0955] A means of obtaining detailed information about the relevant property from a database based on the received property selection information,
[0956] A means for compressing the acquired detailed information in real time and converting it to a streaming format in order to send it to the user's terminal,
[0957] A means for receiving voice input from a user and converting that voice data into text,
[0958] A means of obtaining relevant information from the converted text by executing a database query,
[0959] A means of generating a response using natural language processing based on acquired information and sending it to the user's terminal,
[0960] A system that includes this.
[0961] (Claim 2)
[0962] The system according to claim 1, wherein the user terminal has means for rendering the received compressed detailed information in real time and displaying it as a three-dimensional visual and interactive map.
[0963] (Claim 3)
[0964] The system according to claim 1, further comprising means for noise reduction and accent correction when converting voice input from a user into text using speech recognition technology.
[0965] "Example 1"
[0966] (Claim 1)
[0967] Means for receiving selection information transmitted from the user device,
[0968] A means of obtaining relevant detailed information from an information collection based on the received selection information,
[0969] A means for compressing the acquired detailed information in real time and converting it into a continuous transmission format in order to transmit it to the user device,
[0970] A means for receiving audio from a user and converting that audio data into text information,
[0971] A means of obtaining related information from the converted character information by performing an information collection search,
[0972] A means for generating a response using natural language processing based on acquired information and transmitting it to the user's device,
[0973] A means for receiving continuous audio input from a device and transmitting its contents to a server,
[0974] A system that includes this.
[0975] (Claim 2)
[0976] The system according to claim 1, wherein the user device has means for rendering the received compressed detailed information in real time and displaying it as a three-dimensional visual display and an interactive map.
[0977] (Claim 3)
[0978] The system according to claim 1, which uses speech recognition technology to convert speech from a user into text information, and has means for removing sound interference and correcting pronunciation.
[0979] "Application Example 1"
[0980] (Claim 1)
[0981] A means for receiving object information transmitted from a user device,
[0982] A means for obtaining detailed data of the corresponding object from a data storage device based on the received object information,
[0983] A means for immediately compressing the acquired detailed data and converting it into a continuous playback format in order to transmit it to the user's device,
[0984] A means for receiving voice input from the user and converting that voice data into text information,
[0985] A means for obtaining related information from the converted character information by performing a data storage device search,
[0986] A means of generating an answer using machine learning processing based on acquired information and sending it to the user's device,
[0987] A means of visualizing objects in a three-dimensional virtual environment on the user's device, allowing the user to search for information using voice and obtain immediate answers.
[0988] A system that includes this.
[0989] (Claim 2)
[0990] The system according to claim 1, wherein the user device has means for immediately rendering the received compressed detailed data and displaying it as a three-dimensional visual and an interactive map.
[0991] (Claim 3)
[0992] The system according to claim 1, which uses speech understanding technology to convert voice input from a user into text information, and has means for noise reduction and dialect correction.
[0993] "Example 2 of combining an emotion engine"
[0994] (Claim 1)
[0995] A means for receiving property selection information transmitted from a user terminal,
[0996] A means for obtaining detailed information about the relevant property from a data storage medium based on the received property selection information,
[0997] A means for compressing the acquired detailed information in real time and converting it into a data streaming format in order to send it to the user's terminal,
[0998] A means for receiving voice input from a user and converting that voice data into text,
[0999] A method for performing sentiment analysis on converted text,
[1000] A means of dynamically adjusting the interface and changing the way information is presented according to the user's emotional state,
[1001] A system that includes this.
[1002] (Claim 2)
[1003] The system according to claim 1, wherein the user terminal has means for rendering the received compressed detailed information in real time and displaying it as a spatial digital representation and an interactive map.
[1004] (Claim 3)
[1005] The system according to claim 1, further comprising means for performing noise reduction and pitch correction, and further performing sentiment analysis, when converting voice input from a user into text using speech recognition technology.
[1006] "Application example 2 when combining with an emotional engine"
[1007] (Claim 1)
[1008] Means for receiving real estate information transmitted from a user device,
[1009] A means for obtaining detailed information of the relevant property from a recording device based on the received property information,
[1010] A means for compressing the acquired detailed information in real time and converting it to a streaming format in order to transmit it to the user device,
[1011] A means for receiving voice input from a user and converting that voice data into text,
[1012] Means for obtaining relevant information from the converted text by performing a recorder query,
[1013] A means for generating a response using natural language processing based on acquired information and transmitting it to the user's device,
[1014] A means of determining an emotional state by analyzing a user's voice data using emotion recognition technology,
[1015] A means of adjusting displayed information based on the user's emotions,
[1016] A system that includes this.
[1017] (Claim 2)
[1018] The system according to claim 1, wherein the user device has means for rendering the received compressed detailed information in real time and outputting it as a three-dimensional visual and interactive map.
[1019] (Claim 3)
[1020] The system according to claim 1, further comprising means for removing background noise and correcting speech features when converting voice input from a user into text using speech recognition technology. [Explanation of Symbols]
[1021] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving property selection information transmitted from a user terminal, A means of obtaining detailed information about the relevant property from a database based on the received property selection information, A means for compressing the acquired detailed information in real time and converting it to a streaming format in order to send it to the user's terminal, A means for receiving voice input from a user and converting that voice data into text, A means of obtaining relevant information from the converted text by executing a database query, A means of generating a response using natural language processing based on acquired information and sending it to the user's terminal, A system that includes this.
2. The system according to claim 1, wherein the user terminal has means for rendering the received compressed detailed information in real time and displaying it as a three-dimensional visual and an interactive map.
3. The system according to claim 1, further comprising means for noise reduction and accent correction when converting voice input from a user into text using speech recognition technology.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A