System
The system addresses real-time map generation by using a wearable device with a camera to capture and analyze video data, generating and displaying digital maps with surrounding information, enhancing user experience through real-time updates.
Patent Information
- Application Number
- JP2024120603
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
AI Technical Summary
Conventional map data provision systems lack real-time capabilities, making it difficult to provide instant information based on a user's current location, and there are technical challenges in dynamically generating and updating digital maps to improve the user experience.
A system that utilizes a wearable device with a built-in camera to capture video data, which is analyzed by a server to determine the user's location, generating a real-time digital map that includes surrounding information, and overlays it on the user's field of view, with data compression for efficient transmission.
Enables users to visually confirm the latest information in real-time, improving convenience and user experience by providing immediate and accurate digital map data, including nearby facilities, traffic, and events.
Smart Images

Figure 2026019194000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional map data provision systems lacked real-time capabilities, making it difficult to provide instant information based on a user's current location. Furthermore, the dynamic generation and updating of digital maps to improve the user's visual experience was insufficient. As a result, users had to rely on existing digital maps and navigation systems to obtain the latest information for a specific location, resulting in an unsatisfactory user experience. Furthermore, there were technical challenges in effectively processing dynamic data over a wide range in real time. [Means for solving the problem]
[0005] To address the above-mentioned issues, the present invention provides a system for acquiring video data from a wearable device equipped with a built-in camera device, and receiving and analyzing the video data and location data on a server. Specifically, the system includes means for analyzing the video data transmitted from the wearable device and determining the user's current location, and means for generating a digital map in real time based on the determined current location. The system also includes means for transmitting the generated digital map to the wearable device and overlaying it on the user's field of view. This system provides real-time digital map data including surrounding area, traffic, and event information, enabling immediate and accurate information provision according to the user's location. Furthermore, the system can improve the efficiency of data transmission by compressing the video data. This allows users to always visually confirm the latest information, significantly improving convenience and the user experience when traveling.
[0006] A "camera device" is a device that optically captures images and records them as digital data.
[0007] A "wearable device" is an electronic device that can be worn by the user, is multifunctional, and can be used while on the move.
[0008] "Video data" refers to visual information captured by a camera device expressed in digital form.
[0009] "Location data" is numerical data indicating a specific point or area obtained from a location information system such as a GPS.
[0010] A "server" is a computer system that provides services over a network and has the function of receiving and processing data from clients.
[0011] "Analysis" is a method of calculation and processing to process digital data and extract meaningful information.
[0012] "Real-time" means that data is acquired, analyzed, and provided almost simultaneously, allowing users to receive an immediate response.
[0013] A "digital map" is map data that displays geographic information in a digital format and that can have additional information overlaid on it.
[0014] "Overlay display" is a technique for displaying additional information or visual elements overlaid on top of underlying data.
[0015] "Nearby information" is information about stores, facilities, transportation, etc. that exist around the current location.
[0016] "Traffic information" refers to real-time data on traffic, such as road congestion and the operation status of public transportation.
[0017] "Event information" is information about events, activities, and occasions taking place in a particular area.
[0018] "Compression" is a process for reducing data size and is a technology that enables efficient data transmission. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6]FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram illustrating a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0041] overview
[0042] This system allows users to wear wearable devices such as smart glasses and receive real-time updated digital map information in their field of vision. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and sends it to the wearable device.
[0043] Program processing
[0044] Video acquisition and transmission
[0045] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[0046] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[0047] Video analysis and location
[0048] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[0049] Server: Integrates location data with video analytics results to determine the user's exact location.
[0050] Digital Map Generation
[0051] Server: Based on the location data, a real-time digital map is generated, containing up-to-date information on nearby facilities, traffic, and events.
[0052] Server: Optionally overlay additional data layers (e.g., top-rated restaurants or tourist attractions) on the map.
[0053] Digital map transmission and display
[0054] Server: Compresses the latest digital map data and sends it to the wearable device.
[0055] Terminal: Receives compressed map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[0056] Specific examples
[0057] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[0058] 1. When the user is strolling around a tourist spot
[0059] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[0060] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[0061] Server: Sends map data containing the latest tourist spot information to the smart glasses.
[0062] Device: Overlays tourist attractions and directions onto the user's field of view.
[0063] User: Based on this, you can enjoy sightseeing efficiently.
[0064] 2. When a user is walking around a new city
[0065] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[0066] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[0067] Server: Sends the generated map data to the smart glasses.
[0068] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[0069] Users: Using map information displayed in an easy-to-read format, they can reach their destination efficiently and enjoy exploring the city.
[0070] This will allow users to visually receive information that is updated in real time, greatly improving the convenience of travel and sightseeing.
[0071] The processing flow will be explained below.
[0072] Step 1:
[0073] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[0074] Step 2:
[0075] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[0076] Step 3:
[0077] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[0078] Step 4:
[0079] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[0080] Step 5:
[0081] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[0082] Step 6:
[0083] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0084] Step 7:
[0085] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[0086] Step 8:
[0087] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[0088] Step 9:
[0089] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[0090] Step 10:
[0091] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[0092] Example 1
[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0094] Conventional digital map systems have had difficulty updating information in real time and providing users with accurate location information. As a result, users are unable to efficiently obtain the latest information on their surroundings, traffic, tourist spots, and other information, resulting in low convenience. The present invention aims to solve these problems and provide users with the latest digital map information that is updated in real time.
[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0096] In this invention, the server includes means for analyzing video data, detecting feature points and landmarks, and identifying the user's current location and orientation, means for generating a digital map in real time based on the current location, and means for compressing the generated digital map data and transmitting it to the wearable device, thereby enabling the user to visually receive a digital map that is updated in real time and includes the latest surrounding information, traffic information, and event information.
[0097] "Camera device" refers to a device for capturing video data that is built into a wearable device.
[0098] A "wearable device" is a device worn by the user that has a built-in camera device and GPS module and has the ability to display information within the field of view.
[0099] "Video Data" means data that captures a User's field of view with a camera device and that can be stored or transmitted in digital form.
[0100] "Location data" refers to data indicating the user's current location, typically information obtained from a GPS module.
[0101] A "GPS module" is a hardware component for obtaining geographical location information.
[0102] "Server" refers to a remote computer system that receives data sent from a wearable device and performs analysis and data generation.
[0103] "Feature points" are important points in an image detected by video analysis algorithms and used to identify objects or landmarks.
[0104] A "landmark" refers to a distinctive feature such as a building or terrain that serves as a landmark to identify the user's location and orientation.
[0105] "Real-time" means that data is acquired, processed, and displayed instantly, with minimal delay.
[0106] A "digital map" is a map containing electronically generated geographic information that is used to display a user's current location and surrounding area.
[0107] "Nearby information" refers to information about facilities and features located near the user's current location.
[0108] "Traffic information" refers to information related to travel, such as road congestion conditions and public transportation operation information.
[0109] "Event Information" refers to information about events or activities taking place at a particular location.
[0110] "Compression" is a process for reducing the volume of data and is used to improve transmission efficiency.
[0111] "Overlay" refers to the technology of superimposing information onto the field of view.
[0112] MODE FOR CARRYING OUT THE INVENTION
[0113] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0114] System configuration
[0115] This system consists of the following elements:
[0116] Wearable devices
[0117] 1. Camera Device: A device that captures the user's field of view in real time. The camera device captures high-resolution video and stores or transmits the video data in digital format.
[0118] 2. GPS module: A hardware component for obtaining geographical location information, accurately determining the user's current location.
[0119] 3. Display: A device that displays digital maps and other information, with the ability to overlay it on the user's field of view.
[0120] 4. Communication module: This module transmits video data and location data to the server.
[0121] server
[0122] 1. Data Reception Module: This module receives video and location data sent from the wearable device. The received data is compressed and decoded into H.264 or NMEA format.
[0123] 2. Video analysis algorithm: Using the OpenCV library, feature points and landmarks are detected from the received video data, which allows the user's current position and orientation to be determined.
[0124] 3. Position and orientation determination module: Integrates video analysis results and position data, and uses a Kalman filter to estimate the user's exact position and orientation.
[0125] 4. Digital Map Generation Engine: Generates a digital map in real time based on the user's current location, querying and retrieving information about surroundings, traffic, and events from a database (e.g., PostGIS).
[0126] 5. Data compression module: Compresses the generated digital map data in JSON format and sends it to the wearable device.
[0127] 6. Communication module: A device that transmits compressed digital map data to the wearable device.
[0128] Specific examples of implementation
[0129] When the user is exploring a tourist spot
[0130] As the user walks around a tourist spot, the camera in the smart glasses captures the user's field of view, compresses it, and sends it to the server. At the same time, the GPS module acquires the user's current location and sends it to the server. The server analyzes the received video data, detecting feature points and landmarks to determine the user's current location and orientation. A real-time digital map containing up-to-date tourist spot and route information is then sent to the wearable device. The wearable device then overlays the received information on the user's field of view, allowing the user to enjoy sightseeing efficiently.
[0131] Prompt Sentence Examples
[0132] "Explain how you can provide real-time updated digital map information for a scenario where a user is exploring a tourist destination."
[0133] This invention allows users to visually receive digital maps that are updated in real time and include the latest local, traffic, and event information, greatly improving convenience.
[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0135] Step 1: Capture video data
[0136] Terminal: A camera device captures the user's field of view in real time. The input is an image based on the user's field of view, and the output is video data converted into a digital format. This video data has a resolution of 640x480 pixels and is acquired at 30 frames per second (fps).
[0137] Step 2: Compressing the video data
[0138] Terminal: Compresses the acquired video data in H.264 format. The input is the captured raw video data, and the output is compressed video data in H.264 format. Compression reduces the data size and improves transmission efficiency.
[0139] Step 3: Getting location data
[0140] Device: The built-in GPS module obtains the user's current location. The input is signals from GPS satellites, and the output is location data encoded in NMEA format. This data is obtained by obtaining latitude and longitude information, updated every second.
[0141] Step 4: Sending data
[0142] Terminal: Sends compressed video data and acquired location data to the server. The input is compressed video data and location data, and the output is the data sent to the server. Video data is sent using the UDP protocol, and location data is sent using the TCP protocol to ensure communication reliability and efficiency.
[0143] Step 5: Receive and decode the data
[0144] Server: Decodes the received compressed video data using an H.264 decoder and analyzes the location data. The input is the compressed video data and location data, and the output is the decoded video data and analyzed location information. The decoding and analysis process makes it possible to visualize the video and location.
[0145] Step 6: Video analysis and feature point detection
[0146] Server: Extracts feature points (e.g., ORB feature points) from input video data using the OpenCV library. The input is the decoded video data, and the output is the detected feature points and landmark information. Calculates the user's orientation based on the feature points and identifies landmarks.
[0147] Step 7: Identifying position and orientation
[0148] Server: Integrates feature point and location data and uses a Kalman filter to estimate the user's precise location and orientation. The input is feature point information and latitude and longitude data, and the output is estimated location and orientation information. Identifying the location and orientation clarifies the user's spatial state.
[0149] Step 8: Generate a digital map
[0150] Server: Based on the user's current location, obtains surrounding information, traffic information, and event information from a database and generates a real-time digital map. The input is the estimated location and orientation information and a database query, and the output is the generated digital map data. The map generation engine integrates multi-layer data.
[0151] Step 9: Compressing the digital map
[0152] Server: Compresses the generated digital map data in JSON format. The input is the generated real-time map data, and the output is compressed map data in JSON format. Compression improves the efficiency of data transmission.
[0153] Step 10: Sending the digital map
[0154] Server: Sends compressed digital map data to the wearable device. The input is the compressed data, and the output is the data sent to the wearable device. The transmission uses the TCP protocol to ensure data integrity.
[0155] Step 11: Developing the digital map
[0156] Terminal: Decompresses the received compressed data and renders it for display. The input is compressed map data, and the output is a real-time digital map displayed in the user's field of view.
[0157] Step 12: Digital map overlay display
[0158] Terminal: The expanded map data is overlaid on the user's field of view. The input is the expanded map data, and the output is the map and related information displayed in the user's field of view. This allows the user to visually check the latest map information.
[0159] (Application example 1)
[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0161] Conventional autonomous vehicle systems have struggled to efficiently grasp real-time changes in road conditions and traffic information to improve the safety and efficiency of autonomous driving. Furthermore, the effectiveness of navigation systems is limited because the information obtained while the user is actually driving is only retained for a short period of time. Therefore, there is a need for a system that can reflect the latest information about the vehicle's surroundings in real time, improving the navigation accuracy and safety of autonomous vehicles.
[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0163] In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for displaying the digital map on the display of the wearable device, means for acquiring the vehicle's surroundings from the camera device of the autonomous vehicle and dynamically analyzing the surrounding environment based on location information, and means for reflecting the analysis results in the autonomous vehicle's navigation system in real time. This allows the autonomous vehicle to grasp the latest road conditions and traffic information in real time, enabling safe and efficient operation in response to changes in the environment around the vehicle.
[0164] A "camera device" is an optical device for capturing video data.
[0165] A "wearable device" is an electronic device that can be worn by a user and that acquires and displays data.
[0166] "Video data" is visual information captured using a camera device.
[0167] "Location data" is information indicating the current location obtained using a positioning system such as GPS.
[0168] The "receiving means" refers to a device or mechanism for receiving data from the outside.
[0169] "Means for analyzing" are devices and mechanisms for interpreting and converting acquired data into meaningful information.
[0170] "Current location" means the location where the user or autonomous vehicle is currently located.
[0171] A "digital map" is electronic map data that visually represents location information and related information.
[0172] "Generating means" are devices and mechanisms for producing new data or information.
[0173] A "transmitting means" is a device or mechanism for transmitting data to another device or system.
[0174] A "display" is a device for visually displaying information.
[0175] An "autonomous vehicle" is a vehicle that operates autonomously using artificial intelligence and sensor technology.
[0176] "Surrounding conditions" refers to information about the environment in which an autonomous vehicle operates.
[0177] A "navigation system" is a device or mechanism that guides a vehicle to a destination or provides route guidance based on the vehicle's current location.
[0178] "Dynamic analysis means" refers to devices and mechanisms that analyze data in real time and take immediate action or make decisions based on the results.
[0179] System Configuration
[0180] In this invention, the invention is implemented using a system consisting of the following components:
[0181] 1. Hardware
[0182] Camera device: High-resolution cameras installed in autonomous vehicles are used to acquire video data of the surrounding area.
[0183] GPS module: Uses the GPS system installed in the vehicle to determine the current location.
[0184] Wearable device: A device worn by the user, such as smart glasses, that has a built-in camera and display.
[0185] Server: A high-performance server (e.g. AWS EC2) that receives data, analyzes it, and generates digital maps.
[0186] Communication network: Internet connection for sending and receiving data in real time (e.g., 5G communication).
[0187] 2. Software
[0188] Video analysis algorithm: OpenCV and TensorFlow are used to analyze video data and detect feature points and landmarks.
[0189] Digital map generation software: A program that generates digital maps in real time based on the data received.
[0190] Communication protocol: Protocols such as HTTP and WebSocket for sending and receiving data.
[0191] Overview of operation procedures
[0192] When a user operates an autonomous vehicle, the system captures video and location data from the camera device and GPS module. The data is then compressed and sent to a server. The server analyzes the received data to determine the user's surroundings and current location. The server then generates a digital map in real time and sends the data to the autonomous vehicle's navigation system.
[0193] Data processing details
[0194] 1. Acquisition and transmission of video data: Video data acquired by the camera device is compressed and encoded in real time and transmitted to the server along with GPS data.
[0195] 2. Server-side data analysis: The server analyzes the transmitted video data in real time, detecting features and landmarks to determine the current location. This process uses video analysis algorithms such as OpenCV and TensorFlow.
[0196] 3. Digital map generation: Based on the analysis results and location data, the server generates a digital map that reflects the latest information about the surrounding environment, including traffic conditions, obstacle information, and event information.
[0197] 4. Data transmission and display: The generated digital map is compressed and sent in real time to the autonomous vehicle's navigation system, which then adjusts the vehicle's route accordingly based on the received map information.
[0198] Specific examples
[0199] For example, when a vehicle is traveling on a highway, the camera device and GPS module send video and location data to a server. The server then analyzes the video to detect road congestion and obstacles. Based on this information, the server then generates an optimal driving route in real time and sends it to the vehicle's navigation system, enabling the autonomous vehicle to operate safely and efficiently.
[0200] Prompt Sentence Examples
[0201] "Use your location and video data to generate digital map information that updates in real time, providing up-to-date navigation, including traffic and obstacle information."
[0202] This invention can significantly improve the operational safety and efficiency of autonomous vehicles.
[0203] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0204] Step 1:
[0205] The camera device of the autonomous vehicle captures the surrounding video data in real time. The input from the camera device is raw video frames, which are compressed in JPEG or H.264 format.
[0206] Step 2:
[0207] The GPS module of the autonomous vehicle acquires the current location data. The input from the GPS module is latitude and longitude information, and this data is converted into JSON format.
[0208] Step 3:
[0209] The acquired video data and location data are compressed and integrated into a single concatenated data packet at the autonomous vehicle's terminal. The output from the terminal is a concatenated data of compressed video frames and location information.
[0210] Step 4:
[0211] The terminal sends compressed and consolidated data packets to the server. The data is sent through the communication network and received by the server. The server receives the compressed data packets.
[0212] Step 5:
[0213] The server decodes the received data packets and separates the video data from the location data. The decoded input data is the raw video frame and the current location information.
[0214] Step 6:
[0215] The server uses OpenCV and TensorFlow to analyze the video data. The analysis detects landmarks and feature points, which are then used to identify the user's surroundings. The input is raw video data, and the output is analyzed video information (landmarks and feature points).
[0216] Step 7:
[0217] The server combines the received location data with the video analysis results to determine the exact current location. The output of the combination process is accurate current location information.
[0218] Step 8:
[0219] The server generates a digital map in real time based on the current location data. This process includes surrounding information, traffic information, and event information, and also uses existing map data such as OpenStreetMap. The input is precise current location information and additional data (traffic and event information), and the output is digital map data generated in real time.
[0220] Step 9:
[0221] The server compresses the generated digital map and sends it back to the autonomous vehicle's terminal. The data is transmitted over the communication network. The output of the server is compressed digital map data.
[0222] Step 10:
[0223] The terminal decodes the received compressed digital map data and transmits it to the navigation system of the autonomous vehicle. The input of the terminal is compressed digital map data, and the output is decoded map data.
[0224] Step 11:
[0225] The navigation system displays the decoded map data and adjusts the vehicle's driving route in real time, with the input of the navigation system being the decoded map data and the output being the adjusted driving route.
[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0227] This invention relates to a system that not only provides digital map data in real time but also recognizes the user's emotions and customizes information based on those emotions. This system consists of a wearable terminal with a built-in camera device, a server with an emotion engine, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0228] overview
[0229] Using this system, users can wear wearable devices such as smart glasses and receive not only real-time updated digital map information in their field of vision, but also information customized to the user's emotions. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and then uses an emotion engine to provide information appropriate to the user's emotions.
[0230] Program processing
[0231] Video acquisition and transmission
[0232] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[0233] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[0234] Video analysis and location
[0235] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[0236] Server: Integrates location data with video analytics results to determine the user's exact location.
[0237] Emotion Recognition and Analysis
[0238] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[0239] Server: Based on the recognized emotional state, the digital map display content related to the user's current location and surrounding information is dynamically changed.
[0240] Digital Map Generation
[0241] Server: Based on the location data mentioned above, obtains geographic information around the current location, including surrounding buildings, roads, points of interest (POI), etc.
[0242] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0243] Emotion-based information customization
[0244] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[0245] Digital map transmission and display
[0246] Server: Compresses the latest digital map data and sends it to the wearable device, performing the necessary checks to ensure data integrity.
[0247] Terminal: Receives compressed digital map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[0248] Specific examples
[0249] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[0250] 1. When the user is strolling around a tourist spot
[0251] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[0252] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[0253] Server: Using the emotion engine, analyze the user's emotional state from their facial expressions and voice and recognize that they are relaxed.
[0254] Server: Generates map data containing information about scenic areas and photo spots where users relax.
[0255] Server: Sends the generated digital map data to the smart glasses.
[0256] Device: Overlays tourist attractions and directions onto the user's field of view.
[0257] User: Based on this, you can enjoy sightseeing efficiently.
[0258] 2. When a user is walking around a new city
[0259] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[0260] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[0261] Server: Using an emotion engine, the server analyzes the user's emotional state from their facial expressions and voice and recognizes that they are tired.
[0262] Server: Because the user is tired, it generates map data containing information about cafes and rest spots where the user can relax.
[0263] Server: Sends the generated map data to the smart glasses.
[0264] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[0265] Users: Map information is displayed in an easy-to-read format, allowing them to reach their destination efficiently and enjoy exploring the city.
[0266] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, greatly improving the convenience of travel and sightseeing.
[0267] The processing flow will be explained below.
[0268] Step 1:
[0269] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[0270] Step 2:
[0271] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[0272] Step 3:
[0273] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[0274] Step 4:
[0275] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[0276] Step 5:
[0277] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[0278] Step 6:
[0279] Server: Based on the recognized emotional state, the content displayed on the digital map, including the user's current location and surrounding information, is dynamically changed.
[0280] Step 7:
[0281] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[0282] Step 8:
[0283] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0284] Step 9:
[0285] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[0286] Step 10:
[0287] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[0288] Step 11:
[0289] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[0290] Step 12:
[0291] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[0292] Step 13:
[0293] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[0294] Step 14:
[0295] Users: Receive personalized information based on their emotions and take action accordingly. For example, if they are feeling stressed, they can find a relaxation spot to help them relax.
[0296] Example 2
[0297] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0298] Conventional navigation systems only provide geographical information and are unable to customize information based on the user's emotional state. This makes it difficult to effectively provide users with the information they need in real time. Furthermore, the analysis and integration of video and location data is insufficient, making it difficult to accurately identify the user's current location and direction.
[0299] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing video data and identifying the user's current location, a means for analyzing the user's facial expression and voice data and recognizing the user's emotional state, and a means for customizing information based on the emotional state. This makes it possible to provide information in real time according to the user's emotional state and to identify the user's location accurately.
[0300] A "camera device" is an imaging device built into a wearable device to capture the user's field of view.
[0301] A "wearable device" is a device that can be worn by a user and has a built-in camera device and GPS module.
[0302] "Video data" is data in digital form of a field of view captured by a camera device.
[0303] "Location Data" means information about a user's current location obtained using a GPS module or other location-determining technology.
[0304] "Analysis means" refers to software or hardware functionality for processing received data and extracting specific information.
[0305] "Emotional state" refers to emotional information such as joy, sadness, surprise, etc., obtained by analyzing the user's facial expression and voice data.
[0306] A "digital map" is a digital map that contains geographic information about the surrounding area that is dynamically generated based on the user's current location and facing direction.
[0307] "Nearby information" is information about buildings, roads, points of interest (POIs), etc., in the vicinity of the user's current location.
[0308] "Traffic information" refers to real-time traffic information related to the user's current location and route.
[0309] "Event information" is information about events being held in the vicinity of the user's current location.
[0310] "Compression means" refers to data compression algorithms and techniques for efficiently reducing the volume of large amounts of data.
[0311] An "analysis algorithm" is a specific calculation method or procedure used to analyze video, audio, location data, etc.
[0312] An "emotion recognition engine" is software or hardware that identifies a user's emotional state from facial expressions and voice data.
[0313] A "display" is a display device built into a wearable device that visually presents digital maps and other information to the user.
[0314] This invention is a system that uses a wearable device to analyze a user's movement information and emotional state in real time and provides a customized digital map based on the analysis. This system consists of a wearable device with a built-in camera device, a server with an emotion recognition engine, a server that receives and analyzes video data and location data, and a wearable device for displaying the generated digital map.
[0315] Wearable devices are devices with built-in cameras, such as smart glasses, that capture the user's field of vision and obtain it as video data. This data is compressed in a video compression format such as H.264 and sent to a server. Wearable devices also have built-in GPS modules that obtain the user's current location every second and send this location data to the server. Communication is via 4G / 5G networks, and data is transmitted using the TCP / IP protocol.
[0316] The server runs image processing algorithms on the received video data, for example using image processing libraries such as OpenCV to detect features and landmarks in the video. These features are then combined with the received GPS location data to determine the user's exact location and facing direction using triangulation.
[0317] The server then uses an emotion recognition engine to analyze the user's facial expressions and voice data to recognize their emotional state. The engine can be Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API. Based on this emotional state, the server customizes the information provided to the user.
[0318] Based on the user's current location data, the server obtains surrounding geographical, traffic, and event information from sources such as Google Maps API and OpenStreetMap. This information is integrated to generate a digital map that is updated in real time. Information based on the user's emotional state is displayed on the generated digital map. For example, if the user is tired, nearby cafes and relaxation spots are displayed first, while if the user is happy, tourist spots and event information are displayed first.
[0319] The server compresses the generated digital map data and efficiently transmits it to the wearable device using protocol buffers. The wearable device receives the data, expands the compressed data, and overlays it on the display. This allows the user to view the latest map information and customized information in real time in their field of vision.
[0320] As a concrete example, when a user is strolling around a tourist spot, the smart glasses capture video of the user's field of view and send it to a server. The server analyzes the video data and identifies the user's current location. It then uses an emotion recognition engine to analyze the user's emotional state and recognizes that the user is relaxed. The server generates map data including information on beautiful scenery and photo spots at the tourist spot and sends it to the wearable device. The user can then view tourist spots and route information in their field of view, allowing them to enjoy sightseeing efficiently.
[0321] An example of a prompt is as follows:
[0322] "Find a nearby cafe from your current location"
[0323] "Show me photo spots for tourist attractions"
[0324] "Tell me about traffic in this area."
[0325] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0326] Step 1:
[0327] Terminal: The camera device of the smart glasses captures the user's field of view. The captured video data is compressed into H.264 format by an encoder. This compressed video data is the input. The video data is transmitted to the server via a 4G / 5G network using the TCP / IP protocol. The output is the compressed video data that has arrived at the server.
[0328] Step 2:
[0329] Terminal: The GPS module acquires the user's current location every second. The acquired location data is recorded in NMEA format. This location data is the input and is sent to the server, also using the TCP / IP protocol. The output is the location data that arrives at the server.
[0330] Step 3:
[0331] Server: The server analyzes the video data it receives using an image processing library such as OpenCV. The input is compressed video data, which is first decoded and expanded frame by frame. An image processing algorithm (for example, ORB (Oriented FAST and Rotated BRIEF)) is applied to detect feature points and landmarks in the video. The results of this analysis are output.
[0332] Step 4:
[0333] Server: The server combines the feature points of the analyzed video data with the received location data. The input is feature point data and location data. These are combined using triangulation to determine the user's exact location and facing direction. The result of this processing is the output.
[0334] Step 5:
[0335] Server: The server uses an emotion recognition engine to analyze the user's facial expression and voice data. The input is the user's facial expression data and voice data. Software such as Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API is used to identify the emotional state (happiness, sadness, surprise, etc.). The results of this analysis are the output.
[0336] Step 6:
[0337] Server: Customizes information based on the recognized emotional state. Inputs are emotion recognition results and the user's location data. Based on the user's current location, the server obtains surrounding geographical information, traffic information, and event information from Google Maps API, OpenStreetMap, etc., and generates a digital map. The output is customized digital map data.
[0338] Step 7:
[0339] Server: Compresses the generated digital map data and transmits it efficiently to the wearable device using Protocol Buffers. The input is the customized digital map data. The output is transmitted as compressed data.
[0340] Step 8:
[0341] Terminal: The wearable terminal decompresses the compressed digital map data received from the server and overlays it on the display. The input is compressed digital map data. By displaying the decompressed data on the display, the user can see the latest information in real time in their field of vision. This visual information is the final output.
[0342] (Application example 2)
[0343] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0344] In recent years, advances in wearable devices and emotion analysis technology have garnered attention for providing information tailored to user emotions. However, systems that combine these technologies to provide optimal route guidance in real time for autonomous vehicles are still in the early stages of development. Existing systems struggle to efficiently integrate a user's emotional state and current location information to provide customized information in real time. Therefore, the present invention aims to provide real-time digital map information customized based on the user's emotional state, thereby improving passenger comfort in autonomous vehicles.
[0345] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for analyzing the user's emotional state using an emotion engine, means for customizing the digital map based on the analyzed emotional state, means for displaying the customized digital map on the display of the wearable device, and means installed in an autonomous vehicle for providing optimal route guidance based on the passenger's real-time emotion data and current location. This makes it possible to integrate the user's emotional state and current location information and provide appropriate route guidance and surrounding information within the autonomous vehicle.
[0346] A "camera device" is a device for acquiring video data and is built into a wearable device.
[0347] A "wearable terminal" is a device that can be worn by a user and has a built-in camera device and display.
[0348] "Video Data" means a digital representation of visual information captured by a camera device.
[0349] "Location data" refers to data indicating the current location of a user or autonomous vehicle, and is typically obtained by a GPS module.
[0350] A "server" is a high-performance computing device that receives and analyzes video data and location data.
[0351] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice data to recognize their emotional state.
[0352] A "digital map" is map data that includes geographical information, surrounding area information, traffic information, and other information that is generated in real time.
[0353] "Real-time" refers to the time range in which data is processed and updated immediately, with little to no delay.
[0354] "Customizing" means adjusting or changing the information and functionality provided based on information such as the user's emotional state or current location.
[0355] "Display" means a display device built into a wearable device that visually presents digital maps and other information to the user.
[0356] An "autonomous vehicle" is a vehicle that can move independently without human operation.
[0357] "Route guidance" is guide information that shows the optimal route or method to a destination.
[0358] The present invention relates to a system including a wearable terminal with a built-in camera device, a server with an emotion engine, a server for analyzing video data and location data, and a wearable terminal for displaying a generated digital map, which is installed in an autonomous vehicle and provides customized real-time route guidance based on the passenger's emotional state and current location.
[0359] System configuration
[0360] Wearable device: A device such as smart glasses that incorporates a camera device, display, and GPS module.
[0361] Server: A high-performance server, such as Amazon EC2, is used. An emotion engine (e.g., Affectiva), a video analysis algorithm (e.g., OpenCV), and a data compression library (e.g., zlib) are used.
[0362] Autonomous vehicle: A vehicle that is operated autonomously and carried by a user.
[0363] Program processing
[0364] Video acquisition and transmission
[0365] Terminal: The smart glasses' camera device captures the user's field of view and acquires it as video data. The acquired video data is compressed and sent to the server. In addition, the current location data is acquired from the GPS module and sent to the server.
[0366] Video analysis and location
[0367] Server: Receives video data and analyzes it using a video analytics algorithm (e.g., OpenCV) to detect features and landmarks within the field of view. This data is then combined with GPS location data to determine the current location of the autonomous vehicle.
[0368] Emotion Recognition and Analysis
[0369] Server: An emotion engine (e.g., Affectiva) is used to analyze the user's facial expressions and voice data to recognize their emotional state, which can include happiness, sadness, surprise, etc.
[0370] Digital map generation and customization
[0371] Server: Based on the current location data, the server obtains surrounding geographic information, traffic information, and specific points (restaurants, tourist attractions, rest spots, etc.) in real time and generates a digital map. Furthermore, the server customizes the information displayed based on the user's emotional state, as determined by emotion recognition. For example, if the user is relaxed, it will prioritize displaying tourist attractions and relaxation spots, while if the user is feeling stressed, it will prioritize providing information on the shortest route and rest spots.
[0372] Digital map transmission and display
[0373] Server: Sends compressed digital map data to the device.
[0374] Device: Digital map data is deployed and overlaid on the smart glasses display, allowing users to visually check the latest information in real time.
[0375] Specific usage scenarios
[0376] 1. Tourism scenario:
[0377] When a user walks around a tourist spot, the camera device in the smart glasses captures video data and sends it to a server.
[0378] The server analyzes the video data to determine the current location and transmits map data including information on tourist attractions and directions.
[0379] If the emotion engine determines that the user is relaxed, it will provide customized map data including information on scenic areas and photo spots.
[0380] 2. Urban Walking Scenario:
[0381] As a user walks around a new city, the camera device in the smart glasses captures video data and sends it to a server.
[0382] The server analyzes the video data, determines the current location, and then provides map data including information on nearby popular spots and stores.
[0383] If emotion recognition determines that the user is tired, it will prioritize providing information about relaxation spots and cafes.
[0384] Prompt Sentence Examples
[0385] Here are some examples of prompts for a generative AI model:
[0386] If the user is relaxed:
[0387] "Current emotional state: Relaxed. Prioritize information about nearby tourist attractions and relaxation spots."
[0388] If the user is stressed:
[0389] "Current emotional state: Stress. Prioritize showing quicker routes and directions to relaxation spots."
[0390] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, making travel and sightseeing more comfortable and enjoyable.
[0391] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0392] Step 1:
[0393] Video acquisition and transmission
[0394] The device (smart glasses) uses a camera device to acquire video data within the user's field of view, including the scenery and objects the user is looking at.
[0395] Input: User's field of view
[0396] Output: Video data
[0397] Specific operation: The video data is compressed and sent to the server in real time. At the same time, the current location data is obtained from the GPS module and sent to the server.
[0398] Step 2:
[0399] Video analysis and location
[0400] The server analyzes the received video data using a video analysis algorithm (e.g., OpenCV), which detects feature points and landmarks within the field of view.
[0401] Input: Video data, location data
[0402] Output: Feature point data, landmark data, current location data
[0403] What it does: Detects features and landmarks in the video and combines this data with GPS location data to determine the user's current location.
[0404] Step 3:
[0405] Emotion Recognition and Analysis
[0406] The server uses an emotion engine (e.g., Affectiva) to analyze the user's facial expressions and voice data to recognize their emotional state.
[0407] Input: User's facial expression data, voice data
[0408] Output: Emotional state data
[0409] Specific behavior: Identify emotional states such as happiness, sadness, and surprise from facial expressions and voice, and store that information in a database.
[0410] Step 4:
[0411] Digital map generation
[0412] The server obtains real-time geographical information, traffic information, and data on specific points (restaurants, tourist attractions, rest spots, etc.) based on the current location data, and generates a digital map.
[0413] Input: Current location data, surrounding geographic information, traffic information, point information
[0414] Output: Digital map data
[0415] Specific operation: Integrates information on surrounding buildings, roads, POIs (points of interest), etc. to generate an up-to-date digital map.
[0416] Step 5:
[0417] Customizing digital maps
[0418] The server customizes the information displayed based on the user's emotional state. For example, if the user is relaxed, it will prioritize showing tourist attractions and relaxation spots, while if the user is stressed, it will prioritize showing information about the shortest routes and rest spots.
[0419] Input: Digital map data, emotional state data
[0420] Output: Customized digital map data
[0421] Specific operation: Dynamically change the displayed content of the digital map according to the user's emotional state to provide the most appropriate information to the user.
[0422] Step 6:
[0423] Digital map transmission and display
[0424] The server compresses the customized digital map data and sends it to the smart glasses.
[0425] Input: Customized digital map data
[0426] Output: Compressed digital map data packets
[0427] Specific operation: Data to be overlaid on the display is compressed and sent to the terminal in real time. The terminal then decompresses the received data and displays it on the display.
[0428] Step 7:
[0429] User receipt and use
[0430] Users can check customized digital map information displayed on the smart glasses display and use it when traveling or sightseeing.
[0431] Input: Customized digital map information displayed on the display
[0432] Output: User actions and choices
[0433] Specific actions: Decide on the next action based on visually presented information, such as heading to a displayed relaxation spot.
[0434] This allows users to receive optimal route guidance and information based on real-time updates and their emotional state, making self-driving vehicles more comfortable to use.
[0435] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0436] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0437] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0438] [Second embodiment]
[0439] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0440] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0441] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0442] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0443] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0444] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0445] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0446] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0447] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0448] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0449] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0450] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0451] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0452] overview
[0453] This system allows users to wear wearable devices such as smart glasses and receive real-time updated digital map information in their field of vision. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and sends it to the wearable device.
[0454] Program processing
[0455] Video acquisition and transmission
[0456] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[0457] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[0458] Video analysis and location
[0459] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[0460] Server: Integrates location data with video analytics results to determine the user's exact location.
[0461] Digital Map Generation
[0462] Server: Based on the location data, a real-time digital map is generated, containing up-to-date information on nearby facilities, traffic, and events.
[0463] Server: Optionally overlay additional data layers (e.g., top-rated restaurants or tourist attractions) on the map.
[0464] Digital map transmission and display
[0465] Server: Compresses the latest digital map data and sends it to the wearable device.
[0466] Terminal: Receives compressed map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[0467] Specific examples
[0468] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[0469] 1. When the user is strolling around a tourist spot
[0470] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[0471] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[0472] Server: Sends map data containing the latest tourist spot information to the smart glasses.
[0473] Device: Overlays tourist attractions and directions onto the user's field of view.
[0474] User: Based on this, you can enjoy sightseeing efficiently.
[0475] 2. When a user is walking around a new city
[0476] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[0477] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[0478] Server: Sends the generated map data to the smart glasses.
[0479] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[0480] Users: Using map information displayed in an easy-to-read format, they can reach their destination efficiently and enjoy exploring the city.
[0481] This will allow users to visually receive information that is updated in real time, greatly improving the convenience of travel and sightseeing.
[0482] The processing flow will be explained below.
[0483] Step 1:
[0484] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[0485] Step 2:
[0486] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[0487] Step 3:
[0488] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[0489] Step 4:
[0490] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[0491] Step 5:
[0492] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[0493] Step 6:
[0494] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0495] Step 7:
[0496] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[0497] Step 8:
[0498] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[0499] Step 9:
[0500] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[0501] Step 10:
[0502] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[0503] Example 1
[0504] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0505] Conventional digital map systems have had difficulty updating information in real time and providing users with accurate location information. As a result, users are unable to efficiently obtain the latest information on their surroundings, traffic, tourist spots, and other information, resulting in low convenience. The present invention aims to solve these problems and provide users with the latest digital map information that is updated in real time.
[0506] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0507] In this invention, the server includes means for analyzing video data, detecting feature points and landmarks, and identifying the user's current location and orientation, means for generating a digital map in real time based on the current location, and means for compressing the generated digital map data and transmitting it to the wearable device, thereby enabling the user to visually receive a digital map that is updated in real time and includes the latest surrounding information, traffic information, and event information.
[0508] "Camera device" refers to a device for capturing video data that is built into a wearable device.
[0509] A "wearable device" is a device worn by the user that has a built-in camera device and GPS module and has the ability to display information within the field of view.
[0510] "Video Data" means data that captures a User's field of view with a camera device and that can be stored or transmitted in digital form.
[0511] "Location data" refers to data indicating the user's current location, typically information obtained from a GPS module.
[0512] A "GPS module" is a hardware component for obtaining geographical location information.
[0513] "Server" refers to a remote computer system that receives data sent from a wearable device and performs analysis and data generation.
[0514] "Feature points" are important points in an image detected by video analysis algorithms and used to identify objects or landmarks.
[0515] A "landmark" refers to a distinctive feature such as a building or terrain that serves as a landmark to identify the user's location and orientation.
[0516] "Real-time" means that data is acquired, processed, and displayed instantly, with minimal delay.
[0517] A "digital map" is a map containing electronically generated geographic information that is used to display a user's current location and surrounding area.
[0518] "Nearby information" refers to information about facilities and features located near the user's current location.
[0519] "Traffic information" refers to information related to travel, such as road congestion conditions and public transportation operation information.
[0520] "Event Information" refers to information about events or activities taking place at a particular location.
[0521] "Compression" is a process for reducing the volume of data and is used to improve transmission efficiency.
[0522] "Overlay" refers to the technology of superimposing information onto the field of view.
[0523] MODE FOR CARRYING OUT THE INVENTION
[0524] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0525] System configuration
[0526] This system consists of the following elements:
[0527] Wearable devices
[0528] 1. Camera Device: A device that captures the user's field of view in real time. The camera device captures high-resolution video and stores or transmits the video data in digital format.
[0529] 2. GPS module: A hardware component for obtaining geographical location information, accurately determining the user's current location.
[0530] 3. Display: A device that displays digital maps and other information, with the ability to overlay it on the user's field of view.
[0531] 4. Communication module: This module transmits video data and location data to the server.
[0532] server
[0533] 1. Data Reception Module: This module receives video and location data sent from the wearable device. The received data is compressed and decoded into H.264 or NMEA format.
[0534] 2. Video analysis algorithm: Using the OpenCV library, feature points and landmarks are detected from the received video data, which allows the user's current position and orientation to be determined.
[0535] 3. Position and orientation determination module: Integrates video analysis results and position data, and uses a Kalman filter to estimate the user's exact position and orientation.
[0536] 4. Digital Map Generation Engine: Generates a digital map in real time based on the user's current location, querying and retrieving information about surroundings, traffic, and events from a database (e.g., PostGIS).
[0537] 5. Data compression module: Compresses the generated digital map data in JSON format and sends it to the wearable device.
[0538] 6. Communication module: A device that transmits compressed digital map data to the wearable device.
[0539] Specific examples of implementation
[0540] When the user is exploring a tourist spot
[0541] As the user walks around a tourist spot, the camera in the smart glasses captures the user's field of view, compresses it, and sends it to the server. At the same time, the GPS module acquires the user's current location and sends it to the server. The server analyzes the received video data, detecting feature points and landmarks to determine the user's current location and orientation. A real-time digital map containing up-to-date tourist spot and route information is then sent to the wearable device. The wearable device then overlays the received information on the user's field of view, allowing the user to enjoy sightseeing efficiently.
[0542] Prompt Sentence Examples
[0543] "Explain how you can provide real-time updated digital map information for a scenario where a user is exploring a tourist destination."
[0544] This invention allows users to visually receive digital maps that are updated in real time and include the latest local, traffic, and event information, greatly improving convenience.
[0545] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0546] Step 1: Capture video data
[0547] Terminal: A camera device captures the user's field of view in real time. The input is an image based on the user's field of view, and the output is video data converted into a digital format. This video data has a resolution of 640x480 pixels and is acquired at 30 frames per second (fps).
[0548] Step 2: Compressing the video data
[0549] Terminal: Compresses the acquired video data in H.264 format. The input is the captured raw video data, and the output is compressed video data in H.264 format. Compression reduces the data size and improves transmission efficiency.
[0550] Step 3: Getting location data
[0551] Device: The built-in GPS module obtains the user's current location. The input is signals from GPS satellites, and the output is location data encoded in NMEA format. This data is obtained by obtaining latitude and longitude information, updated every second.
[0552] Step 4: Sending data
[0553] Terminal: Sends compressed video data and acquired location data to the server. The input is compressed video data and location data, and the output is the data sent to the server. Video data is sent using the UDP protocol, and location data is sent using the TCP protocol to ensure communication reliability and efficiency.
[0554] Step 5: Receive and decode the data
[0555] Server: Decodes the received compressed video data using an H.264 decoder and analyzes the location data. The input is the compressed video data and location data, and the output is the decoded video data and analyzed location information. The decoding and analysis process makes it possible to visualize the video and location.
[0556] Step 6: Video analysis and feature point detection
[0557] Server: Extracts feature points (e.g., ORB feature points) from input video data using the OpenCV library. The input is the decoded video data, and the output is the detected feature points and landmark information. Calculates the user's orientation based on the feature points and identifies landmarks.
[0558] Step 7: Identifying position and orientation
[0559] Server: Integrates feature point and location data and uses a Kalman filter to estimate the user's precise location and orientation. The input is feature point information and latitude and longitude data, and the output is estimated location and orientation information. Identifying the location and orientation clarifies the user's spatial state.
[0560] Step 8: Generate a digital map
[0561] Server: Based on the user's current location, obtains surrounding information, traffic information, and event information from a database and generates a real-time digital map. The input is the estimated location and orientation information and a database query, and the output is the generated digital map data. The map generation engine integrates multi-layer data.
[0562] Step 9: Compressing the digital map
[0563] Server: Compresses the generated digital map data in JSON format. The input is the generated real-time map data, and the output is compressed map data in JSON format. Compression improves the efficiency of data transmission.
[0564] Step 10: Sending the digital map
[0565] Server: Sends compressed digital map data to the wearable device. The input is the compressed data, and the output is the data sent to the wearable device. The transmission uses the TCP protocol to ensure data integrity.
[0566] Step 11: Developing the digital map
[0567] Terminal: Decompresses the received compressed data and renders it for display. The input is compressed map data, and the output is a real-time digital map displayed in the user's field of view.
[0568] Step 12: Digital map overlay display
[0569] Terminal: The expanded map data is overlaid on the user's field of view. The input is the expanded map data, and the output is the map and related information displayed in the user's field of view. This allows the user to visually check the latest map information.
[0570] (Application example 1)
[0571] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0572] Conventional autonomous vehicle systems have struggled to efficiently grasp real-time changes in road conditions and traffic information to improve the safety and efficiency of autonomous driving. Furthermore, the effectiveness of navigation systems is limited because the information obtained while the user is actually driving is only retained for a short period of time. Therefore, there is a need for a system that can reflect the latest information about the vehicle's surroundings in real time, improving the navigation accuracy and safety of autonomous vehicles.
[0573] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0574] In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for displaying the digital map on the display of the wearable device, means for acquiring the vehicle's surroundings from the camera device of the autonomous vehicle and dynamically analyzing the surrounding environment based on location information, and means for reflecting the analysis results in the autonomous vehicle's navigation system in real time. This allows the autonomous vehicle to grasp the latest road conditions and traffic information in real time, enabling safe and efficient operation in response to changes in the environment around the vehicle.
[0575] A "camera device" is an optical device for capturing video data.
[0576] A "wearable device" is an electronic device that can be worn by a user and that acquires and displays data.
[0577] "Video data" is visual information captured using a camera device.
[0578] "Location data" is information indicating the current location obtained using a positioning system such as GPS.
[0579] The "receiving means" refers to a device or mechanism for receiving data from the outside.
[0580] "Means for analyzing" are devices and mechanisms for interpreting and converting acquired data into meaningful information.
[0581] "Current location" means the location where the user or autonomous vehicle is currently located.
[0582] A "digital map" is electronic map data that visually represents location information and related information.
[0583] "Generating means" are devices and mechanisms for producing new data or information.
[0584] A "transmitting means" is a device or mechanism for transmitting data to another device or system.
[0585] A "display" is a device for visually displaying information.
[0586] An "autonomous vehicle" is a vehicle that operates autonomously using artificial intelligence and sensor technology.
[0587] "Surrounding conditions" refers to information about the environment in which an autonomous vehicle operates.
[0588] A "navigation system" is a device or mechanism that guides a vehicle to a destination or provides route guidance based on the vehicle's current location.
[0589] "Dynamic analysis means" refers to devices and mechanisms that analyze data in real time and take immediate action or make decisions based on the results.
[0590] System Configuration
[0591] In this invention, the invention is implemented using a system consisting of the following components:
[0592] 1. Hardware
[0593] Camera device: High-resolution cameras installed in autonomous vehicles are used to acquire video data of the surrounding area.
[0594] GPS module: Uses the GPS system installed in the vehicle to determine the current location.
[0595] Wearable device: A device worn by the user, such as smart glasses, that has a built-in camera and display.
[0596] Server: A high-performance server (e.g. AWS EC2) that receives data, analyzes it, and generates digital maps.
[0597] Communication network: Internet connection for sending and receiving data in real time (e.g., 5G communication).
[0598] 2. Software
[0599] Video analysis algorithm: OpenCV and TensorFlow are used to analyze video data and detect feature points and landmarks.
[0600] Digital map generation software: A program that generates digital maps in real time based on the data received.
[0601] Communication protocol: Protocols such as HTTP and WebSocket for sending and receiving data.
[0602] Overview of operation procedures
[0603] When a user operates an autonomous vehicle, the system captures video and location data from the camera device and GPS module. The data is then compressed and sent to a server. The server analyzes the received data to determine the user's surroundings and current location. The server then generates a digital map in real time and sends the data to the autonomous vehicle's navigation system.
[0604] Data processing details
[0605] 1. Acquisition and transmission of video data: Video data acquired by the camera device is compressed and encoded in real time and transmitted to the server along with GPS data.
[0606] 2. Server-side data analysis: The server analyzes the transmitted video data in real time, detecting features and landmarks to determine the current location. This process uses video analysis algorithms such as OpenCV and TensorFlow.
[0607] 3. Digital map generation: Based on the analysis results and location data, the server generates a digital map that reflects the latest information about the surrounding environment, including traffic conditions, obstacle information, and event information.
[0608] 4. Data transmission and display: The generated digital map is compressed and sent in real time to the autonomous vehicle's navigation system, which then adjusts the vehicle's route accordingly based on the received map information.
[0609] Specific examples
[0610] For example, when a vehicle is traveling on a highway, the camera device and GPS module send video and location data to a server. The server then analyzes the video to detect road congestion and obstacles. Based on this information, the server then generates an optimal driving route in real time and sends it to the vehicle's navigation system, enabling the autonomous vehicle to operate safely and efficiently.
[0611] Prompt Sentence Examples
[0612] "Use your location and video data to generate digital map information that updates in real time, providing up-to-date navigation, including traffic and obstacle information."
[0613] This invention can significantly improve the operational safety and efficiency of autonomous vehicles.
[0614] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0615] Step 1:
[0616] The camera device of the autonomous vehicle captures the surrounding video data in real time. The input from the camera device is raw video frames, which are compressed in JPEG or H.264 format.
[0617] Step 2:
[0618] The GPS module of the autonomous vehicle acquires the current location data. The input from the GPS module is latitude and longitude information, and this data is converted into JSON format.
[0619] Step 3:
[0620] The acquired video data and location data are compressed and integrated into a single concatenated data packet at the autonomous vehicle's terminal. The output from the terminal is a concatenated data of compressed video frames and location information.
[0621] Step 4:
[0622] The terminal sends compressed and consolidated data packets to the server. The data is sent through the communication network and received by the server. The server receives the compressed data packets.
[0623] Step 5:
[0624] The server decodes the received data packets and separates the video data from the location data. The decoded input data is the raw video frame and the current location information.
[0625] Step 6:
[0626] The server uses OpenCV and TensorFlow to analyze the video data. The analysis detects landmarks and feature points, which are then used to identify the user's surroundings. The input is raw video data, and the output is analyzed video information (landmarks and feature points).
[0627] Step 7:
[0628] The server combines the received location data with the video analysis results to determine the exact current location. The output of the combination process is accurate current location information.
[0629] Step 8:
[0630] The server generates a digital map in real time based on the current location data. This process includes surrounding information, traffic information, and event information, and also uses existing map data such as OpenStreetMap. The input is precise current location information and additional data (traffic and event information), and the output is digital map data generated in real time.
[0631] Step 9:
[0632] The server compresses the generated digital map and sends it back to the autonomous vehicle's terminal. The data is transmitted over the communication network. The output of the server is compressed digital map data.
[0633] Step 10:
[0634] The terminal decodes the received compressed digital map data and transmits it to the navigation system of the autonomous vehicle. The input of the terminal is compressed digital map data, and the output is decoded map data.
[0635] Step 11:
[0636] The navigation system displays the decoded map data and adjusts the vehicle's driving route in real time, with the input of the navigation system being the decoded map data and the output being the adjusted driving route.
[0637] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0638] This invention relates to a system that not only provides digital map data in real time but also recognizes the user's emotions and customizes information based on those emotions. This system consists of a wearable terminal with a built-in camera device, a server with an emotion engine, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0639] overview
[0640] Using this system, users can wear wearable devices such as smart glasses and receive not only real-time updated digital map information in their field of vision, but also information customized to the user's emotions. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and then uses an emotion engine to provide information appropriate to the user's emotions.
[0641] Program processing
[0642] Video acquisition and transmission
[0643] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[0644] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[0645] Video analysis and location
[0646] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[0647] Server: Integrates location data with video analytics results to determine the user's exact location.
[0648] Emotion Recognition and Analysis
[0649] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[0650] Server: Based on the recognized emotional state, the digital map display content related to the user's current location and surrounding information is dynamically changed.
[0651] Digital Map Generation
[0652] Server: Based on the location data mentioned above, obtains geographic information around the current location, including surrounding buildings, roads, points of interest (POI), etc.
[0653] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0654] Emotion-based information customization
[0655] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[0656] Digital map transmission and display
[0657] Server: Compresses the latest digital map data and sends it to the wearable device, performing the necessary checks to ensure data integrity.
[0658] Terminal: Receives compressed digital map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[0659] Specific examples
[0660] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[0661] 1. When the user is strolling around a tourist spot
[0662] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[0663] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[0664] Server: Using the emotion engine, analyze the user's emotional state from their facial expressions and voice and recognize that they are relaxed.
[0665] Server: Generates map data containing information about scenic areas and photo spots where users relax.
[0666] Server: Sends the generated digital map data to the smart glasses.
[0667] Device: Overlays tourist attractions and directions onto the user's field of view.
[0668] User: Based on this, you can enjoy sightseeing efficiently.
[0669] 2. When a user is walking around a new city
[0670] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[0671] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[0672] Server: Using an emotion engine, the server analyzes the user's emotional state from their facial expressions and voice and recognizes that they are tired.
[0673] Server: Because the user is tired, it generates map data containing information about cafes and rest spots where the user can relax.
[0674] Server: Sends the generated map data to the smart glasses.
[0675] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[0676] Users: Map information is displayed in an easy-to-read format, allowing them to reach their destination efficiently and enjoy exploring the city.
[0677] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, greatly improving the convenience of travel and sightseeing.
[0678] The processing flow will be explained below.
[0679] Step 1:
[0680] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[0681] Step 2:
[0682] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[0683] Step 3:
[0684] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[0685] Step 4:
[0686] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[0687] Step 5:
[0688] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[0689] Step 6:
[0690] Server: Based on the recognized emotional state, the content displayed on the digital map, including the user's current location and surrounding information, is dynamically changed.
[0691] Step 7:
[0692] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[0693] Step 8:
[0694] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0695] Step 9:
[0696] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[0697] Step 10:
[0698] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[0699] Step 11:
[0700] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[0701] Step 12:
[0702] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[0703] Step 13:
[0704] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[0705] Step 14:
[0706] Users: Receive personalized information based on their emotions and take action accordingly. For example, if they are feeling stressed, they can find a relaxation spot to help them relax.
[0707] Example 2
[0708] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0709] Conventional navigation systems only provide geographical information and are unable to customize information based on the user's emotional state. This makes it difficult to effectively provide users with the information they need in real time. Furthermore, the analysis and integration of video and location data is insufficient, making it difficult to accurately identify the user's current location and direction.
[0710] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing video data and identifying the user's current location, a means for analyzing the user's facial expression and voice data and recognizing the user's emotional state, and a means for customizing information based on the emotional state. This makes it possible to provide information in real time according to the user's emotional state and to identify the user's location accurately.
[0711] A "camera device" is an imaging device built into a wearable device to capture the user's field of view.
[0712] A "wearable device" is a device that can be worn by a user and has a built-in camera device and GPS module.
[0713] "Video data" is data in digital form of a field of view captured by a camera device.
[0714] "Location Data" means information about a user's current location obtained using a GPS module or other location-determining technology.
[0715] "Analysis means" refers to software or hardware functionality for processing received data and extracting specific information.
[0716] "Emotional state" refers to emotional information such as joy, sadness, surprise, etc., obtained by analyzing the user's facial expression and voice data.
[0717] A "digital map" is a digital map that contains geographic information about the surrounding area that is dynamically generated based on the user's current location and facing direction.
[0718] "Nearby information" is information about buildings, roads, points of interest (POIs), etc., in the vicinity of the user's current location.
[0719] "Traffic information" refers to real-time traffic information related to the user's current location and route.
[0720] "Event information" is information about events being held in the vicinity of the user's current location.
[0721] "Compression means" refers to data compression algorithms and techniques for efficiently reducing the volume of large amounts of data.
[0722] An "analysis algorithm" is a specific calculation method or procedure used to analyze video, audio, location data, etc.
[0723] An "emotion recognition engine" is software or hardware that identifies a user's emotional state from facial expressions and voice data.
[0724] A "display" is a display device built into a wearable device that visually presents digital maps and other information to the user.
[0725] This invention is a system that uses a wearable device to analyze a user's movement information and emotional state in real time and provides a customized digital map based on the analysis. This system consists of a wearable device with a built-in camera device, a server with an emotion recognition engine, a server that receives and analyzes video data and location data, and a wearable device for displaying the generated digital map.
[0726] Wearable devices are devices with built-in cameras, such as smart glasses, that capture the user's field of vision and obtain it as video data. This data is compressed in a video compression format such as H.264 and sent to a server. Wearable devices also have built-in GPS modules that obtain the user's current location every second and send this location data to the server. Communication is via 4G / 5G networks, and data is transmitted using the TCP / IP protocol.
[0727] The server runs image processing algorithms on the received video data, for example using image processing libraries such as OpenCV to detect features and landmarks in the video. These features are then combined with the received GPS location data to determine the user's exact location and facing direction using triangulation.
[0728] The server then uses an emotion recognition engine to analyze the user's facial expressions and voice data to recognize their emotional state. The engine can be Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API. Based on this emotional state, the server customizes the information provided to the user.
[0729] Based on the user's current location data, the server obtains surrounding geographical, traffic, and event information from sources such as Google Maps API and OpenStreetMap. This information is integrated to generate a digital map that is updated in real time. Information based on the user's emotional state is displayed on the generated digital map. For example, if the user is tired, nearby cafes and relaxation spots are displayed first, while if the user is happy, tourist spots and event information are displayed first.
[0730] The server compresses the generated digital map data and efficiently transmits it to the wearable device using protocol buffers. The wearable device receives the data, expands the compressed data, and overlays it on the display. This allows the user to view the latest map information and customized information in real time in their field of vision.
[0731] As a concrete example, when a user is strolling around a tourist spot, the smart glasses capture video of the user's field of view and send it to a server. The server analyzes the video data and identifies the user's current location. It then uses an emotion recognition engine to analyze the user's emotional state and recognizes that the user is relaxed. The server generates map data including information on beautiful scenery and photo spots at the tourist spot and sends it to the wearable device. The user can then view tourist spots and route information in their field of view, allowing them to enjoy sightseeing efficiently.
[0732] An example of a prompt is as follows:
[0733] "Find a nearby cafe from your current location"
[0734] "Show me photo spots for tourist attractions"
[0735] "Tell me about traffic in this area."
[0736] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0737] Step 1:
[0738] Terminal: The camera device of the smart glasses captures the user's field of view. The captured video data is compressed into H.264 format by an encoder. This compressed video data is the input. The video data is transmitted to the server via a 4G / 5G network using the TCP / IP protocol. The output is the compressed video data that has arrived at the server.
[0739] Step 2:
[0740] Terminal: The GPS module acquires the user's current location every second. The acquired location data is recorded in NMEA format. This location data is the input and is sent to the server, also using the TCP / IP protocol. The output is the location data that arrives at the server.
[0741] Step 3:
[0742] Server: The server analyzes the video data it receives using an image processing library such as OpenCV. The input is compressed video data, which is first decoded and expanded frame by frame. An image processing algorithm (for example, ORB (Oriented FAST and Rotated BRIEF)) is applied to detect feature points and landmarks in the video. The results of this analysis are output.
[0743] Step 4:
[0744] Server: The server combines the feature points of the analyzed video data with the received location data. The input is feature point data and location data. These are combined using triangulation to determine the user's exact location and facing direction. The result of this processing is the output.
[0745] Step 5:
[0746] Server: The server uses an emotion recognition engine to analyze the user's facial expression and voice data. The input is the user's facial expression data and voice data. Software such as Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API is used to identify the emotional state (happiness, sadness, surprise, etc.). The results of this analysis are the output.
[0747] Step 6:
[0748] Server: Customizes information based on the recognized emotional state. Inputs are emotion recognition results and the user's location data. Based on the user's current location, the server obtains surrounding geographical information, traffic information, and event information from Google Maps API, OpenStreetMap, etc., and generates a digital map. The output is customized digital map data.
[0749] Step 7:
[0750] Server: Compresses the generated digital map data and transmits it efficiently to the wearable device using Protocol Buffers. The input is the customized digital map data. The output is transmitted as compressed data.
[0751] Step 8:
[0752] Terminal: The wearable terminal decompresses the compressed digital map data received from the server and overlays it on the display. The input is compressed digital map data. By displaying the decompressed data on the display, the user can see the latest information in real time in their field of vision. This visual information is the final output.
[0753] (Application example 2)
[0754] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0755] In recent years, advances in wearable devices and emotion analysis technology have garnered attention for providing information tailored to user emotions. However, systems that combine these technologies to provide optimal route guidance in real time for autonomous vehicles are still in the early stages of development. Existing systems struggle to efficiently integrate a user's emotional state and current location information to provide customized information in real time. Therefore, the present invention aims to provide real-time digital map information customized based on the user's emotional state, thereby improving passenger comfort in autonomous vehicles.
[0756] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for analyzing the user's emotional state using an emotion engine, means for customizing the digital map based on the analyzed emotional state, means for displaying the customized digital map on the display of the wearable device, and means installed in an autonomous vehicle for providing optimal route guidance based on the passenger's real-time emotion data and current location. This makes it possible to integrate the user's emotional state and current location information and provide appropriate route guidance and surrounding information within the autonomous vehicle.
[0757] A "camera device" is a device for acquiring video data and is built into a wearable device.
[0758] A "wearable terminal" is a device that can be worn by a user and has a built-in camera device and display.
[0759] "Video Data" means a digital representation of visual information captured by a camera device.
[0760] "Location data" refers to data indicating the current location of a user or autonomous vehicle, and is typically obtained by a GPS module.
[0761] A "server" is a high-performance computing device that receives and analyzes video data and location data.
[0762] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice data to recognize their emotional state.
[0763] A "digital map" is map data that includes geographical information, surrounding area information, traffic information, and other information that is generated in real time.
[0764] "Real-time" refers to the time range in which data is processed and updated immediately, with little to no delay.
[0765] "Customizing" means adjusting or changing the information and functionality provided based on information such as the user's emotional state or current location.
[0766] "Display" means a display device built into a wearable device that visually presents digital maps and other information to the user.
[0767] An "autonomous vehicle" is a vehicle that can move independently without human operation.
[0768] "Route guidance" is guide information that shows the optimal route or method to a destination.
[0769] The present invention relates to a system including a wearable terminal with a built-in camera device, a server with an emotion engine, a server for analyzing video data and location data, and a wearable terminal for displaying a generated digital map, which is installed in an autonomous vehicle and provides customized real-time route guidance based on the passenger's emotional state and current location.
[0770] System configuration
[0771] Wearable device: A device such as smart glasses that incorporates a camera device, display, and GPS module.
[0772] Server: A high-performance server, such as Amazon EC2, is used. An emotion engine (e.g., Affectiva), a video analysis algorithm (e.g., OpenCV), and a data compression library (e.g., zlib) are used.
[0773] Autonomous vehicle: A vehicle that is operated autonomously and carried by a user.
[0774] Program processing
[0775] Video acquisition and transmission
[0776] Terminal: The smart glasses' camera device captures the user's field of view and acquires it as video data. The acquired video data is compressed and sent to the server. In addition, the current location data is acquired from the GPS module and sent to the server.
[0777] Video analysis and location
[0778] Server: Receives video data and analyzes it using a video analytics algorithm (e.g., OpenCV) to detect features and landmarks within the field of view. This data is then combined with GPS location data to determine the current location of the autonomous vehicle.
[0779] Emotion Recognition and Analysis
[0780] Server: An emotion engine (e.g., Affectiva) is used to analyze the user's facial expressions and voice data to recognize their emotional state, which can include happiness, sadness, surprise, etc.
[0781] Digital map generation and customization
[0782] Server: Based on the current location data, the server obtains surrounding geographic information, traffic information, and specific points (restaurants, tourist attractions, rest spots, etc.) in real time and generates a digital map. Furthermore, the server customizes the information displayed based on the user's emotional state, as determined by emotion recognition. For example, if the user is relaxed, it will prioritize displaying tourist attractions and relaxation spots, while if the user is feeling stressed, it will prioritize providing information on the shortest route and rest spots.
[0783] Digital map transmission and display
[0784] Server: Sends compressed digital map data to the device.
[0785] Device: Digital map data is deployed and overlaid on the smart glasses display, allowing users to visually check the latest information in real time.
[0786] Specific usage scenarios
[0787] 1. Tourism scenario:
[0788] When a user walks around a tourist spot, the camera device in the smart glasses captures video data and sends it to a server.
[0789] The server analyzes the video data to determine the current location and transmits map data including information on tourist attractions and directions.
[0790] If the emotion engine determines that the user is relaxed, it will provide customized map data including information on scenic areas and photo spots.
[0791] 2. Urban Walking Scenario:
[0792] As a user walks around a new city, the camera device in the smart glasses captures video data and sends it to a server.
[0793] The server analyzes the video data, determines the current location, and then provides map data including information on nearby popular spots and stores.
[0794] If emotion recognition determines that the user is tired, it will prioritize providing information about relaxation spots and cafes.
[0795] Prompt Sentence Examples
[0796] Here are some examples of prompts for a generative AI model:
[0797] If the user is relaxed:
[0798] "Current emotional state: Relaxed. Prioritize information about nearby tourist attractions and relaxation spots."
[0799] If the user is stressed:
[0800] "Current emotional state: Stress. Prioritize showing quicker routes and directions to relaxation spots."
[0801] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, making travel and sightseeing more comfortable and enjoyable.
[0802] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0803] Step 1:
[0804] Video acquisition and transmission
[0805] The device (smart glasses) uses a camera device to acquire video data within the user's field of view, including the scenery and objects the user is looking at.
[0806] Input: User's field of view
[0807] Output: Video data
[0808] Specific operation: The video data is compressed and sent to the server in real time. At the same time, the current location data is obtained from the GPS module and sent to the server.
[0809] Step 2:
[0810] Video analysis and location
[0811] The server analyzes the received video data using a video analysis algorithm (e.g., OpenCV), which detects feature points and landmarks within the field of view.
[0812] Input: Video data, location data
[0813] Output: Feature point data, landmark data, current location data
[0814] What it does: Detects features and landmarks in the video and combines this data with GPS location data to determine the user's current location.
[0815] Step 3:
[0816] Emotion Recognition and Analysis
[0817] The server uses an emotion engine (e.g., Affectiva) to analyze the user's facial expressions and voice data to recognize their emotional state.
[0818] Input: User's facial expression data, voice data
[0819] Output: Emotional state data
[0820] Specific behavior: Identify emotional states such as happiness, sadness, and surprise from facial expressions and voice, and store that information in a database.
[0821] Step 4:
[0822] Digital map generation
[0823] The server obtains real-time geographical information, traffic information, and data on specific points (restaurants, tourist attractions, rest spots, etc.) based on the current location data, and generates a digital map.
[0824] Input: Current location data, surrounding geographic information, traffic information, point information
[0825] Output: Digital map data
[0826] Specific operation: Integrates information on surrounding buildings, roads, POIs (points of interest), etc. to generate an up-to-date digital map.
[0827] Step 5:
[0828] Customizing digital maps
[0829] The server customizes the information displayed based on the user's emotional state. For example, if the user is relaxed, it will prioritize showing tourist attractions and relaxation spots, while if the user is stressed, it will prioritize showing information about the shortest routes and rest spots.
[0830] Input: Digital map data, emotional state data
[0831] Output: Customized digital map data
[0832] Specific operation: Dynamically change the displayed content of the digital map according to the user's emotional state to provide the most appropriate information to the user.
[0833] Step 6:
[0834] Digital map transmission and display
[0835] The server compresses the customized digital map data and sends it to the smart glasses.
[0836] Input: Customized digital map data
[0837] Output: Compressed digital map data packets
[0838] Specific operation: Data to be overlaid on the display is compressed and sent to the terminal in real time. The terminal then decompresses the received data and displays it on the display.
[0839] Step 7:
[0840] User receipt and use
[0841] Users can check customized digital map information displayed on the smart glasses display and use it when traveling or sightseeing.
[0842] Input: Customized digital map information displayed on the display
[0843] Output: User actions and choices
[0844] Specific actions: Decide on the next action based on visually presented information, such as heading to a displayed relaxation spot.
[0845] This allows users to receive optimal route guidance and information based on real-time updates and their emotional state, making self-driving vehicles more comfortable to use.
[0846] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0847] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0848] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0849] [Third embodiment]
[0850] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0851] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0852] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0853] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0854] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0855] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0856] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0857] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0858] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0859] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0860] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0861] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0862] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0863] overview
[0864] This system allows users to wear wearable devices such as smart glasses and receive real-time updated digital map information in their field of vision. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and sends it to the wearable device.
[0865] Program processing
[0866] Video acquisition and transmission
[0867] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[0868] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[0869] Video analysis and location
[0870] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[0871] Server: Integrates location data with video analytics results to determine the user's exact location.
[0872] Digital Map Generation
[0873] Server: Based on the location data, a real-time digital map is generated, containing up-to-date information on nearby facilities, traffic, and events.
[0874] Server: Optionally overlay additional data layers (e.g., top-rated restaurants or tourist attractions) on the map.
[0875] Digital map transmission and display
[0876] Server: Compresses the latest digital map data and sends it to the wearable device.
[0877] Terminal: Receives compressed map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[0878] Specific examples
[0879] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[0880] 1. When the user is strolling around a tourist spot
[0881] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[0882] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[0883] Server: Sends map data containing the latest tourist spot information to the smart glasses.
[0884] Device: Overlays tourist attractions and directions onto the user's field of view.
[0885] User: Based on this, you can enjoy sightseeing efficiently.
[0886] 2. When a user is walking around a new city
[0887] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[0888] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[0889] Server: Sends the generated map data to the smart glasses.
[0890] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[0891] Users: Using map information displayed in an easy-to-read format, they can reach their destination efficiently and enjoy exploring the city.
[0892] This will allow users to visually receive information that is updated in real time, greatly improving the convenience of travel and sightseeing.
[0893] The processing flow will be explained below.
[0894] Step 1:
[0895] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[0896] Step 2:
[0897] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[0898] Step 3:
[0899] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[0900] Step 4:
[0901] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[0902] Step 5:
[0903] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[0904] Step 6:
[0905] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[0906] Step 7:
[0907] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[0908] Step 8:
[0909] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[0910] Step 9:
[0911] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[0912] Step 10:
[0913] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[0914] Example 1
[0915] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0916] Conventional digital map systems have had difficulty updating information in real time and providing users with accurate location information. As a result, users are unable to efficiently obtain the latest information on their surroundings, traffic, tourist spots, and other information, resulting in low convenience. The present invention aims to solve these problems and provide users with the latest digital map information that is updated in real time.
[0917] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0918] In this invention, the server includes means for analyzing video data, detecting feature points and landmarks, and identifying the user's current location and orientation, means for generating a digital map in real time based on the current location, and means for compressing the generated digital map data and transmitting it to the wearable device, thereby enabling the user to visually receive a digital map that is updated in real time and includes the latest surrounding information, traffic information, and event information.
[0919] "Camera device" refers to a device for capturing video data that is built into a wearable device.
[0920] A "wearable device" is a device worn by the user that has a built-in camera device and GPS module and has the ability to display information within the field of view.
[0921] "Video Data" means data that captures a User's field of view with a camera device and that can be stored or transmitted in digital form.
[0922] "Location data" refers to data indicating the user's current location, typically information obtained from a GPS module.
[0923] A "GPS module" is a hardware component for obtaining geographical location information.
[0924] "Server" refers to a remote computer system that receives data sent from a wearable device and performs analysis and data generation.
[0925] "Feature points" are important points in an image detected by video analysis algorithms and used to identify objects or landmarks.
[0926] A "landmark" refers to a distinctive feature such as a building or terrain that serves as a landmark to identify the user's location and orientation.
[0927] "Real-time" means that data is acquired, processed, and displayed instantly, with minimal delay.
[0928] A "digital map" is a map containing electronically generated geographic information that is used to display a user's current location and surrounding area.
[0929] "Nearby information" refers to information about facilities and features located near the user's current location.
[0930] "Traffic information" refers to information related to travel, such as road congestion conditions and public transportation operation information.
[0931] "Event Information" refers to information about events or activities taking place at a particular location.
[0932] "Compression" is a process for reducing the volume of data and is used to improve transmission efficiency.
[0933] "Overlay" refers to the technology of superimposing information onto the field of view.
[0934] MODE FOR CARRYING OUT THE INVENTION
[0935] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[0936] System configuration
[0937] This system consists of the following elements:
[0938] Wearable devices
[0939] 1. Camera Device: A device that captures the user's field of view in real time. The camera device captures high-resolution video and stores or transmits the video data in digital format.
[0940] 2. GPS module: A hardware component for obtaining geographical location information, accurately determining the user's current location.
[0941] 3. Display: A device that displays digital maps and other information, with the ability to overlay it on the user's field of view.
[0942] 4. Communication module: This module transmits video data and location data to the server.
[0943] server
[0944] 1. Data Reception Module: This module receives video and location data sent from the wearable device. The received data is compressed and decoded into H.264 or NMEA format.
[0945] 2. Video analysis algorithm: Using the OpenCV library, feature points and landmarks are detected from the received video data, which allows the user's current position and orientation to be determined.
[0946] 3. Position and orientation determination module: Integrates video analysis results and position data, and uses a Kalman filter to estimate the user's exact position and orientation.
[0947] 4. Digital Map Generation Engine: Generates a digital map in real time based on the user's current location, querying and retrieving information about surroundings, traffic, and events from a database (e.g., PostGIS).
[0948] 5. Data compression module: Compresses the generated digital map data in JSON format and sends it to the wearable device.
[0949] 6. Communication module: A device that transmits compressed digital map data to the wearable device.
[0950] Specific examples of implementation
[0951] When the user is exploring a tourist spot
[0952] As the user walks around a tourist spot, the camera in the smart glasses captures the user's field of view, compresses it, and sends it to the server. At the same time, the GPS module acquires the user's current location and sends it to the server. The server analyzes the received video data, detecting feature points and landmarks to determine the user's current location and orientation. A real-time digital map containing up-to-date tourist spot and route information is then sent to the wearable device. The wearable device then overlays the received information on the user's field of view, allowing the user to enjoy sightseeing efficiently.
[0953] Prompt Sentence Examples
[0954] "Explain how you can provide real-time updated digital map information for a scenario where a user is exploring a tourist destination."
[0955] This invention allows users to visually receive digital maps that are updated in real time and include the latest local, traffic, and event information, greatly improving convenience.
[0956] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0957] Step 1: Capture video data
[0958] Terminal: A camera device captures the user's field of view in real time. The input is an image based on the user's field of view, and the output is video data converted into a digital format. This video data has a resolution of 640x480 pixels and is acquired at 30 frames per second (fps).
[0959] Step 2: Compressing the video data
[0960] Terminal: Compresses the acquired video data in H.264 format. The input is the captured raw video data, and the output is compressed video data in H.264 format. Compression reduces the data size and improves transmission efficiency.
[0961] Step 3: Getting location data
[0962] Device: The built-in GPS module obtains the user's current location. The input is signals from GPS satellites, and the output is location data encoded in NMEA format. This data is obtained by obtaining latitude and longitude information, updated every second.
[0963] Step 4: Sending data
[0964] Terminal: Sends compressed video data and acquired location data to the server. The input is compressed video data and location data, and the output is the data sent to the server. Video data is sent using the UDP protocol, and location data is sent using the TCP protocol to ensure communication reliability and efficiency.
[0965] Step 5: Receive and decode the data
[0966] Server: Decodes the received compressed video data using an H.264 decoder and analyzes the location data. The input is the compressed video data and location data, and the output is the decoded video data and analyzed location information. The decoding and analysis process makes it possible to visualize the video and location.
[0967] Step 6: Video analysis and feature point detection
[0968] Server: Extracts feature points (e.g., ORB feature points) from input video data using the OpenCV library. The input is the decoded video data, and the output is the detected feature points and landmark information. Calculates the user's orientation based on the feature points and identifies landmarks.
[0969] Step 7: Identifying position and orientation
[0970] Server: Integrates feature point and location data and uses a Kalman filter to estimate the user's precise location and orientation. The input is feature point information and latitude and longitude data, and the output is estimated location and orientation information. Identifying the location and orientation clarifies the user's spatial state.
[0971] Step 8: Generate a digital map
[0972] Server: Based on the user's current location, obtains surrounding information, traffic information, and event information from a database and generates a real-time digital map. The input is the estimated location and orientation information and a database query, and the output is the generated digital map data. The map generation engine integrates multi-layer data.
[0973] Step 9: Compressing the digital map
[0974] Server: Compresses the generated digital map data in JSON format. The input is the generated real-time map data, and the output is compressed map data in JSON format. Compression improves the efficiency of data transmission.
[0975] Step 10: Sending the digital map
[0976] Server: Sends compressed digital map data to the wearable device. The input is the compressed data, and the output is the data sent to the wearable device. The transmission uses the TCP protocol to ensure data integrity.
[0977] Step 11: Developing the digital map
[0978] Terminal: Decompresses the received compressed data and renders it for display. The input is compressed map data, and the output is a real-time digital map displayed in the user's field of view.
[0979] Step 12: Digital map overlay display
[0980] Terminal: The expanded map data is overlaid on the user's field of view. The input is the expanded map data, and the output is the map and related information displayed in the user's field of view. This allows the user to visually check the latest map information.
[0981] (Application example 1)
[0982] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0983] Conventional autonomous vehicle systems have struggled to efficiently grasp real-time changes in road conditions and traffic information to improve the safety and efficiency of autonomous driving. Furthermore, the effectiveness of navigation systems is limited because the information obtained while the user is actually driving is only retained for a short period of time. Therefore, there is a need for a system that can reflect the latest information about the vehicle's surroundings in real time, improving the navigation accuracy and safety of autonomous vehicles.
[0984] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0985] In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for displaying the digital map on the display of the wearable device, means for acquiring the vehicle's surroundings from the camera device of the autonomous vehicle and dynamically analyzing the surrounding environment based on location information, and means for reflecting the analysis results in the autonomous vehicle's navigation system in real time. This allows the autonomous vehicle to grasp the latest road conditions and traffic information in real time, enabling safe and efficient operation in response to changes in the environment around the vehicle.
[0986] A "camera device" is an optical device for capturing video data.
[0987] A "wearable device" is an electronic device that can be worn by a user and that acquires and displays data.
[0988] "Video data" is visual information captured using a camera device.
[0989] "Location data" is information indicating the current location obtained using a positioning system such as GPS.
[0990] The "receiving means" refers to a device or mechanism for receiving data from the outside.
[0991] "Means for analyzing" are devices and mechanisms for interpreting and converting acquired data into meaningful information.
[0992] "Current location" means the location where the user or autonomous vehicle is currently located.
[0993] A "digital map" is electronic map data that visually represents location information and related information.
[0994] "Generating means" are devices and mechanisms for producing new data or information.
[0995] A "transmitting means" is a device or mechanism for transmitting data to another device or system.
[0996] A "display" is a device for visually displaying information.
[0997] An "autonomous vehicle" is a vehicle that operates autonomously using artificial intelligence and sensor technology.
[0998] "Surrounding conditions" refers to information about the environment in which an autonomous vehicle operates.
[0999] A "navigation system" is a device or mechanism that guides a vehicle to a destination or provides route guidance based on the vehicle's current location.
[1000] "Dynamic analysis means" refers to devices and mechanisms that analyze data in real time and take immediate action or make decisions based on the results.
[1001] System Configuration
[1002] In this invention, the invention is implemented using a system consisting of the following components:
[1003] 1. Hardware
[1004] Camera device: High-resolution cameras installed in autonomous vehicles are used to acquire video data of the surrounding area.
[1005] GPS module: Uses the GPS system installed in the vehicle to determine the current location.
[1006] Wearable device: A device worn by the user, such as smart glasses, that has a built-in camera and display.
[1007] Server: A high-performance server (e.g. AWS EC2) that receives data, analyzes it, and generates digital maps.
[1008] Communication network: Internet connection for sending and receiving data in real time (e.g., 5G communication).
[1009] 2. Software
[1010] Video analysis algorithm: OpenCV and TensorFlow are used to analyze video data and detect feature points and landmarks.
[1011] Digital map generation software: A program that generates digital maps in real time based on the data received.
[1012] Communication protocol: Protocols such as HTTP and WebSocket for sending and receiving data.
[1013] Overview of operation procedures
[1014] When a user operates an autonomous vehicle, the system captures video and location data from the camera device and GPS module. The data is then compressed and sent to a server. The server analyzes the received data to determine the user's surroundings and current location. The server then generates a digital map in real time and sends the data to the autonomous vehicle's navigation system.
[1015] Data processing details
[1016] 1. Acquisition and transmission of video data: Video data acquired by the camera device is compressed and encoded in real time and transmitted to the server along with GPS data.
[1017] 2. Server-side data analysis: The server analyzes the transmitted video data in real time, detecting features and landmarks to determine the current location. This process uses video analysis algorithms such as OpenCV and TensorFlow.
[1018] 3. Digital map generation: Based on the analysis results and location data, the server generates a digital map that reflects the latest information about the surrounding environment, including traffic conditions, obstacle information, and event information.
[1019] 4. Data transmission and display: The generated digital map is compressed and sent in real time to the autonomous vehicle's navigation system, which then adjusts the vehicle's route accordingly based on the received map information.
[1020] Specific examples
[1021] For example, when a vehicle is traveling on a highway, the camera device and GPS module send video and location data to a server. The server then analyzes the video to detect road congestion and obstacles. Based on this information, the server then generates an optimal driving route in real time and sends it to the vehicle's navigation system, enabling the autonomous vehicle to operate safely and efficiently.
[1022] Prompt Sentence Examples
[1023] "Use your location and video data to generate digital map information that updates in real time, providing up-to-date navigation, including traffic and obstacle information."
[1024] This invention can significantly improve the operational safety and efficiency of autonomous vehicles.
[1025] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1026] Step 1:
[1027] The camera device of the autonomous vehicle captures the surrounding video data in real time. The input from the camera device is raw video frames, which are compressed in JPEG or H.264 format.
[1028] Step 2:
[1029] The GPS module of the autonomous vehicle acquires the current location data. The input from the GPS module is latitude and longitude information, and this data is converted into JSON format.
[1030] Step 3:
[1031] The acquired video data and location data are compressed and integrated into a single concatenated data packet at the autonomous vehicle's terminal. The output from the terminal is a concatenated data of compressed video frames and location information.
[1032] Step 4:
[1033] The terminal sends compressed and consolidated data packets to the server. The data is sent through the communication network and received by the server. The server receives the compressed data packets.
[1034] Step 5:
[1035] The server decodes the received data packets and separates the video data from the location data. The decoded input data is the raw video frame and the current location information.
[1036] Step 6:
[1037] The server uses OpenCV and TensorFlow to analyze the video data. The analysis detects landmarks and feature points, which are then used to identify the user's surroundings. The input is raw video data, and the output is analyzed video information (landmarks and feature points).
[1038] Step 7:
[1039] The server combines the received location data with the video analysis results to determine the exact current location. The output of the combination process is accurate current location information.
[1040] Step 8:
[1041] The server generates a digital map in real time based on the current location data. This process includes surrounding information, traffic information, and event information, and also uses existing map data such as OpenStreetMap. The input is precise current location information and additional data (traffic and event information), and the output is digital map data generated in real time.
[1042] Step 9:
[1043] The server compresses the generated digital map and sends it back to the autonomous vehicle's terminal. The data is transmitted over the communication network. The output of the server is compressed digital map data.
[1044] Step 10:
[1045] The terminal decodes the received compressed digital map data and transmits it to the navigation system of the autonomous vehicle. The input of the terminal is compressed digital map data, and the output is decoded map data.
[1046] Step 11:
[1047] The navigation system displays the decoded map data and adjusts the vehicle's driving route in real time, with the input of the navigation system being the decoded map data and the output being the adjusted driving route.
[1048] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1049] This invention relates to a system that not only provides digital map data in real time but also recognizes the user's emotions and customizes information based on those emotions. This system consists of a wearable terminal with a built-in camera device, a server with an emotion engine, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[1050] overview
[1051] Using this system, users can wear wearable devices such as smart glasses and receive not only real-time updated digital map information in their field of vision, but also information customized to the user's emotions. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and then uses an emotion engine to provide information appropriate to the user's emotions.
[1052] Program processing
[1053] Video acquisition and transmission
[1054] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[1055] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[1056] Video analysis and location
[1057] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[1058] Server: Integrates location data with video analytics results to determine the user's exact location.
[1059] Emotion Recognition and Analysis
[1060] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[1061] Server: Based on the recognized emotional state, the digital map display content related to the user's current location and surrounding information is dynamically changed.
[1062] Digital Map Generation
[1063] Server: Based on the location data mentioned above, obtains geographic information around the current location, including surrounding buildings, roads, points of interest (POI), etc.
[1064] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[1065] Emotion-based information customization
[1066] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[1067] Digital map transmission and display
[1068] Server: Compresses the latest digital map data and sends it to the wearable device, performing the necessary checks to ensure data integrity.
[1069] Terminal: Receives compressed digital map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[1070] Specific examples
[1071] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[1072] 1. When the user is strolling around a tourist spot
[1073] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[1074] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[1075] Server: Using the emotion engine, analyze the user's emotional state from their facial expressions and voice and recognize that they are relaxed.
[1076] Server: Generates map data containing information about scenic areas and photo spots where users relax.
[1077] Server: Sends the generated digital map data to the smart glasses.
[1078] Device: Overlays tourist attractions and directions onto the user's field of view.
[1079] User: Based on this, you can enjoy sightseeing efficiently.
[1080] 2. When a user is walking around a new city
[1081] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[1082] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[1083] Server: Using an emotion engine, the server analyzes the user's emotional state from their facial expressions and voice and recognizes that they are tired.
[1084] Server: Because the user is tired, it generates map data containing information about cafes and rest spots where the user can relax.
[1085] Server: Sends the generated map data to the smart glasses.
[1086] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[1087] Users: Map information is displayed in an easy-to-read format, allowing them to reach their destination efficiently and enjoy exploring the city.
[1088] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, greatly improving the convenience of travel and sightseeing.
[1089] The processing flow will be explained below.
[1090] Step 1:
[1091] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[1092] Step 2:
[1093] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[1094] Step 3:
[1095] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[1096] Step 4:
[1097] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[1098] Step 5:
[1099] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[1100] Step 6:
[1101] Server: Based on the recognized emotional state, the content displayed on the digital map, including the user's current location and surrounding information, is dynamically changed.
[1102] Step 7:
[1103] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[1104] Step 8:
[1105] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[1106] Step 9:
[1107] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[1108] Step 10:
[1109] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[1110] Step 11:
[1111] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[1112] Step 12:
[1113] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[1114] Step 13:
[1115] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[1116] Step 14:
[1117] Users: Receive personalized information based on their emotions and take action accordingly. For example, if they are feeling stressed, they can find a relaxation spot to help them relax.
[1118] Example 2
[1119] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1120] Conventional navigation systems only provide geographical information and are unable to customize information based on the user's emotional state. This makes it difficult to effectively provide users with the information they need in real time. Furthermore, the analysis and integration of video and location data is insufficient, making it difficult to accurately identify the user's current location and direction.
[1121] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing video data and identifying the user's current location, a means for analyzing the user's facial expression and voice data and recognizing the user's emotional state, and a means for customizing information based on the emotional state. This makes it possible to provide information in real time according to the user's emotional state and to identify the user's location accurately.
[1122] A "camera device" is an imaging device built into a wearable device to capture the user's field of view.
[1123] A "wearable device" is a device that can be worn by a user and has a built-in camera device and GPS module.
[1124] "Video data" is data in digital form of a field of view captured by a camera device.
[1125] "Location Data" means information about a user's current location obtained using a GPS module or other location-determining technology.
[1126] "Analysis means" refers to software or hardware functionality for processing received data and extracting specific information.
[1127] "Emotional state" refers to emotional information such as joy, sadness, surprise, etc., obtained by analyzing the user's facial expression and voice data.
[1128] A "digital map" is a digital map that contains geographic information about the surrounding area that is dynamically generated based on the user's current location and facing direction.
[1129] "Nearby information" is information about buildings, roads, points of interest (POIs), etc., in the vicinity of the user's current location.
[1130] "Traffic information" refers to real-time traffic information related to the user's current location and route.
[1131] "Event information" is information about events being held in the vicinity of the user's current location.
[1132] "Compression means" refers to data compression algorithms and techniques for efficiently reducing the volume of large amounts of data.
[1133] An "analysis algorithm" is a specific calculation method or procedure used to analyze video, audio, location data, etc.
[1134] An "emotion recognition engine" is software or hardware that identifies a user's emotional state from facial expressions and voice data.
[1135] A "display" is a display device built into a wearable device that visually presents digital maps and other information to the user.
[1136] This invention is a system that uses a wearable device to analyze a user's movement information and emotional state in real time and provides a customized digital map based on the analysis. This system consists of a wearable device with a built-in camera device, a server with an emotion recognition engine, a server that receives and analyzes video data and location data, and a wearable device for displaying the generated digital map.
[1137] Wearable devices are devices with built-in cameras, such as smart glasses, that capture the user's field of vision and obtain it as video data. This data is compressed in a video compression format such as H.264 and sent to a server. Wearable devices also have built-in GPS modules that obtain the user's current location every second and send this location data to the server. Communication is via 4G / 5G networks, and data is transmitted using the TCP / IP protocol.
[1138] The server runs image processing algorithms on the received video data, for example using image processing libraries such as OpenCV to detect features and landmarks in the video. These features are then combined with the received GPS location data to determine the user's exact location and facing direction using triangulation.
[1139] The server then uses an emotion recognition engine to analyze the user's facial expressions and voice data to recognize their emotional state. The engine can be Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API. Based on this emotional state, the server customizes the information provided to the user.
[1140] Based on the user's current location data, the server obtains surrounding geographical, traffic, and event information from sources such as Google Maps API and OpenStreetMap. This information is integrated to generate a digital map that is updated in real time. Information based on the user's emotional state is displayed on the generated digital map. For example, if the user is tired, nearby cafes and relaxation spots are displayed first, while if the user is happy, tourist spots and event information are displayed first.
[1141] The server compresses the generated digital map data and efficiently transmits it to the wearable device using protocol buffers. The wearable device receives the data, expands the compressed data, and overlays it on the display. This allows the user to view the latest map information and customized information in real time in their field of vision.
[1142] As a concrete example, when a user is strolling around a tourist spot, the smart glasses capture video of the user's field of view and send it to a server. The server analyzes the video data and identifies the user's current location. It then uses an emotion recognition engine to analyze the user's emotional state and recognizes that the user is relaxed. The server generates map data including information on beautiful scenery and photo spots at the tourist spot and sends it to the wearable device. The user can then view tourist spots and route information in their field of view, allowing them to enjoy sightseeing efficiently.
[1143] An example of a prompt is as follows:
[1144] "Find a nearby cafe from your current location"
[1145] "Show me photo spots for tourist attractions"
[1146] "Tell me about traffic in this area."
[1147] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1148] Step 1:
[1149] Terminal: The camera device of the smart glasses captures the user's field of view. The captured video data is compressed into H.264 format by an encoder. This compressed video data is the input. The video data is transmitted to the server via a 4G / 5G network using the TCP / IP protocol. The output is the compressed video data that has arrived at the server.
[1150] Step 2:
[1151] Terminal: The GPS module acquires the user's current location every second. The acquired location data is recorded in NMEA format. This location data is the input and is sent to the server, also using the TCP / IP protocol. The output is the location data that arrives at the server.
[1152] Step 3:
[1153] Server: The server analyzes the video data it receives using an image processing library such as OpenCV. The input is compressed video data, which is first decoded and expanded frame by frame. An image processing algorithm (for example, ORB (Oriented FAST and Rotated BRIEF)) is applied to detect feature points and landmarks in the video. The results of this analysis are output.
[1154] Step 4:
[1155] Server: The server combines the feature points of the analyzed video data with the received location data. The input is feature point data and location data. These are combined using triangulation to determine the user's exact location and facing direction. The result of this processing is the output.
[1156] Step 5:
[1157] Server: The server uses an emotion recognition engine to analyze the user's facial expression and voice data. The input is the user's facial expression data and voice data. Software such as Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API is used to identify the emotional state (happiness, sadness, surprise, etc.). The results of this analysis are the output.
[1158] Step 6:
[1159] Server: Customizes information based on the recognized emotional state. Inputs are emotion recognition results and the user's location data. Based on the user's current location, the server obtains surrounding geographical information, traffic information, and event information from Google Maps API, OpenStreetMap, etc., and generates a digital map. The output is customized digital map data.
[1160] Step 7:
[1161] Server: Compresses the generated digital map data and transmits it efficiently to the wearable device using Protocol Buffers. The input is the customized digital map data. The output is transmitted as compressed data.
[1162] Step 8:
[1163] Terminal: The wearable terminal decompresses the compressed digital map data received from the server and overlays it on the display. The input is compressed digital map data. By displaying the decompressed data on the display, the user can see the latest information in real time in their field of vision. This visual information is the final output.
[1164] (Application example 2)
[1165] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1166] In recent years, advances in wearable devices and emotion analysis technology have garnered attention for providing information tailored to user emotions. However, systems that combine these technologies to provide optimal route guidance in real time for autonomous vehicles are still in the early stages of development. Existing systems struggle to efficiently integrate a user's emotional state and current location information to provide customized information in real time. Therefore, the present invention aims to provide real-time digital map information customized based on the user's emotional state, thereby improving passenger comfort in autonomous vehicles.
[1167] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for analyzing the user's emotional state using an emotion engine, means for customizing the digital map based on the analyzed emotional state, means for displaying the customized digital map on the display of the wearable device, and means installed in an autonomous vehicle for providing optimal route guidance based on the passenger's real-time emotion data and current location. This makes it possible to integrate the user's emotional state and current location information and provide appropriate route guidance and surrounding information within the autonomous vehicle.
[1168] A "camera device" is a device for acquiring video data and is built into a wearable device.
[1169] A "wearable terminal" is a device that can be worn by a user and has a built-in camera device and display.
[1170] "Video Data" means a digital representation of visual information captured by a camera device.
[1171] "Location data" refers to data indicating the current location of a user or autonomous vehicle, and is typically obtained by a GPS module.
[1172] A "server" is a high-performance computing device that receives and analyzes video data and location data.
[1173] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice data to recognize their emotional state.
[1174] A "digital map" is map data that includes geographical information, surrounding area information, traffic information, and other information that is generated in real time.
[1175] "Real-time" refers to the time range in which data is processed and updated immediately, with little to no delay.
[1176] "Customizing" means adjusting or changing the information and functionality provided based on information such as the user's emotional state or current location.
[1177] "Display" means a display device built into a wearable device that visually presents digital maps and other information to the user.
[1178] An "autonomous vehicle" is a vehicle that can move independently without human operation.
[1179] "Route guidance" is guide information that shows the optimal route or method to a destination.
[1180] The present invention relates to a system including a wearable terminal with a built-in camera device, a server with an emotion engine, a server for analyzing video data and location data, and a wearable terminal for displaying a generated digital map, which is installed in an autonomous vehicle and provides customized real-time route guidance based on the passenger's emotional state and current location.
[1181] System configuration
[1182] Wearable device: A device such as smart glasses that incorporates a camera device, display, and GPS module.
[1183] Server: A high-performance server, such as Amazon EC2, is used. An emotion engine (e.g., Affectiva), a video analysis algorithm (e.g., OpenCV), and a data compression library (e.g., zlib) are used.
[1184] Autonomous vehicle: A vehicle that is operated autonomously and carried by a user.
[1185] Program processing
[1186] Video acquisition and transmission
[1187] Terminal: The smart glasses' camera device captures the user's field of view and acquires it as video data. The acquired video data is compressed and sent to the server. In addition, the current location data is acquired from the GPS module and sent to the server.
[1188] Video analysis and location
[1189] Server: Receives video data and analyzes it using a video analytics algorithm (e.g., OpenCV) to detect features and landmarks within the field of view. This data is then combined with GPS location data to determine the current location of the autonomous vehicle.
[1190] Emotion Recognition and Analysis
[1191] Server: An emotion engine (e.g., Affectiva) is used to analyze the user's facial expressions and voice data to recognize their emotional state, which can include happiness, sadness, surprise, etc.
[1192] Digital map generation and customization
[1193] Server: Based on the current location data, the server obtains surrounding geographic information, traffic information, and specific points (restaurants, tourist attractions, rest spots, etc.) in real time and generates a digital map. Furthermore, the server customizes the information displayed based on the user's emotional state, as determined by emotion recognition. For example, if the user is relaxed, it will prioritize displaying tourist attractions and relaxation spots, while if the user is feeling stressed, it will prioritize providing information on the shortest route and rest spots.
[1194] Digital map transmission and display
[1195] Server: Sends compressed digital map data to the device.
[1196] Device: Digital map data is deployed and overlaid on the smart glasses display, allowing users to visually check the latest information in real time.
[1197] Specific usage scenarios
[1198] 1. Tourism scenario:
[1199] When a user walks around a tourist spot, the camera device in the smart glasses captures video data and sends it to a server.
[1200] The server analyzes the video data to determine the current location and transmits map data including information on tourist attractions and directions.
[1201] If the emotion engine determines that the user is relaxed, it will provide customized map data including information on scenic areas and photo spots.
[1202] 2. Urban Walking Scenario:
[1203] As a user walks around a new city, the camera device in the smart glasses captures video data and sends it to a server.
[1204] The server analyzes the video data, determines the current location, and then provides map data including information on nearby popular spots and stores.
[1205] If emotion recognition determines that the user is tired, it will prioritize providing information about relaxation spots and cafes.
[1206] Prompt Sentence Examples
[1207] Here are some examples of prompts for a generative AI model:
[1208] If the user is relaxed:
[1209] "Current emotional state: Relaxed. Prioritize information about nearby tourist attractions and relaxation spots."
[1210] If the user is stressed:
[1211] "Current emotional state: Stress. Prioritize showing quicker routes and directions to relaxation spots."
[1212] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, making travel and sightseeing more comfortable and enjoyable.
[1213] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1214] Step 1:
[1215] Video acquisition and transmission
[1216] The device (smart glasses) uses a camera device to acquire video data within the user's field of view, including the scenery and objects the user is looking at.
[1217] Input: User's field of view
[1218] Output: Video data
[1219] Specific operation: The video data is compressed and sent to the server in real time. At the same time, the current location data is obtained from the GPS module and sent to the server.
[1220] Step 2:
[1221] Video analysis and location
[1222] The server analyzes the received video data using a video analysis algorithm (e.g., OpenCV), which detects feature points and landmarks within the field of view.
[1223] Input: Video data, location data
[1224] Output: Feature point data, landmark data, current location data
[1225] What it does: Detects features and landmarks in the video and combines this data with GPS location data to determine the user's current location.
[1226] Step 3:
[1227] Emotion Recognition and Analysis
[1228] The server uses an emotion engine (e.g., Affectiva) to analyze the user's facial expressions and voice data to recognize their emotional state.
[1229] Input: User's facial expression data, voice data
[1230] Output: Emotional state data
[1231] Specific behavior: Identify emotional states such as happiness, sadness, and surprise from facial expressions and voice, and store that information in a database.
[1232] Step 4:
[1233] Digital map generation
[1234] The server obtains real-time geographical information, traffic information, and data on specific points (restaurants, tourist attractions, rest spots, etc.) based on the current location data, and generates a digital map.
[1235] Input: Current location data, surrounding geographic information, traffic information, point information
[1236] Output: Digital map data
[1237] Specific operation: Integrates information on surrounding buildings, roads, POIs (points of interest), etc. to generate an up-to-date digital map.
[1238] Step 5:
[1239] Customizing digital maps
[1240] The server customizes the information displayed based on the user's emotional state. For example, if the user is relaxed, it will prioritize showing tourist attractions and relaxation spots, while if the user is stressed, it will prioritize showing information about the shortest routes and rest spots.
[1241] Input: Digital map data, emotional state data
[1242] Output: Customized digital map data
[1243] Specific operation: Dynamically change the displayed content of the digital map according to the user's emotional state to provide the most appropriate information to the user.
[1244] Step 6:
[1245] Digital map transmission and display
[1246] The server compresses the customized digital map data and sends it to the smart glasses.
[1247] Input: Customized digital map data
[1248] Output: Compressed digital map data packets
[1249] Specific operation: Data to be overlaid on the display is compressed and sent to the terminal in real time. The terminal then decompresses the received data and displays it on the display.
[1250] Step 7:
[1251] User receipt and use
[1252] Users can check customized digital map information displayed on the smart glasses display and use it when traveling or sightseeing.
[1253] Input: Customized digital map information displayed on the display
[1254] Output: User actions and choices
[1255] Specific actions: Decide on the next action based on visually presented information, such as heading to a displayed relaxation spot.
[1256] This allows users to receive optimal route guidance and information based on real-time updates and their emotional state, making self-driving vehicles more comfortable to use.
[1257] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1258] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1259] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1260] [Fourth embodiment]
[1261] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1262] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1263] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1264] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1265] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1266] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1267] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1268] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1269] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1270] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1271] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1272] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1273] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1274] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[1275] overview
[1276] This system allows users to wear wearable devices such as smart glasses and receive real-time updated digital map information in their field of vision. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and sends it to the wearable device.
[1277] Program processing
[1278] Video acquisition and transmission
[1279] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[1280] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[1281] Video analysis and location
[1282] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[1283] Server: Integrates location data with video analytics results to determine the user's exact location.
[1284] Digital Map Generation
[1285] Server: Based on the location data, a real-time digital map is generated, containing up-to-date information on nearby facilities, traffic, and events.
[1286] Server: Optionally overlay additional data layers (e.g., top-rated restaurants or tourist attractions) on the map.
[1287] Digital map transmission and display
[1288] Server: Compresses the latest digital map data and sends it to the wearable device.
[1289] Terminal: Receives compressed map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[1290] Specific examples
[1291] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[1292] 1. When the user is strolling around a tourist spot
[1293] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[1294] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[1295] Server: Sends map data containing the latest tourist spot information to the smart glasses.
[1296] Device: Overlays tourist attractions and directions onto the user's field of view.
[1297] User: Based on this, you can enjoy sightseeing efficiently.
[1298] 2. When a user is walking around a new city
[1299] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[1300] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[1301] Server: Sends the generated map data to the smart glasses.
[1302] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[1303] Users: Using map information displayed in an easy-to-read format, they can reach their destination efficiently and enjoy exploring the city.
[1304] This will allow users to visually receive information that is updated in real time, greatly improving the convenience of travel and sightseeing.
[1305] The processing flow will be explained below.
[1306] Step 1:
[1307] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[1308] Step 2:
[1309] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[1310] Step 3:
[1311] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[1312] Step 4:
[1313] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[1314] Step 5:
[1315] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[1316] Step 6:
[1317] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[1318] Step 7:
[1319] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[1320] Step 8:
[1321] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[1322] Step 9:
[1323] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[1324] Step 10:
[1325] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[1326] Example 1
[1327] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1328] Conventional digital map systems have had difficulty updating information in real time and providing users with accurate location information. As a result, users are unable to efficiently obtain the latest information on their surroundings, traffic, tourist spots, and other information, resulting in low convenience. The present invention aims to solve these problems and provide users with the latest digital map information that is updated in real time.
[1329] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1330] In this invention, the server includes means for analyzing video data, detecting feature points and landmarks, and identifying the user's current location and orientation, means for generating a digital map in real time based on the current location, and means for compressing the generated digital map data and transmitting it to the wearable device, thereby enabling the user to visually receive a digital map that is updated in real time and includes the latest surrounding information, traffic information, and event information.
[1331] "Camera device" refers to a device for capturing video data that is built into a wearable device.
[1332] A "wearable device" is a device worn by the user that has a built-in camera device and GPS module and has the ability to display information within the field of view.
[1333] "Video Data" means data that captures a User's field of view with a camera device and that can be stored or transmitted in digital form.
[1334] "Location data" refers to data indicating the user's current location, typically information obtained from a GPS module.
[1335] A "GPS module" is a hardware component for obtaining geographical location information.
[1336] "Server" refers to a remote computer system that receives data sent from a wearable device and performs analysis and data generation.
[1337] "Feature points" are important points in an image detected by video analysis algorithms and used to identify objects or landmarks.
[1338] A "landmark" refers to a distinctive feature such as a building or terrain that serves as a landmark to identify the user's location and orientation.
[1339] "Real-time" means that data is acquired, processed, and displayed instantly, with minimal delay.
[1340] A "digital map" is a map containing electronically generated geographic information that is used to display a user's current location and surrounding area.
[1341] "Nearby information" refers to information about facilities and features located near the user's current location.
[1342] "Traffic information" refers to information related to travel, such as road congestion conditions and public transportation operation information.
[1343] "Event Information" refers to information about events or activities taking place at a particular location.
[1344] "Compression" is a process for reducing the volume of data and is used to improve transmission efficiency.
[1345] "Overlay" refers to the technology of superimposing information onto the field of view.
[1346] MODE FOR CARRYING OUT THE INVENTION
[1347] This invention relates to a system that provides digital map data in real time. This system consists of a wearable terminal with a built-in camera device, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[1348] System configuration
[1349] This system consists of the following elements:
[1350] Wearable devices
[1351] 1. Camera Device: A device that captures the user's field of view in real time. The camera device captures high-resolution video and stores or transmits the video data in digital format.
[1352] 2. GPS module: A hardware component for obtaining geographical location information, accurately determining the user's current location.
[1353] 3. Display: A device that displays digital maps and other information, with the ability to overlay it on the user's field of view.
[1354] 4. Communication module: This module transmits video data and location data to the server.
[1355] server
[1356] 1. Data Reception Module: This module receives video and location data sent from the wearable device. The received data is compressed and decoded into H.264 or NMEA format.
[1357] 2. Video analysis algorithm: Using the OpenCV library, feature points and landmarks are detected from the received video data, which allows the user's current position and orientation to be determined.
[1358] 3. Position and orientation determination module: Integrates video analysis results and position data, and uses a Kalman filter to estimate the user's exact position and orientation.
[1359] 4. Digital Map Generation Engine: Generates a digital map in real time based on the user's current location, querying and retrieving information about surroundings, traffic, and events from a database (e.g., PostGIS).
[1360] 5. Data compression module: Compresses the generated digital map data in JSON format and sends it to the wearable device.
[1361] 6. Communication module: A device that transmits compressed digital map data to the wearable device.
[1362] Specific examples of implementation
[1363] When the user is exploring a tourist spot
[1364] As the user walks around a tourist spot, the camera in the smart glasses captures the user's field of view, compresses it, and sends it to the server. At the same time, the GPS module acquires the user's current location and sends it to the server. The server analyzes the received video data, detecting feature points and landmarks to determine the user's current location and orientation. A real-time digital map containing up-to-date tourist spot and route information is then sent to the wearable device. The wearable device then overlays the received information on the user's field of view, allowing the user to enjoy sightseeing efficiently.
[1365] Prompt Sentence Examples
[1366] "Explain how you can provide real-time updated digital map information for a scenario where a user is exploring a tourist destination."
[1367] This invention allows users to visually receive digital maps that are updated in real time and include the latest local, traffic, and event information, greatly improving convenience.
[1368] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1369] Step 1: Capture video data
[1370] Terminal: A camera device captures the user's field of view in real time. The input is an image based on the user's field of view, and the output is video data converted into a digital format. This video data has a resolution of 640x480 pixels and is acquired at 30 frames per second (fps).
[1371] Step 2: Compressing the video data
[1372] Terminal: Compresses the acquired video data in H.264 format. The input is the captured raw video data, and the output is compressed video data in H.264 format. Compression reduces the data size and improves transmission efficiency.
[1373] Step 3: Getting location data
[1374] Device: The built-in GPS module obtains the user's current location. The input is signals from GPS satellites, and the output is location data encoded in NMEA format. This data is obtained by obtaining latitude and longitude information, updated every second.
[1375] Step 4: Sending data
[1376] Terminal: Sends compressed video data and acquired location data to the server. The input is compressed video data and location data, and the output is the data sent to the server. Video data is sent using the UDP protocol, and location data is sent using the TCP protocol to ensure communication reliability and efficiency.
[1377] Step 5: Receive and decode the data
[1378] Server: Decodes the received compressed video data using an H.264 decoder and analyzes the location data. The input is the compressed video data and location data, and the output is the decoded video data and analyzed location information. The decoding and analysis process makes it possible to visualize the video and location.
[1379] Step 6: Video analysis and feature point detection
[1380] Server: Extracts feature points (e.g., ORB feature points) from input video data using the OpenCV library. The input is the decoded video data, and the output is the detected feature points and landmark information. Calculates the user's orientation based on the feature points and identifies landmarks.
[1381] Step 7: Identifying position and orientation
[1382] Server: Integrates feature point and location data and uses a Kalman filter to estimate the user's precise location and orientation. The input is feature point information and latitude and longitude data, and the output is estimated location and orientation information. Identifying the location and orientation clarifies the user's spatial state.
[1383] Step 8: Generate a digital map
[1384] Server: Based on the user's current location, obtains surrounding information, traffic information, and event information from a database and generates a real-time digital map. The input is the estimated location and orientation information and a database query, and the output is the generated digital map data. The map generation engine integrates multi-layer data.
[1385] Step 9: Compressing the digital map
[1386] Server: Compresses the generated digital map data in JSON format. The input is the generated real-time map data, and the output is compressed map data in JSON format. Compression improves the efficiency of data transmission.
[1387] Step 10: Sending the digital map
[1388] Server: Sends compressed digital map data to the wearable device. The input is the compressed data, and the output is the data sent to the wearable device. The transmission uses the TCP protocol to ensure data integrity.
[1389] Step 11: Developing the digital map
[1390] Terminal: Decompresses the received compressed data and renders it for display. The input is compressed map data, and the output is a real-time digital map displayed in the user's field of view.
[1391] Step 12: Digital map overlay display
[1392] Terminal: The expanded map data is overlaid on the user's field of view. The input is the expanded map data, and the output is the map and related information displayed in the user's field of view. This allows the user to visually check the latest map information.
[1393] (Application example 1)
[1394] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1395] Conventional autonomous vehicle systems have struggled to efficiently grasp real-time changes in road conditions and traffic information to improve the safety and efficiency of autonomous driving. Furthermore, the effectiveness of navigation systems is limited because the information obtained while the user is actually driving is only retained for a short period of time. Therefore, there is a need for a system that can reflect the latest information about the vehicle's surroundings in real time, improving the navigation accuracy and safety of autonomous vehicles.
[1396] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1397] In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for displaying the digital map on the display of the wearable device, means for acquiring the vehicle's surroundings from the camera device of the autonomous vehicle and dynamically analyzing the surrounding environment based on location information, and means for reflecting the analysis results in the autonomous vehicle's navigation system in real time. This allows the autonomous vehicle to grasp the latest road conditions and traffic information in real time, enabling safe and efficient operation in response to changes in the environment around the vehicle.
[1398] A "camera device" is an optical device for capturing video data.
[1399] A "wearable device" is an electronic device that can be worn by a user and that acquires and displays data.
[1400] "Video data" is visual information captured using a camera device.
[1401] "Location data" is information indicating the current location obtained using a positioning system such as GPS.
[1402] The "receiving means" refers to a device or mechanism for receiving data from the outside.
[1403] "Means for analyzing" are devices and mechanisms for interpreting and converting acquired data into meaningful information.
[1404] "Current location" means the location where the user or autonomous vehicle is currently located.
[1405] A "digital map" is electronic map data that visually represents location information and related information.
[1406] "Generating means" are devices and mechanisms for producing new data or information.
[1407] A "transmitting means" is a device or mechanism for transmitting data to another device or system.
[1408] A "display" is a device for visually displaying information.
[1409] An "autonomous vehicle" is a vehicle that operates autonomously using artificial intelligence and sensor technology.
[1410] "Surrounding conditions" refers to information about the environment in which an autonomous vehicle operates.
[1411] A "navigation system" is a device or mechanism that guides a vehicle to a destination or provides route guidance based on the vehicle's current location.
[1412] "Dynamic analysis means" refers to devices and mechanisms that analyze data in real time and take immediate action or make decisions based on the results.
[1413] System Configuration
[1414] In this invention, the invention is implemented using a system consisting of the following components:
[1415] 1. Hardware
[1416] Camera device: High-resolution cameras installed in autonomous vehicles are used to acquire video data of the surrounding area.
[1417] GPS module: Uses the GPS system installed in the vehicle to determine the current location.
[1418] Wearable device: A device worn by the user, such as smart glasses, that has a built-in camera and display.
[1419] Server: A high-performance server (e.g. AWS EC2) that receives data, analyzes it, and generates digital maps.
[1420] Communication network: Internet connection for sending and receiving data in real time (e.g., 5G communication).
[1421] 2. Software
[1422] Video analysis algorithm: OpenCV and TensorFlow are used to analyze video data and detect feature points and landmarks.
[1423] Digital map generation software: A program that generates digital maps in real time based on the data received.
[1424] Communication protocol: Protocols such as HTTP and WebSocket for sending and receiving data.
[1425] Overview of operation procedures
[1426] When a user operates an autonomous vehicle, the system captures video and location data from the camera device and GPS module. The data is then compressed and sent to a server. The server analyzes the received data to determine the user's surroundings and current location. The server then generates a digital map in real time and sends the data to the autonomous vehicle's navigation system.
[1427] Data processing details
[1428] 1. Acquisition and transmission of video data: Video data acquired by the camera device is compressed and encoded in real time and transmitted to the server along with GPS data.
[1429] 2. Server-side data analysis: The server analyzes the transmitted video data in real time, detecting features and landmarks to determine the current location. This process uses video analysis algorithms such as OpenCV and TensorFlow.
[1430] 3. Digital map generation: Based on the analysis results and location data, the server generates a digital map that reflects the latest information about the surrounding environment, including traffic conditions, obstacle information, and event information.
[1431] 4. Data transmission and display: The generated digital map is compressed and sent in real time to the autonomous vehicle's navigation system, which then adjusts the vehicle's route accordingly based on the received map information.
[1432] Specific examples
[1433] For example, when a vehicle is traveling on a highway, the camera device and GPS module send video and location data to a server. The server then analyzes the video to detect road congestion and obstacles. Based on this information, the server then generates an optimal driving route in real time and sends it to the vehicle's navigation system, enabling the autonomous vehicle to operate safely and efficiently.
[1434] Prompt Sentence Examples
[1435] "Use your location and video data to generate digital map information that updates in real time, providing up-to-date navigation, including traffic and obstacle information."
[1436] This invention can significantly improve the operational safety and efficiency of autonomous vehicles.
[1437] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1438] Step 1:
[1439] The camera device of the autonomous vehicle captures the surrounding video data in real time. The input from the camera device is raw video frames, which are compressed in JPEG or H.264 format.
[1440] Step 2:
[1441] The GPS module of the autonomous vehicle acquires the current location data. The input from the GPS module is latitude and longitude information, and this data is converted into JSON format.
[1442] Step 3:
[1443] The acquired video data and location data are compressed and integrated into a single concatenated data packet at the autonomous vehicle's terminal. The output from the terminal is a concatenated data of compressed video frames and location information.
[1444] Step 4:
[1445] The terminal sends compressed and consolidated data packets to the server. The data is sent through the communication network and received by the server. The server receives the compressed data packets.
[1446] Step 5:
[1447] The server decodes the received data packets and separates the video data from the location data. The decoded input data is the raw video frame and the current location information.
[1448] Step 6:
[1449] The server uses OpenCV and TensorFlow to analyze the video data. The analysis detects landmarks and feature points, which are then used to identify the user's surroundings. The input is raw video data, and the output is analyzed video information (landmarks and feature points).
[1450] Step 7:
[1451] The server combines the received location data with the video analysis results to determine the exact current location. The output of the combination process is accurate current location information.
[1452] Step 8:
[1453] The server generates a digital map in real time based on the current location data. This process includes surrounding information, traffic information, and event information, and also uses existing map data such as OpenStreetMap. The input is precise current location information and additional data (traffic and event information), and the output is digital map data generated in real time.
[1454] Step 9:
[1455] The server compresses the generated digital map and sends it back to the autonomous vehicle's terminal. The data is transmitted over the communication network. The output of the server is compressed digital map data.
[1456] Step 10:
[1457] The terminal decodes the received compressed digital map data and transmits it to the navigation system of the autonomous vehicle. The input of the terminal is compressed digital map data, and the output is decoded map data.
[1458] Step 11:
[1459] The navigation system displays the decoded map data and adjusts the vehicle's driving route in real time, with the input of the navigation system being the decoded map data and the output being the adjusted driving route.
[1460] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1461] This invention relates to a system that not only provides digital map data in real time but also recognizes the user's emotions and customizes information based on those emotions. This system consists of a wearable terminal with a built-in camera device, a server with an emotion engine, a server that receives and analyzes video data and location data, and a wearable terminal that displays the generated digital map.
[1462] overview
[1463] Using this system, users can wear wearable devices such as smart glasses and receive not only real-time updated digital map information in their field of vision, but also information customized to the user's emotions. The server analyzes the user's current location and dynamically generates a digital map including information on nearby areas, traffic, events, etc., and then uses an emotion engine to provide information appropriate to the user's emotions.
[1464] Program processing
[1465] Video acquisition and transmission
[1466] Device: The camera device captures the user's field of vision while wearing the smart glasses, and the video is then converted into digital data, which is then compressed and sent to the server.
[1467] Device: The user's current location is obtained from the GPS module and sent to the server as well.
[1468] Video analysis and location
[1469] Server: The received video data is fed into an analysis algorithm to detect feature points and landmarks in the video, thereby determining the user's current location and facing direction.
[1470] Server: Integrates location data with video analytics results to determine the user's exact location.
[1471] Emotion Recognition and Analysis
[1472] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[1473] Server: Based on the recognized emotional state, the digital map display content related to the user's current location and surrounding information is dynamically changed.
[1474] Digital Map Generation
[1475] Server: Based on the location data mentioned above, obtains geographic information around the current location, including surrounding buildings, roads, points of interest (POI), etc.
[1476] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[1477] Emotion-based information customization
[1478] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[1479] Digital map transmission and display
[1480] Server: Compresses the latest digital map data and sends it to the wearable device, performing the necessary checks to ensure data integrity.
[1481] Terminal: Receives compressed digital map data, decompresses it, and overlays it on the display, allowing users to see the latest information in their field of vision.
[1482] Specific examples
[1483] As a specific example, consider a scenario in which a user is strolling around a tourist spot.
[1484] 1. When the user is strolling around a tourist spot
[1485] Device: As the user walks around the tourist spot, the smart glasses capture images of their surroundings and send them to the server.
[1486] Server: Analyzes the video, identifies the user's current location, and generates map information including up-to-date information on tourist attractions and routes.
[1487] Server: Using the emotion engine, analyze the user's emotional state from their facial expressions and voice and recognize that they are relaxed.
[1488] Server: Generates map data containing information about scenic areas and photo spots where users relax.
[1489] Server: Sends the generated digital map data to the smart glasses.
[1490] Device: Overlays tourist attractions and directions onto the user's field of view.
[1491] User: Based on this, you can enjoy sightseeing efficiently.
[1492] 2. When a user is walking around a new city
[1493] Device: As the user walks through the new city, the smart glasses capture images of their surroundings and send them to a server.
[1494] Server: Analyzes the video, identifies the user's current location, and generates map information including information on nearby popular spots and stores.
[1495] Server: Using an emotion engine, the server analyzes the user's emotional state from their facial expressions and voice and recognizes that they are tired.
[1496] Server: Because the user is tired, it generates map data containing information about cafes and rest spots where the user can relax.
[1497] Server: Sends the generated map data to the smart glasses.
[1498] On the device: Overlays directions and nearby points of interest onto the user's field of view.
[1499] Users: Map information is displayed in an easy-to-read format, allowing them to reach their destination efficiently and enjoy exploring the city.
[1500] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, greatly improving the convenience of travel and sightseeing.
[1501] The processing flow will be explained below.
[1502] Step 1:
[1503] Terminal: The smart glasses camera device is activated and captures the user's field of view in real time. The captured video is compressed and sent to the server.
[1504] Step 2:
[1505] Device: The device obtains the user's current location from the built-in GPS module. This location data is also sent to the server.
[1506] Step 3:
[1507] Server: The received video data and location data are input into the analysis system. First, specific landmarks and feature points are detected from the video data, and the location is confirmed based on this.
[1508] Step 4:
[1509] Server: Integrates the video data analysis results with location data to determine the user's exact location and facing direction, thereby providing accurate visibility information for the user.
[1510] Step 5:
[1511] Server: Utilizes an emotion engine to analyze the user's facial expressions and voice data to recognize the user's emotional state (happiness, sadness, surprise, etc.).
[1512] Step 6:
[1513] Server: Based on the recognized emotional state, the content displayed on the digital map, including the user's current location and surrounding information, is dynamically changed.
[1514] Step 7:
[1515] Server: Based on the identified location, retrieves geographic information around the current location, including surrounding buildings, roads, points of interest (POIs), etc.
[1516] Step 8:
[1517] Server: Generates a digital map with real-time updated neighborhood, traffic, and event information, including key information related to the user's current location and facing direction.
[1518] Step 9:
[1519] Server: Dynamically change the information displayed on the digital map depending on the user's emotional state. For example, if the user is feeling stressed, prioritize displaying information about relaxation spots and cafes.
[1520] Step 10:
[1521] Server: The generated digital map data is compressed and sent to the wearable device, where necessary checks are performed to ensure data integrity.
[1522] Step 11:
[1523] Terminal: Receives compressed digital map data, decodes it, and decompresses it, preparing the data for overlay display.
[1524] Step 12:
[1525] Device: The wearable device displays a digital map overlay that is tailored to the user's current location and facing direction.
[1526] Step 13:
[1527] Users can view digital map information that is updated in real time and make the necessary decisions and move around, allowing them to reach their destination efficiently and take action by utilizing information about their surroundings.
[1528] Step 14:
[1529] Users: Receive personalized information based on their emotions and take action accordingly. For example, if they are feeling stressed, they can find a relaxation spot to help them relax.
[1530] Example 2
[1531] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1532] Conventional navigation systems only provide geographical information and are unable to customize information based on the user's emotional state. This makes it difficult to effectively provide users with the information they need in real time. Furthermore, the analysis and integration of video and location data is insufficient, making it difficult to accurately identify the user's current location and direction.
[1533] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for analyzing video data and identifying the user's current location, a means for analyzing the user's facial expression and voice data and recognizing the user's emotional state, and a means for customizing information based on the emotional state. This makes it possible to provide information in real time according to the user's emotional state and to identify the user's location accurately.
[1534] A "camera device" is an imaging device built into a wearable device to capture the user's field of view.
[1535] A "wearable device" is a device that can be worn by a user and has a built-in camera device and GPS module.
[1536] "Video data" is data in digital form of a field of view captured by a camera device.
[1537] "Location Data" means information about a user's current location obtained using a GPS module or other location-determining technology.
[1538] "Analysis means" refers to software or hardware functionality for processing received data and extracting specific information.
[1539] "Emotional state" refers to emotional information such as joy, sadness, surprise, etc., obtained by analyzing the user's facial expression and voice data.
[1540] A "digital map" is a digital map that contains geographic information about the surrounding area that is dynamically generated based on the user's current location and facing direction.
[1541] "Nearby information" is information about buildings, roads, points of interest (POIs), etc., in the vicinity of the user's current location.
[1542] "Traffic information" refers to real-time traffic information related to the user's current location and route.
[1543] "Event information" is information about events being held in the vicinity of the user's current location.
[1544] "Compression means" refers to data compression algorithms and techniques for efficiently reducing the volume of large amounts of data.
[1545] An "analysis algorithm" is a specific calculation method or procedure used to analyze video, audio, location data, etc.
[1546] An "emotion recognition engine" is software or hardware that identifies a user's emotional state from facial expressions and voice data.
[1547] A "display" is a display device built into a wearable device that visually presents digital maps and other information to the user.
[1548] This invention is a system that uses a wearable device to analyze a user's movement information and emotional state in real time and provides a customized digital map based on the analysis. This system consists of a wearable device with a built-in camera device, a server with an emotion recognition engine, a server that receives and analyzes video data and location data, and a wearable device for displaying the generated digital map.
[1549] Wearable devices are devices with built-in cameras, such as smart glasses, that capture the user's field of vision and obtain it as video data. This data is compressed in a video compression format such as H.264 and sent to a server. Wearable devices also have built-in GPS modules that obtain the user's current location every second and send this location data to the server. Communication is via 4G / 5G networks, and data is transmitted using the TCP / IP protocol.
[1550] The server runs image processing algorithms on the received video data, for example using image processing libraries such as OpenCV to detect features and landmarks in the video. These features are then combined with the received GPS location data to determine the user's exact location and facing direction using triangulation.
[1551] The server then uses an emotion recognition engine to analyze the user's facial expressions and voice data to recognize their emotional state. The engine can be Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API. Based on this emotional state, the server customizes the information provided to the user.
[1552] Based on the user's current location data, the server obtains surrounding geographical, traffic, and event information from sources such as Google Maps API and OpenStreetMap. This information is integrated to generate a digital map that is updated in real time. Information based on the user's emotional state is displayed on the generated digital map. For example, if the user is tired, nearby cafes and relaxation spots are displayed first, while if the user is happy, tourist spots and event information are displayed first.
[1553] The server compresses the generated digital map data and efficiently transmits it to the wearable device using protocol buffers. The wearable device receives the data, expands the compressed data, and overlays it on the display. This allows the user to view the latest map information and customized information in real time in their field of vision.
[1554] As a concrete example, when a user is strolling around a tourist spot, the smart glasses capture video of the user's field of view and send it to a server. The server analyzes the video data and identifies the user's current location. It then uses an emotion recognition engine to analyze the user's emotional state and recognizes that the user is relaxed. The server generates map data including information on beautiful scenery and photo spots at the tourist spot and sends it to the wearable device. The user can then view tourist spots and route information in their field of view, allowing them to enjoy sightseeing efficiently.
[1555] An example of a prompt is as follows:
[1556] "Find a nearby cafe from your current location"
[1557] "Show me photo spots for tourist attractions"
[1558] "Tell me about traffic in this area."
[1559] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1560] Step 1:
[1561] Terminal: The camera device of the smart glasses captures the user's field of view. The captured video data is compressed into H.264 format by an encoder. This compressed video data is the input. The video data is transmitted to the server via a 4G / 5G network using the TCP / IP protocol. The output is the compressed video data that has arrived at the server.
[1562] Step 2:
[1563] Terminal: The GPS module acquires the user's current location every second. The acquired location data is recorded in NMEA format. This location data is the input and is sent to the server, also using the TCP / IP protocol. The output is the location data that arrives at the server.
[1564] Step 3:
[1565] Server: The server analyzes the video data it receives using an image processing library such as OpenCV. The input is compressed video data, which is first decoded and expanded frame by frame. An image processing algorithm (for example, ORB (Oriented FAST and Rotated BRIEF)) is applied to detect feature points and landmarks in the video. The results of this analysis are output.
[1566] Step 4:
[1567] Server: The server combines the feature points of the analyzed video data with the received location data. The input is feature point data and location data. These are combined using triangulation to determine the user's exact location and facing direction. The result of this processing is the output.
[1568] Step 5:
[1569] Server: The server uses an emotion recognition engine to analyze the user's facial expression and voice data. The input is the user's facial expression data and voice data. Software such as Microsoft's Cognitive Services or Google's Cloud Speech-to-Text API is used to identify the emotional state (happiness, sadness, surprise, etc.). The results of this analysis are the output.
[1570] Step 6:
[1571] Server: Customizes information based on the recognized emotional state. Inputs are emotion recognition results and the user's location data. Based on the user's current location, the server obtains surrounding geographical information, traffic information, and event information from Google Maps API, OpenStreetMap, etc., and generates a digital map. The output is customized digital map data.
[1572] Step 7:
[1573] Server: Compresses the generated digital map data and transmits it efficiently to the wearable device using Protocol Buffers. The input is the customized digital map data. The output is transmitted as compressed data.
[1574] Step 8:
[1575] Terminal: The wearable terminal decompresses the compressed digital map data received from the server and overlays it on the display. The input is compressed digital map data. By displaying the decompressed data on the display, the user can see the latest information in real time in their field of vision. This visual information is the final output.
[1576] (Application example 2)
[1577] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1578] In recent years, advances in wearable devices and emotion analysis technology have garnered attention for providing information tailored to user emotions. However, systems that combine these technologies to provide optimal route guidance in real time for autonomous vehicles are still in the early stages of development. Existing systems struggle to efficiently integrate a user's emotional state and current location information to provide customized information in real time. Therefore, the present invention aims to provide real-time digital map information customized based on the user's emotional state, thereby improving passenger comfort in autonomous vehicles.
[1579] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring video data from a wearable device equipped with a camera device, means for receiving video data and location data transmitted from the wearable device, means for analyzing the video data and identifying the user's current location, means for generating a digital map in real time based on the current location, means for transmitting the digital map to the wearable device, means for analyzing the user's emotional state using an emotion engine, means for customizing the digital map based on the analyzed emotional state, means for displaying the customized digital map on the display of the wearable device, and means installed in an autonomous vehicle for providing optimal route guidance based on the passenger's real-time emotion data and current location. This makes it possible to integrate the user's emotional state and current location information and provide appropriate route guidance and surrounding information within the autonomous vehicle.
[1580] A "camera device" is a device for acquiring video data and is built into a wearable device.
[1581] A "wearable terminal" is a device that can be worn by a user and has a built-in camera device and display.
[1582] "Video Data" means a digital representation of visual information captured by a camera device.
[1583] "Location data" refers to data indicating the current location of a user or autonomous vehicle, and is typically obtained by a GPS module.
[1584] A "server" is a high-performance computing device that receives and analyzes video data and location data.
[1585] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice data to recognize their emotional state.
[1586] A "digital map" is map data that includes geographical information, surrounding area information, traffic information, and other information that is generated in real time.
[1587] "Real-time" refers to the time range in which data is processed and updated immediately, with little to no delay.
[1588] "Customizing" means adjusting or changing the information and functionality provided based on information such as the user's emotional state or current location.
[1589] "Display" means a display device built into a wearable device that visually presents digital maps and other information to the user.
[1590] An "autonomous vehicle" is a vehicle that can move independently without human operation.
[1591] "Route guidance" is guide information that shows the optimal route or method to a destination.
[1592] The present invention relates to a system including a wearable terminal with a built-in camera device, a server with an emotion engine, a server for analyzing video data and location data, and a wearable terminal for displaying a generated digital map, which is installed in an autonomous vehicle and provides customized real-time route guidance based on the passenger's emotional state and current location.
[1593] System configuration
[1594] Wearable device: A device such as smart glasses that incorporates a camera device, display, and GPS module.
[1595] Server: A high-performance server, such as Amazon EC2, is used. An emotion engine (e.g., Affectiva), a video analysis algorithm (e.g., OpenCV), and a data compression library (e.g., zlib) are used.
[1596] Autonomous vehicle: A vehicle that is operated autonomously and carried by a user.
[1597] Program processing
[1598] Video acquisition and transmission
[1599] Terminal: The smart glasses' camera device captures the user's field of view and acquires it as video data. The acquired video data is compressed and sent to the server. In addition, the current location data is acquired from the GPS module and sent to the server.
[1600] Video analysis and location
[1601] Server: Receives video data and analyzes it using a video analytics algorithm (e.g., OpenCV) to detect features and landmarks within the field of view. This data is then combined with GPS location data to determine the current location of the autonomous vehicle.
[1602] Emotion Recognition and Analysis
[1603] Server: An emotion engine (e.g., Affectiva) is used to analyze the user's facial expressions and voice data to recognize their emotional state, which can include happiness, sadness, surprise, etc.
[1604] Digital map generation and customization
[1605] Server: Based on the current location data, the server obtains surrounding geographic information, traffic information, and specific points (restaurants, tourist attractions, rest spots, etc.) in real time and generates a digital map. Furthermore, the server customizes the information displayed based on the user's emotional state, as determined by emotion recognition. For example, if the user is relaxed, it will prioritize displaying tourist attractions and relaxation spots, while if the user is feeling stressed, it will prioritize providing information on the shortest route and rest spots.
[1606] Digital map transmission and display
[1607] Server: Sends compressed digital map data to the device.
[1608] Device: Digital map data is deployed and overlaid on the smart glasses display, allowing users to visually check the latest information in real time.
[1609] Specific usage scenarios
[1610] 1. Tourism scenario:
[1611] When a user walks around a tourist spot, the camera device in the smart glasses captures video data and sends it to a server.
[1612] The server analyzes the video data to determine the current location and transmits map data including information on tourist attractions and directions.
[1613] If the emotion engine determines that the user is relaxed, it will provide customized map data including information on scenic areas and photo spots.
[1614] 2. Urban Walking Scenario:
[1615] As a user walks around a new city, the camera device in the smart glasses captures video data and sends it to a server.
[1616] The server analyzes the video data, determines the current location, and then provides map data including information on nearby popular spots and stores.
[1617] If emotion recognition determines that the user is tired, it will prioritize providing information about relaxation spots and cafes.
[1618] Prompt Sentence Examples
[1619] Here are some examples of prompts for a generative AI model:
[1620] If the user is relaxed:
[1621] "Current emotional state: Relaxed. Prioritize information about nearby tourist attractions and relaxation spots."
[1622] If the user is stressed:
[1623] "Current emotional state: Stress. Prioritize showing quicker routes and directions to relaxation spots."
[1624] This allows users to not only visually receive information that is updated in real time, but also obtain optimal information based on their emotional state, making travel and sightseeing more comfortable and enjoyable.
[1625] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1626] Step 1:
[1627] Video acquisition and transmission
[1628] The device (smart glasses) uses a camera device to acquire video data within the user's field of view, including the scenery and objects the user is looking at.
[1629] Input: User's field of view
[1630] Output: Video data
[1631] Specific operation: The video data is compressed and sent to the server in real time. At the same time, the current location data is obtained from the GPS module and sent to the server.
[1632] Step 2:
[1633] Video analysis and location
[1634] The server analyzes the received video data using a video analysis algorithm (e.g., OpenCV), which detects feature points and landmarks within the field of view.
[1635] Input: Video data, location data
[1636] Output: Feature point data, landmark data, current location data
[1637] What it does: Detects features and landmarks in the video and combines this data with GPS location data to determine the user's current location.
[1638] Step 3:
[1639] Emotion Recognition and Analysis
[1640] The server uses an emotion engine (e.g., Affectiva) to analyze the user's facial expressions and voice data to recognize their emotional state.
[1641] Input: User's facial expression data, voice data
[1642] Output: Emotional state data
[1643] Specific behavior: Identify emotional states such as happiness, sadness, and surprise from facial expressions and voice, and store that information in a database.
[1644] Step 4:
[1645] Digital map generation
[1646] The server obtains real-time geographical information, traffic information, and data on specific points (restaurants, tourist attractions, rest spots, etc.) based on the current location data, and generates a digital map.
[1647] Input: Current location data, surrounding geographic information, traffic information, point information
[1648] Output: Digital map data
[1649] Specific operation: Integrates information on surrounding buildings, roads, POIs (points of interest), etc. to generate an up-to-date digital map.
[1650] Step 5:
[1651] Customizing digital maps
[1652] The server customizes the information displayed based on the user's emotional state. For example, if the user is relaxed, it will prioritize showing tourist attractions and relaxation spots, while if the user is stressed, it will prioritize showing information about the shortest routes and rest spots.
[1653] Input: Digital map data, emotional state data
[1654] Output: Customized digital map data
[1655] Specific operation: Dynamically change the displayed content of the digital map according to the user's emotional state to provide the most appropriate information to the user.
[1656] Step 6:
[1657] Digital map transmission and display
[1658] The server compresses the customized digital map data and sends it to the smart glasses.
[1659] Input: Customized digital map data
[1660] Output: Compressed digital map data packets
[1661] Specific operation: Data to be overlaid on the display is compressed and sent to the terminal in real time. The terminal then decompresses the received data and displays it on the display.
[1662] Step 7:
[1663] User receipt and use
[1664] Users can check customized digital map information displayed on the smart glasses display and use it when traveling or sightseeing.
[1665] Input: Customized digital map information displayed on the display
[1666] Output: User actions and choices
[1667] Specific actions: Decide on the next action based on visually presented information, such as heading to a displayed relaxation spot.
[1668] This allows users to receive optimal route guidance and information based on real-time updates and their emotional state, making self-driving vehicles more comfortable to use.
[1669] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1670] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1671] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1672] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1673] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1674] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1675] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1676] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1677] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1678] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1679] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1680] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1681] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1682] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1683] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1684] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1685] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1686] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1687] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1688] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1689] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1690] The following is further disclosed regarding the above embodiment.
[1691] (Claim 1)
[1692] A means for acquiring video data from a wearable terminal having a built-in camera device;
[1693] means for receiving video data and position data transmitted from the wearable device;
[1694] means for analyzing the video data and identifying the user's current location;
[1695] means for generating a digital map in real time based on the current location;
[1696] means for transmitting the digital map to the wearable device;
[1697] means for displaying the digital map on a display of the wearable terminal;
[1698] A system including:
[1699] (Claim 2)
[1700] 2. The system according to claim 1, wherein the digital map includes local area information, traffic information, and event information.
[1701] (Claim 3)
[1702] 10. The system of claim 1, further comprising: means for compressing video data acquired from a camera device of the wearable terminal.
[1703] "Example 1"
[1704] (Claim 1)
[1705] A means for acquiring video data from a wearable terminal having a built-in camera device;
[1706] means for receiving video data and position data transmitted from the wearable device;
[1707] means for compressing the video data and transmitting it to a server;
[1708] means for acquiring the location data from a GPS module and transmitting the data to a server;
[1709] means for analyzing the video data, detecting feature points and landmarks, and identifying the user's current position and orientation;
[1710] means for generating a digital map in real time based on the current location;
[1711] means for compressing the digital map data generated in real time and transmitting the compressed data to the wearable device;
[1712] means for overlaying and displaying the digital map on the display of the wearable device;
[1713] A system including:
[1714] (Claim 2)
[1715] 2. The system according to claim 1, wherein the digital map includes local area information, traffic information, and event information.
[1716] (Claim 3)
[1717] The system of claim 1 further comprising means for determining a user's orientation using the feature points and landmarks.
[1718] "Application Example 1"
[1719] (Claim 1)
[1720] A means for acquiring video data from a wearable terminal having a built-in camera device;
[1721] means for receiving video data and position data transmitted from the wearable device;
[1722] means for analyzing the video data and identifying the user's current location;
[1723] means for generating a digital map in real time based on the current location;
[1724] means for transmitting the digital map to the wearable device;
[1725] means for displaying the digital map on a display of the wearable terminal;
[1726] a means for acquiring the surrounding situation of the vehicle from a camera device of the autonomous driving vehicle and dynamically analyzing the surrounding environment based on the position information;
[1727] a means for reflecting the analysis results in a navigation system for an autonomous vehicle in real time;
[1728] A system including:
[1729] (Claim 2)
[1730] 2. The system according to claim 1, wherein the digital map includes local area information, traffic information, and event information.
[1731] (Claim 3)
[1732] 10. The system of claim 1, further comprising: means for compressing video data acquired from a camera device of the wearable terminal.
[1733] "Example 2: Combining Emotion Engines"
[1734] (Claim 1)
[1735] A means for acquiring video data from a wearable terminal having a built-in camera device;
[1736] means for receiving video data and position data transmitted from the wearable device;
[1737] means for analyzing the video data and identifying the user's current location;
[1738] means for analyzing the user's facial expressions and voice data to recognize the user's emotional state;
[1739] means for customizing information based on said emotional state;
[1740] means for generating a digital map in real time based on the current location;
[1741] means for transmitting the digital map to the wearable device;
[1742] means for displaying the digital map on a display of the wearable terminal;
[1743] A system including:
[1744] (Claim 2)
[1745] 2. The system according to claim 1, wherein the digital map includes local area information, traffic information, and event information.
[1746] (Claim 3)
[1747] 10. The system of claim 1, further comprising: means for compressing video data acquired from a camera device of the wearable terminal.
[1748] "Application example 2 when combining emotion engines"
[1749] (Claim 1)
[1750] A means for acquiring video data from a wearable terminal having a built-in camera device;
[1751] means for receiving video data and position data transmitted from the wearable device;
[1752] means for analyzing the video data and identifying the user's current location;
[1753] means for generating a digital map in real time based on the current location;
[1754] means for transmitting the digital map to the wearable device;
[1755] means for analyzing the emotional state of a user using an emotion engine;
[1756] means for customizing the digital map based on the analyzed emotional state;
[1757] means for displaying the customized digital map on a display of the wearable terminal;
[1758] It will be installed in autonomous vehicles and provide optimal route guidance based on passengers' real-time emotional data and current location.
[1759] A system including:
[1760] (Claim 2)
[1761] 2. The system according to claim 1, wherein the digital map includes local area information, traffic information, and event information.
[1762] (Claim 3)
[1763] 10. The system of claim 1, further comprising: means for compressing video data acquired from a camera device of the wearable terminal. [Explanation of symbols]
[1764] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for acquiring video data from a wearable terminal having a built-in camera device; means for receiving video data and position data transmitted from the wearable device; means for analyzing the video data and identifying the user's current location; means for generating a digital map in real time based on the current location; means for transmitting the digital map to the wearable device; means for displaying the digital map on a display of the wearable terminal; A system including:
2. 2. The system of claim 1, wherein the digital map includes neighborhood information, traffic information, and event information.
3. The system of claim 1 , further comprising: means for compressing video data acquired from a camera device of the wearable terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A