system
The system addresses the challenge of overlooked signs and signals by recreating past accidents using augmented reality, enhancing driver awareness and safety through integrated audio-visual warnings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Conventional traffic safety measures are inadequate in preventing accidents and near misses at dangerous locations due to drivers overlooking signs and signals, particularly for those without a good sense of direction, as they do not effectively convey the characteristics of accident-prone areas.
A system that utilizes location information of past accidents and near misses to recreate past accident situations using video generation technology, displayed via augmented reality, providing multi-stage warnings to drivers through audio and visuals.
Enhances driver awareness of potential dangers by visually and audibly alerting them to accident-prone locations, promoting safe driving habits and reducing the risk of accidents.
Smart Images

Figure 2026069164000001_ABST
Abstract
Description
Technical Field
[0001] The technology of this disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional traffic safety measures, there is a problem that it is difficult to sufficiently prevent accidents and near misses at dangerous locations due to drivers overlooking signs and signals. Especially for drivers without a good sense of direction, they are likely to face unpredictable dangers because they do not understand the characteristics of accident-prone locations. Therefore, in order to promote safe driving and prevent accidents, more effective means of attracting attention are required.
Means for Solving the Problems
[0005] This invention is a system that acquires location information of past accidents and near misses, uses video generation technology to recreate past accident situations based on those locations, and displays this visually to the driver using augmented reality technology. As a result, the system alerts the driver when the vehicle is approaching a dangerous location, supporting safe driving. This system provides information in a visually recognizable format for the driver while driving and provides multi-stage warnings including audio and visuals, enabling the driver to quickly detect and respond to danger.
[0006] "Information on locations where past accidents or near misses occurred" refers to specific geographical data on locations where traffic accidents or near misses have occurred in the past.
[0007] "Video generation technology" refers to the technology that generates dynamic video from source data such as text and still images.
[0008] Augmented reality technology refers to technology that overlays computer-generated information onto the real world, providing information in an augmented form.
[0009] The term "driver" refers to the person who is responsible for operating the vehicle and driving it safely to its destination.
[0010] "Vehicle's current location" refers to the real-time position of a moving object, determined using position measurement technologies such as GPS.
[0011] "Warning" refers to the act of providing a warning about a specific danger or situation requiring attention, thereby improving people's awareness. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Embodiments for Carrying Out the Invention
[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0014] First, the language used in the following description will be explained.
[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention relates to a system that supports safe driving by providing drivers with information on past accident and near-miss locations when they are actually driving a vehicle. This system mainly consists of a server, terminals, and users.
[0034] First, the server collects information on the locations of past accidents and near misses from traffic management agencies and other information sources. This allows it to store information in a database about locations where accidents frequently occur. The server also uses video generation technology to create simulated videos of past accident situations based on the collected document records and video data. These reconstructed videos are important for visually conveying the actual circumstances of an accident to drivers.
[0035] Next, the terminal's role is to utilize accident-prone location data downloaded from the server while driving. The terminal uses GPS to measure the vehicle's current location in real time and issues a warning when approaching an accident-prone area. Specifically, it uses augmented reality technology to display reconstructed accident footage on the windshield, visually alerting the driver to the danger. At this time, voice guidance and visual alerts can also be used to draw the driver's attention.
[0036] The driver, as the user, then receives alerts from the device and takes precautions to drive safely. By receiving hazard information through sight or sound, they can prevent accidents by taking appropriate deceleration and checking their surroundings. With this system, drivers can drive safely even in unfamiliar territory, as if they were familiar with the area.
[0037] The system of the present invention, configured in this manner, aims to prevent accidents from occurring at accident-prone locations and to comprehensively support safe driving.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] The server periodically collects data on past traffic accidents and near misses from traffic management agencies and public databases, and stores this information in a database. The collected data includes the date and time of the accident, the location, detailed circumstances of the accident, and its cause.
[0041] Step 2:
[0042] The server uses video generation technology based on collected data to create videos that recreate past accident situations. In this process, it analyzes document records and fragmented video data to perform simulations as detailed as possible. The generated videos reflect vehicle movements, surrounding obstacles, and traffic signal conditions.
[0043] Step 3:
[0044] The device stores information on accident-prone locations downloaded from the server and continuously measures the vehicle's current location in real time using GPS to determine which accident-prone location it is approaching.
[0045] Step 4:
[0046] When the device detects that it is approaching a dangerous location, it retrieves relevant reconstructed video footage from a server. The retrieved footage is displayed on the windshield using augmented reality technology, providing the driver with a visual alert. Furthermore, it uses audio and animation to further draw attention.
[0047] Step 5:
[0048] Users should review the reenactment videos and warnings provided by their devices and strive to drive safely. Specifically, they should take appropriate driving actions, such as slowing down or carefully checking their surroundings, to prevent accidents.
[0049] In this way, each step works in conjunction to realize the overall function of the system, providing a safer driving environment.
[0050] (Example 1)
[0051] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0052] In recent years, traffic accidents have remained a significant social problem, and there are concerns that the risk of accidents increases, especially when driving in unfamiliar areas. Conventional navigation systems primarily provide route guidance to destinations, but they have difficulty considering information about past accident locations. Therefore, there is a need for systems that support safe driving and prevent accidents in accident-prone areas.
[0053] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0054] In this invention, the server includes means for receiving past accident information and having a database for storing such information, means for analyzing the data and detecting frequent accidents at specific locations, and means for generating simulated videos based on the analyzed accident information using generative AI technology. This enables drivers to receive visual and auditory warnings integrated with actual driving conditions when approaching locations with a high risk of accidents.
[0055] "Past accident information" refers to detailed information about traffic accidents that have occurred in the past, including the date, time, location, and cause.
[0056] A "database" refers to an electronic information storage device that accumulates and organizes information according to a specific method, making it easily searchable.
[0057] "Generative AI technology" refers to technology that uses artificial intelligence to generate new data and information based on underlying data.
[0058] A "simulated video" refers to a virtual video generated based on the actual accident situation, which visually recreates the situation.
[0059] Augmented reality technology refers to a technology that adds or supplements information by overlaying computer graphics (CG) onto visual information from the real world.
[0060] "Driver" refers to the person operating a vehicle.
[0061] "Combined audio and visual warnings" refers to the act of using both audio messages and visual information to alert the driver to danger.
[0062] This invention is a system that supports safe driving by providing drivers with information on past accident and near-miss locations while they are operating a vehicle. This system mainly consists of a server, terminals, and users.
[0063] The server receives past accident information from traffic management agencies via API and securely stores this information in a database. It analyzes the data using programming languages such as Python and the Pandas library to identify accident-prone locations. Furthermore, it utilizes generative AI technology to generate simulated videos based on past accident information. These videos visually recreate the situation at dangerous locations, conveying the realistic risks to drivers.
[0064] The terminal receives accident information and simulated video footage transmitted from the server. Using the vehicle's GPS system, it determines the vehicle's current location in real time and warns the driver by comparing it to accident-prone areas. The terminal also utilizes augmented reality technology to display simulated video footage on the windshield. In addition to visual warnings, voice messages using voice guidance software describe the hazards.
[0065] The driver, as the user, accepts the warnings from the terminal and takes appropriate action to ensure safe driving. For example, if the driver approaches a specific intersection in a city they are visiting for the first time, visual and auditory warnings from the terminal allow them to quickly slow down and carefully navigate the intersection. This system aims to significantly improve the safety of vehicle driving.
[0066] A concrete example is that when a driver is driving in a new city, they can drive with confidence even on unfamiliar roads by referring to simulated video and audio warnings displayed on the terminal. This allows the driver to drive safely as if they were familiar with the area. An example of a prompt would be: "Please describe in detail the AI system that generates reconstructed videos of traffic accidents. Please also describe the specific hardware and software usage, and how to alert the driver."
[0067] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0068] Step 1:
[0069] The server receives historical accident information from traffic management agencies via API. This information includes the date, time, location, and cause of the accident. The received data is stored in a database to prepare for future analysis. Data is entered via API in JSON format and stored in the database in an organized format.
[0070] Step 2:
[0071] The server uses the Python Pandas library to analyze past accident data stored in a database. This analysis uses statistical methods to determine whether accidents are frequent in specific areas or intersections, and identifies accident-prone locations. The input is accident data stored in the database, and the output is a list of accident-prone locations.
[0072] Step 3:
[0073] The server uses generational AI technology to generate videos that simulate past accident situations based on identified accident-prone locations. During this generation process, the AI model is prompted with detailed accident information and outputs realistic accident footage. The videos are intended to visually convey the dangers to drivers.
[0074] Step 4:
[0075] The server sends the generated simulated video and a list of accident-prone locations to the terminal. This transmission uses the SSL / TLS encryption protocol to ensure secure data delivery. The input data consists of the generated video and location list, and the output is the secure delivery of data to the terminal.
[0076] Step 5:
[0077] The terminal processes the received accident information and video. It uses the vehicle's GPS to determine its real-time current location and compares it to accident-prone areas. The input is GPS data and a list of accident locations, and the output is the proximity to accident-prone areas.
[0078] Step 6:
[0079] The device utilizes augmented reality technology to display simulated images on the windshield for the driver. During this process, voice guidance is used to alert the driver through both visual and auditory means. The input for the display consists of video and audio data received from a server, while the output consists of a warning display and an audio message for the driver.
[0080] Step 7:
[0081] The user, the driver, receives alerts from the terminal and strives to drive safely. Based on the simulated images and audio warnings displayed by the terminal, they take appropriate action or slow down the vehicle to prevent accidents. In this process, the input is warning data from the terminal, and the output is the driver's safe driving behavior.
[0082] (Application Example 1)
[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] To improve safety in autonomous vehicles, it is necessary to proactively identify hazards at locations where past accidents have occurred or where warnings are needed, and to support immediate responses by vehicle users and vehicle control systems. In particular, when driving in unfamiliar locations, there is a need for technology that can prevent accidents and near misses resulting from a lack of prior knowledge, and support safe driving.
[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0086] In this invention, the server includes means for acquiring information on past accidents and locations requiring warnings, means for generating reconstructed images based on said location information, and means for displaying the reconstructed images using augmented reality technology. This makes it possible to provide information that supports safe driving in real time and intuitively through a head-mounted display installed in an autonomous vehicle.
[0087] "Information on past accidents and locations requiring caution" refers to data on places where traffic accidents or other hazards have occurred frequently, and is intended to help identify potential dangers at these locations in advance.
[0088] "Reenactment footage" refers to video data that simulates and visually presents past accident situations or situations requiring caution.
[0089] Augmented reality technology is a technology that overlays virtual information onto images of the real world, providing users with an integrated experience of reality and virtuality.
[0090] "Means for measuring current position and monitoring distance" refers to a technological function that measures the vehicle's current location and continuously monitors its relative position to past accident sites.
[0091] A "head-mounted display" is a display device worn by the user to cover their field of vision, enabling the display of augmented reality content directly within their field of view.
[0092] "Means for issuing warnings to vehicle control devices or users" refers to methods for issuing audible or visual warnings to the vehicle's control system or users to draw their attention when the vehicle approaches a dangerous location.
[0093] The system that implements this application consists of a server, a terminal, and a user.
[0094] First, the server collects information on past accidents and locations requiring warnings from traffic management agencies and other sources, and stores it in a database. Next, the server analyzes the data and uses a generative AI model to generate reenactment videos that simulate past accidents and near misses. These reenactment videos are used to visually demonstrate how the accidents occurred.
[0095] The terminal is installed in the autonomous vehicle and uses GPS to obtain the vehicle's current location in real time. The terminal compares the current location information with accident-prone location data downloaded from a server and monitors whether the vehicle is approaching a dangerous location. If it is approaching, it displays a reconstructed image using augmented reality technology via a head-mounted display and alerts the driver and vehicle control system with an audio warning.
[0096] The user (driver) or vehicle control system takes deceleration or other safety measures as needed based on the information presented. This system allows drivers to operate safely in unfamiliar areas, as if they were familiar with the region, by utilizing past accident data.
[0097] As a concrete example, when an autonomous vehicle approaches a complex intersection in a busy area, a reenactment of a past accident is displayed on the head-mounted display, and the autonomous driving system analyzes this information to adjust its speed and take measures to pass through safely.
[0098] An example of a prompt to input into the generating AI model is: "Based on past traffic accident data, recreate and visualize the dangerous conditions at the specified location."
[0099] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0100] Step 1:
[0101] The server retrieves information on past accidents and locations requiring warnings from traffic management agencies and other sources. It uses data from traffic databases and external sources as input to extract location information. As output, it organizes the retrieved location information into a data format and stores it in an internal database.
[0102] Step 2:
[0103] The server uses a generative AI model to generate video footage that recreates past accident situations. In this process, it uses past accident data as input and prompts the AI model with the message, "Based on past traffic accident data, recreate and visualize the dangerous situation at the specified location," before outputting video data. The generated video data is then converted into a format suitable for AR display.
[0104] Step 3:
[0105] The terminal uses an onboard GPS module to obtain the current location of the autonomous vehicle. It receives GPS signals as input and obtains the coordinate information of the current location as output. This coordinate information is used for comparison with accident-prone area data.
[0106] Step 4:
[0107] The device compares the current location with location information based on past accident data and calculates the distance. It uses the current location coordinates and accident location coordinate data as input and calculates the distance as output. This allows it to determine if the vehicle is approaching an accident-prone area.
[0108] Step 5:
[0109] The device displays augmented reality images via a head-mounted display when the vehicle approaches a dangerous location, providing a visual warning to the user. Using reconstructed video data and the user's gaze information as input, the device presents the images at the optimal position within the user's field of view. The output provides intuitively easy-to-understand visual information.
[0110] Step 6:
[0111] Users perform safe driving actions based on visual information and audio warnings from their devices. They receive visual and audio data as input and select appropriate deceleration and steering actions as output. This entire process enables safe driving even in unfamiliar areas.
[0112] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0113] This invention relates to a system that effectively notifies drivers of various dangers they may face while driving, based on past accident information and real-time emotional states. The system mainly consists of three components: a server, a terminal, and a user, as well as an emotion engine.
[0114] First, the server collects data on past accidents and near misses from traffic management agencies and other information sources, and stores the location information and accident details in a database. Based on this, it creates videos that recreate past accident situations using video generation technology. These videos are important for realistically conveying the actual accident situation to drivers.
[0115] Next, the device uses data on accident-prone locations transmitted from the server to track the vehicle's current location in real time via GPS. When the vehicle approaches an accident site, the device displays a reconstructed video to the driver using augmented reality technology to warn them. Simultaneously, the device is equipped with an emotion engine that uses cameras and sensors to analyze the user's emotions in real time based on their facial expressions and voice. Once the emotional state is read, the warning information is customized according to that emotional state and communicated to the driver in the most optimal way.
[0116] For example, if a user shows signs of fatigue or stress, the device reduces the amount of information displayed and softens the voice guidance to reassure the driver. If the device detects that the driver is prone to distraction, it can increase the frequency of displaying reenactment videos to enhance the driver's concentration. In this way, the emotion engine dynamically adjusts the displayed content and intensity to provide the most effective feedback for the driver.
[0117] This system allows users to receive safe driving assistance tailored to their emotional state. By receiving alerts, drivers can reduce the risk of accidents and maintain safe driving. In this way, the present invention utilizes emotion recognition technology to realize more personalized driving assistance.
[0118] The following describes the processing flow.
[0119] Step 1:
[0120] The server retrieves location information, accident details, and the date and time of past traffic accidents and near misses from traffic management agencies and public databases. This data is stored in the database and used as the basis for accident reconstruction videos using video generation technology.
[0121] Step 2:
[0122] The server uses video generation technology based on the collected information to create videos that recreate past accident situations. These videos realistically reproduce vehicle movements, traffic signals, and surrounding conditions, visually alerting drivers to potential dangers in advance.
[0123] Step 3:
[0124] The terminal downloads accident location and reconstructed video data from the server and installs it in the vehicle. The terminal has a built-in GPS function to constantly measure the vehicle's current location.
[0125] Step 4:
[0126] The device detects in real time when the vehicle is approaching an accident site and displays a reconstructed video using augmented reality technology to the driver. The displayed content includes the circumstances of past accidents and serves to alert the driver.
[0127] Step 5:
[0128] An emotion engine is built into the device, using cameras and sensors to analyze the user's facial expressions and tone of voice, recognizing their emotional state in real time. This allows for real-time monitoring of emotional changes while driving.
[0129] Step 6:
[0130] Based on the output of the emotion engine, the device adjusts the content and display method of the reenacted video according to the driver's emotional state. For example, if the driver is showing signs of stress, the information is simplified and calming voice guidance is used.
[0131] Step 7:
[0132] Users can recognize visual and auditory alerts received from their devices and drive safely. By selecting appropriate driving actions based on the information provided, drivers can reduce the risk of accidents.
[0133] (Example 2)
[0134] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0135] In recent years, driver assistance systems have attracted attention as a means of reducing traffic accidents. However, conventional systems often do not utilize past accident data and do not provide information that is tailored to the driver's real-time emotional state. As a result, the timing and method of warnings may not be optimal for the driver. This invention aims to solve these problems by providing a driver assistance system that utilizes past accident data and is optimized according to the driver's emotional state.
[0136] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0137] In this invention, the server includes means for collecting location information of past hazards and near misses from information sources, means for recreating past situations using video generation technology, and means for analyzing the emotional state of the driver assister and adjusting the content presented. This enables real-time, personalized driving assistance based on past accident information.
[0138] "Information sources" refer to sources from which information on past accidents and near misses is obtained, including traffic management agencies and other data providers.
[0139] "Location information" refers to data about the physical locations where past accidents or near misses occurred.
[0140] "Video generation technology" refers to techniques for recreating past accident situations on a computer, and includes computer graphics and simulation techniques.
[0141] Augmented reality technology refers to the technology of overlaying computer-generated elements onto images of the real world.
[0142] A "driver's assistant" refers to the person operating the vehicle, usually known as the driver.
[0143] A "sensor" refers to a device that captures a user's facial expressions and voice and converts them into digital data.
[0144] "Emotional state" refers to the mental state of the driver assistance provider, and includes factors such as fatigue, stress, and concentration.
[0145] A "warning or advisory" refers to a warning or information provided to a driver assistance provider in specific situations.
[0146] The system of the present invention aims to notify a driver assistance user of hazards they may face while driving, based on past accident information and real-time emotional states. It primarily consists of three components: a server, a terminal, and a user, along with an emotion engine.
[0147] The server collects information on past hazards and near-miss locations from information sources such as traffic management agencies and stores this information in a database. It accumulates detailed information such as the date, time, location, and summary of accidents, and uses video generation technology to recreate past accident situations. The videos generated using CG technology and simulation tools are stored in the server's storage. Through this process, the server can provide basic data for gaining a concrete overview of accident situations.
[0148] The terminal receives data on accident-prone locations from a server and tracks the vehicle's position in real time using GPS. If the vehicle approaches a dangerous location, the terminal uses augmented reality technology to present a reconstructed image to the driver assistance system. In addition, the terminal's built-in camera and voice sensors analyze the user's facial expressions and voice, and an emotion engine evaluates their emotional state. Based on this evaluation, the warning information is customized for the driver assistance system. For example, if the user is showing signs of fatigue, the system will reduce the amount of information presented and provide gentler voice guidance to support the driver assistance system.
[0149] Users can visually receive this augmented reality-based information while driving. As their emotional state is analyzed, the frequency of the recreated images and the intensity of the notifications are adjusted, allowing users to receive information tailored to their emotional state. This system enables driver assistants to detect potential hazards early and maintain safe driving.
[0150] As a concrete example, you can use prompt statements like the following to test the emotion engine's output: "Create a driving assistance message appropriate for when the user is tired and has low concentration," or "Suggest a calming message to display when the user is stressed."
[0151] In summary, this system provides support that combines the driver's real-time emotional state with past accident information, thereby creating a safer driving environment.
[0152] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0153] Step 1:
[0154] The server retrieves numerical and location information related to past accidents and near-miss incidents from information sources. It uses the raw data received from each information source as input. This data is converted into a format suitable for database storage and then saved to the database. This process outputs the basic data that will later be used to generate videos.
[0155] Step 2:
[0156] The server uses video generation technology to create a simulated video based on accident location information and other data. It utilizes detailed accident information stored in a database as input. Applying computer graphics (CG) technology, it outputs a video that recreates the actual accident situation. This video is necessary to visually communicate the accident situation to the driver assistance system.
[0157] Step 3:
[0158] The terminal acquires accident-prone location data provided by the server and uses GPS data as input to determine the vehicle's location in real time. When the current location approaches an accident site, it uses augmented reality technology to display a reconstructed video. This outputs information to alert the driver assistance provider.
[0159] Step 4:
[0160] The device's emotion engine uses camera and voice sensors to analyze the user's emotional state from their facial expressions and voice. This sensor data is used as input to calculate the user's fatigue and stress levels. The analysis results are output and used in later steps to customize alert information.
[0161] Step 5:
[0162] The device optimizes the alert information provided to the driver assistance system based on the analyzed emotional state. Using the emotional analysis results as input, it determines the appropriate frequency of voice guidance and visual displays for the driver assistance system. For example, if fatigue is detected, it conveys only the essential information in a calm voice. The output of this process is an alert message optimized for the driver assistance system.
[0163] Step 6:
[0164] The user receives augmented reality images and audio guidance presented from the device. This allows them to enjoy real-time driving assistance based on past accident information and their own emotional state. As an output, direct visual and auditory alerts are promoted for the user.
[0165] (Application Example 2)
[0166] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0167] There is a need to effectively notify drivers of various dangers they may face while driving, based on past accident information and their real-time emotional state. However, conventional systems have struggled to provide optimal feedback tailored to the emotional state of individual drivers. The present invention aims to provide a system that can reduce accident risk and support safe driving by analyzing the driver's emotional state and adjusting danger notifications and warnings based on this analysis.
[0168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0169] In this invention, the server includes means for acquiring location information of past dangerous events, means for recreating past events based on said location information using a video generation method, and means for providing the recreated events to the operator using an augmented reality method. This makes it possible to monitor the distance between the current location of a moving vehicle while it is in motion and the location of past dangerous events, issue a warning to the operator when it approaches a dangerous location, and adjust the content of the warning by analyzing the operator's emotional state.
[0170] "Location information of past dangerous events" refers to data on the geographical locations where dangerous incidents such as accidents or near misses occurred in the past.
[0171] "Image generation techniques" is a general term for technologies that use computer graphics or live-action footage to visually recreate past events.
[0172] Augmented reality techniques are technologies that overlay additional digital information onto information from the real world and present it to the user.
[0173] The term "operator" often refers to the person using the system while driving, i.e., the driver of the vehicle, but in a broader sense, it includes all people who operate the system.
[0174] "Mobile entities" refer to means of moving that change position, such as vehicles and automobiles, and include those with engines.
[0175] "Analyzing emotional state" is the process of analyzing a user's facial expressions, voice, and other physiological responses to infer their mental state at that moment.
[0176] "Adjusting warning content" means optimizing the method and amount of information presented in a warning, according to the user's emotions and circumstances.
[0177] This system acquires information on past dangerous events and provides effective warnings to operators based on their emotional state. Its main components are a server, terminals, and users.
[0178] The server collects location information of past hazardous events from traffic management agencies and other sources and stores it in a database. Based on this information, it creates videos that recreate past events using video generation techniques.
[0179] The terminal uses information transmitted from the server to track the vehicle's current location using a GPS module. When the vehicle approaches a dangerous area, it displays a reconstructed image to the operator using augmented reality techniques and issues a warning. The terminal also has an emotion engine that analyzes the user's emotions in real time via its built-in camera and voice sensors, and can adjust the warning content according to the user's state.
[0180] For example, if the device analyzes that the user is feeling fatigued or stressed, it will reduce the amount of warning information and soften the voice guidance. Conversely, if it determines that concentration is needed, it will increase the number of times the reenactment video is displayed and employ techniques to maximize attention.
[0181] An example of a prompt in a generative AI model is, "Please suggest the most appropriate real-time notification method for safe driving assistance based on past accident information and emotional state." By using this prompt, the system can generate and provide more appropriate feedback to the operator.
[0182] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0183] Step 1:
[0184] The server collects location data on past hazardous events from traffic management agencies and other information sources. Based on this input data, the server updates the database of hazardous events, organizing and storing the location information. This process provides the basic data for events that should be reproduced.
[0185] Step 2:
[0186] The server uses collected location data to create videos that recreate past dangerous events using a video generation method. This video generation process uses location data as input, performs 3D modeling and simulation, and outputs visually recreated videos. This operation generates content that can be used as a warning to the operator.
[0187] Step 3:
[0188] The device acquires location data and reconstructed video footage transmitted from the server. It uses GPS to track the vehicle's current location in real time and compares this location information with the data from the server. Based on this comparison, if the vehicle is approaching a dangerous location, a determination is output.
[0189] Step 4:
[0190] The device uses augmented reality techniques to display recreated images within the operator's field of view. By using cameras and sensors, it overlays digital information onto the real world, providing warnings to the operator. This process creates visual information that blends reality and virtuality.
[0191] Step 5:
[0192] The device uses its built-in camera and voice sensors to analyze the operator's facial expressions and voice, and an emotion engine evaluates their emotional state. It uses video and audio data as input, executes an algorithm to identify the emotional state, and outputs an emotion evaluation. This evaluation is used to adjust warnings.
[0193] Step 6:
[0194] The device adjusts the amount and presentation of warning information according to the user's emotional state. For example, if the user is fatigued, the voice guidance becomes gentler and visual information is reduced. Conversely, if it is determined that concentration is needed, the warning is strengthened. This final warning content is delivered to the user in the most optimal way.
[0195] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0196] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0197] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0198] [Second Embodiment]
[0199] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0200] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0201] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0202] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0203] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0204] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0205] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0206] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0207] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0208] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0209] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0210] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0211] This invention relates to a system that supports safe driving by providing drivers with information on past accident and near-miss locations when they are actually driving a vehicle. This system mainly consists of a server, terminals, and users.
[0212] First, the server collects information on the locations of past accidents and near misses from traffic management agencies and other information sources. This allows it to store information in a database about locations where accidents frequently occur. The server also uses video generation technology to create simulated videos of past accident situations based on the collected document records and video data. These reconstructed videos are important for visually conveying the actual circumstances of an accident to drivers.
[0213] Next, the terminal's role is to utilize accident-prone location data downloaded from the server while driving. The terminal uses GPS to measure the vehicle's current location in real time and issues a warning when approaching an accident-prone area. Specifically, it uses augmented reality technology to display reconstructed accident footage on the windshield, visually alerting the driver to the danger. At this time, voice guidance and visual alerts can also be used to draw the driver's attention.
[0214] The driver, as the user, then receives alerts from the device and takes precautions to drive safely. By receiving hazard information through sight or sound, they can prevent accidents by taking appropriate deceleration and checking their surroundings. With this system, drivers can drive safely even in unfamiliar territory, as if they were familiar with the area.
[0215] The system of the present invention, configured in this manner, aims to prevent accidents from occurring at accident-prone locations and to comprehensively support safe driving.
[0216] The following describes the processing flow.
[0217] Step 1:
[0218] The server periodically collects data on past traffic accidents and near misses from traffic management agencies and public databases, and stores this information in a database. The collected data includes the date and time of the accident, the location, detailed circumstances of the accident, and its cause.
[0219] Step 2:
[0220] The server uses video generation technology based on collected data to create videos that recreate past accident situations. In this process, it analyzes document records and fragmented video data to perform simulations as detailed as possible. The generated videos reflect vehicle movements, surrounding obstacles, and traffic signal conditions.
[0221] Step 3:
[0222] The device stores information on accident-prone locations downloaded from the server and continuously measures the vehicle's current location in real time using GPS to determine which accident-prone location it is approaching.
[0223] Step 4:
[0224] When the device detects that it is approaching a dangerous location, it retrieves relevant reconstructed video footage from a server. The retrieved footage is displayed on the windshield using augmented reality technology, providing the driver with a visual alert. Furthermore, it uses audio and animation to further draw attention.
[0225] Step 5:
[0226] Users should review the reenactment videos and warnings provided by their devices and strive to drive safely. Specifically, they should take appropriate driving actions, such as slowing down or carefully checking their surroundings, to prevent accidents.
[0227] In this way, each step works in conjunction to realize the overall function of the system, providing a safer driving environment.
[0228] (Example 1)
[0229] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0230] In recent years, traffic accidents have remained a significant social problem, and there are concerns that the risk of accidents increases, especially when driving in unfamiliar areas. Conventional navigation systems primarily provide route guidance to destinations, but they have difficulty considering information about past accident locations. Therefore, there is a need for systems that support safe driving and prevent accidents in accident-prone areas.
[0231] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0232] In this invention, the server includes means for receiving past accident information and having a database for storing such information, means for analyzing the data and detecting frequent accidents at specific locations, and means for generating simulated videos based on the analyzed accident information using generative AI technology. This enables drivers to receive visual and auditory warnings integrated with actual driving conditions when approaching locations with a high risk of accidents.
[0233] "Past accident information" refers to detailed information about traffic accidents that have occurred in the past, including the date, time, location, and cause.
[0234] A "database" refers to an electronic information storage device that accumulates and organizes information according to a specific method, making it easily searchable.
[0235] "Generative AI technology" refers to technology that uses artificial intelligence to generate new data and information based on underlying data.
[0236] A "simulated video" refers to a virtual video generated based on the actual accident situation, which visually recreates the situation.
[0237] Augmented reality technology refers to a technology that adds or supplements information by overlaying computer graphics (CG) onto visual information from the real world.
[0238] "Driver" refers to the person operating a vehicle.
[0239] "Combined audio and visual warnings" refers to the act of using both audio messages and visual information to alert the driver to danger.
[0240] This invention is a system that supports safe driving by providing drivers with information on past accident and near-miss locations while they are operating a vehicle. This system mainly consists of a server, terminals, and users.
[0241] The server receives past accident information from traffic management agencies via API and securely stores this information in a database. It analyzes the data using programming languages such as Python and the Pandas library to identify accident-prone locations. Furthermore, it utilizes generative AI technology to generate simulated videos based on past accident information. These videos visually recreate the situation at dangerous locations, conveying the realistic risks to drivers.
[0242] The terminal receives accident information and simulated video footage transmitted from the server. Using the vehicle's GPS system, it determines the vehicle's current location in real time and warns the driver by comparing it to accident-prone areas. The terminal also utilizes augmented reality technology to display simulated video footage on the windshield. In addition to visual warnings, voice messages using voice guidance software describe the hazards.
[0243] The driver, as the user, accepts the warnings from the terminal and takes appropriate action to ensure safe driving. For example, if the driver approaches a specific intersection in a city they are visiting for the first time, visual and auditory warnings from the terminal allow them to quickly slow down and carefully navigate the intersection. This system aims to significantly improve the safety of vehicle driving.
[0244] A concrete example is that when a driver is driving in a new city, they can drive with confidence even on unfamiliar roads by referring to simulated video and audio warnings displayed on the terminal. This allows the driver to drive safely as if they were familiar with the area. An example of a prompt would be: "Please describe in detail the AI system that generates reconstructed videos of traffic accidents. Please also describe the specific hardware and software usage, and how to alert the driver."
[0245] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0246] Step 1:
[0247] The server receives historical accident information from traffic management agencies via API. This information includes the date, time, location, and cause of the accident. The received data is stored in a database to prepare for future analysis. Data is entered via API in JSON format and stored in the database in an organized format.
[0248] Step 2:
[0249] The server uses the Python Pandas library to analyze past accident data stored in a database. This analysis uses statistical methods to determine whether accidents are frequent in specific areas or intersections, and identifies accident-prone locations. The input is accident data stored in the database, and the output is a list of accident-prone locations.
[0250] Step 3:
[0251] The server uses generational AI technology to generate videos that simulate past accident situations based on identified accident-prone locations. During this generation process, the AI model is prompted with detailed accident information and outputs realistic accident footage. The videos are intended to visually convey the dangers to drivers.
[0252] Step 4:
[0253] The server sends the generated simulated video and a list of accident-prone locations to the terminal. This transmission uses SSL / TLS encryption protocols to ensure secure data delivery. The input data consists of the generated video and location list, and the output is the secure delivery of data to the terminal.
[0254] Step 5:
[0255] The terminal processes the received accident information and video. It uses the vehicle's GPS to determine its real-time current location and compares it to accident-prone areas. The input is GPS data and a list of accident locations, and the output is the proximity to accident-prone areas.
[0256] Step 6:
[0257] The device utilizes augmented reality technology to display simulated images on the windshield for the driver. During this process, voice guidance is used to alert the driver through both visual and auditory means. The input for the display consists of video and audio data received from a server, while the output consists of a warning display and an audio message for the driver.
[0258] Step 7:
[0259] The user, the driver, receives alerts from the terminal and strives to drive safely. Based on the simulated images and audio warnings displayed by the terminal, they take appropriate action or slow down the vehicle to prevent accidents. In this process, the input is warning data from the terminal, and the output is the driver's safe driving behavior.
[0260] (Application Example 1)
[0261] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0262] To improve safety in autonomous vehicles, it is necessary to proactively identify hazards at locations where past accidents have occurred or where warnings are needed, and to support immediate responses by vehicle users and vehicle control systems. In particular, when driving in unfamiliar locations, there is a need for technology that can prevent accidents and near misses resulting from a lack of prior knowledge, and support safe driving.
[0263] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0264] In this invention, the server includes means for acquiring information on past accidents and locations requiring warnings, means for generating reconstructed images based on said location information, and means for displaying the reconstructed images using augmented reality technology. This makes it possible to provide information that supports safe driving in real time and intuitively through a head-mounted display installed in an autonomous vehicle.
[0265] "Information on past accidents and locations requiring caution" refers to data on places where traffic accidents or other hazards have occurred frequently, and is intended to help identify potential dangers at these locations in advance.
[0266] "Reenactment footage" refers to video data that simulates and visually presents past accident situations or situations requiring caution.
[0267] Augmented reality technology is a technology that overlays virtual information onto images of the real world, providing users with an integrated experience of reality and virtuality.
[0268] "Means for measuring current position and monitoring distance" refers to a technological function that measures the vehicle's current location and continuously monitors its relative position to past accident sites.
[0269] A "head-mounted display" is a display device worn by the user to cover their field of vision, enabling the display of augmented reality content directly within their field of view.
[0270] "Means for issuing warnings to vehicle control devices or users" refers to methods for issuing audible or visual warnings to the vehicle's control system or users to draw their attention when the vehicle approaches a dangerous location.
[0271] The system that implements this application consists of a server, a terminal, and a user.
[0272] First, the server collects information on past accidents and locations requiring warnings from traffic management agencies and other sources, and stores it in a database. Next, the server analyzes the data and uses a generative AI model to generate reenactment videos that simulate past accidents and near misses. These reenactment videos are used to visually demonstrate how the accidents occurred.
[0273] The terminal is installed in the autonomous vehicle and uses GPS to obtain the vehicle's current location in real time. The terminal compares the current location information with accident-prone location data downloaded from a server and monitors whether the vehicle is approaching a dangerous location. If it is approaching, it displays a reconstructed image using augmented reality technology via a head-mounted display and alerts the driver and vehicle control system with an audio warning.
[0274] The user (driver) or vehicle control system takes deceleration or other safety measures as needed based on the information presented. This system allows drivers to operate safely in unfamiliar areas, as if they were familiar with the region, by utilizing past accident data.
[0275] As a concrete example, when an autonomous vehicle approaches a complex intersection in a busy area, a reenactment of a past accident is displayed on the head-mounted display, and the autonomous driving system analyzes this information to adjust its speed and take measures to pass through safely.
[0276] An example of a prompt to input into the generating AI model is: "Based on past traffic accident data, recreate and visualize the dangerous conditions at the specified location."
[0277] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0278] Step 1:
[0279] The server obtains information on past accidents and locations where attention is required from traffic management agencies and other information sources. Using the data from the traffic database and external information sources as input, it extracts the location information. As output, it organizes the obtained location information into a data format and stores it in the internal database.
[0280] Step 2:
[0281] The server uses a generative AI model to generate a video that reproduces past accident situations. At this time, using past accident data as input, it gives the AI model the prompt sentence "Please reproduce and visualize the dangerous situations at the specified location based on the past traffic accident data." and outputs video data. The generated video data is converted into a format suitable for AR display.
[0282] Step 3:
[0283] The terminal uses the in-vehicle GPS module to obtain the current position of the autonomous vehicle. Receiving the GPS signal as input, it obtains the coordinate information of the current location as output. This coordinate information is used to compare with the accident-prone location data.
[0284] Step 4:
[0285] The terminal collates the location information based on the current location and past accident information and calculates the distance. Using the coordinate of the current location and the coordinate data of the accident location as input, it calculates the distance as output. Thereby, it determines whether the vehicle is approaching an accident-prone location.
[0286] Step 5:
[0287] When the vehicle approaches a dangerous location, the terminal displays an augmented reality video via a head-mounted display to visually warn the user. Using the reproduced video data and the user's gaze information as input, it presents the video at an optimal position within the user's field of view. As output, intuitive and easily understandable visual information is presented.
[0288] Step 6:
[0289] Users perform safe driving actions based on visual information and audio warnings from their devices. They receive visual and audio data as input and select appropriate deceleration and steering actions as output. This entire process enables safe driving even in unfamiliar areas.
[0290] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0291] This invention relates to a system that effectively notifies drivers of various dangers they may face while driving, based on past accident information and real-time emotional states. The system mainly consists of three components: a server, a terminal, and a user, as well as an emotion engine.
[0292] First, the server collects data on past accidents and near misses from traffic management agencies and other information sources, and stores the location information and accident details in a database. Based on this, it creates videos that recreate past accident situations using video generation technology. These videos are important for realistically conveying the actual accident situation to drivers.
[0293] Next, the device uses data on accident-prone locations transmitted from the server to track the vehicle's current location in real time via GPS. When the vehicle approaches an accident site, the device displays a reconstructed video to the driver using augmented reality technology to warn them. Simultaneously, the device is equipped with an emotion engine that uses cameras and sensors to analyze the user's emotions in real time based on their facial expressions and voice. Once the emotional state is read, the warning information is customized according to that emotional state and communicated to the driver in the most optimal way.
[0294] For example, if a user shows signs of fatigue or stress, the device reduces the amount of information displayed and softens the voice guidance to reassure the driver. If the device detects that the driver is prone to distraction, it can increase the frequency of displaying reenactment videos to enhance the driver's concentration. In this way, the emotion engine dynamically adjusts the displayed content and intensity to provide the most effective feedback for the driver.
[0295] This system allows users to receive safe driving assistance tailored to their emotional state. By receiving alerts, drivers can reduce the risk of accidents and maintain safe driving. In this way, the present invention utilizes emotion recognition technology to realize more personalized driving assistance.
[0296] The following describes the processing flow.
[0297] Step 1:
[0298] The server retrieves location information, accident details, and the date and time of past traffic accidents and near misses from traffic management agencies and public databases. This data is stored in the database and used as the basis for accident reconstruction videos using video generation technology.
[0299] Step 2:
[0300] The server uses video generation technology based on the collected information to create videos that recreate past accident situations. These videos realistically reproduce vehicle movements, traffic signals, and surrounding conditions, visually alerting drivers to potential dangers in advance.
[0301] Step 3:
[0302] The terminal downloads accident location and reconstructed video data from the server and installs it in the vehicle. The terminal has a built-in GPS function to constantly measure the vehicle's current location.
[0303] Step 4:
[0304] The terminal determines in real time that the bicycle has approached the accident site and displays a reproduced video using augmented reality technology to the driver. The content to be displayed includes the situation of past accidents, prompting the driver to pay attention.
[0305] Step 5:
[0306] An emotion engine is incorporated into the terminal, which analyzes the user's facial expressions and voice tones using a camera and sensors, and recognizes the emotional state in real time. Thereby, the emotional changes during driving are monitored in real time.
[0307] Step 6:
[0308] Based on the output of the emotion engine, the terminal adjusts the content and display method of the reproduced video according to the driver's emotional state. For example, when the driver shows stress, the information is made concise and a calming voice guide is used.
[0309] Step 7:
[0310] The user confirms the visual and auditory warnings received from the terminal and conducts safe driving. The driver can reduce the risk of accidents by selecting appropriate driving actions based on the provided information.
[0311] (Example 2)
[0312] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0313] <000^986>In recent years, driver assistance systems have attracted attention as a means of reducing traffic accidents. However, conventional systems often do not utilize past accident data and do not provide information that is tailored to the driver's real-time emotional state. As a result, the timing and method of warnings may not be optimal for the driver. This invention aims to solve these problems by providing a driver assistance system that utilizes past accident data and is optimized according to the driver's emotional state.
[0314] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0315] In this invention, the server includes means for collecting location information of past hazards and near misses from information sources, means for recreating past situations using video generation technology, and means for analyzing the emotional state of the driver assister and adjusting the content presented. This enables real-time, personalized driving assistance based on past accident information.
[0316] "Information sources" refer to sources from which information on past accidents and near misses is obtained, including traffic management agencies and other data providers.
[0317] "Location information" refers to data about the physical locations where past accidents or near misses occurred.
[0318] "Video generation technology" refers to techniques for recreating past accident situations on a computer, and includes computer graphics and simulation techniques.
[0319] Augmented reality technology refers to the technology of overlaying computer-generated elements onto images of the real world.
[0320] A "driver's assistant" refers to the person operating the vehicle, usually known as the driver.
[0321] A "sensor" refers to a device that captures a user's facial expressions and voice and converts them into digital data.
[0322] "Emotional state" refers to the mental state of the driver assistance provider, and includes factors such as fatigue, stress, and concentration.
[0323] A "warning or advisory" refers to a warning or information provided to a driver assistance provider in specific situations.
[0324] The system of the present invention aims to notify a driver assistance user of hazards they may face while driving, based on past accident information and real-time emotional states. It primarily consists of three components: a server, a terminal, and a user, along with an emotion engine.
[0325] The server collects information on past hazards and near-miss locations from information sources such as traffic management agencies and stores this information in a database. It accumulates detailed information such as the date, time, location, and summary of accidents, and uses video generation technology to recreate past accident situations. The videos generated using CG technology and simulation tools are stored in the server's storage. Through this process, the server can provide basic data for gaining a concrete overview of accident situations.
[0326] The terminal receives data on accident-prone locations from a server and tracks the vehicle's position in real time using GPS. If the vehicle approaches a dangerous location, the terminal uses augmented reality technology to present a reconstructed image to the driver assistance system. In addition, the terminal's built-in camera and voice sensors analyze the user's facial expressions and voice, and an emotion engine evaluates their emotional state. Based on this evaluation, the warning information is customized for the driver assistance system. For example, if the user is showing signs of fatigue, the system will reduce the amount of information presented and provide gentler voice guidance to support the driver assistance system.
[0327] Users can visually receive this augmented reality-based information while driving. As their emotional state is analyzed, the frequency of the recreated images and the intensity of the notifications are adjusted, allowing users to receive information tailored to their emotional state. This system enables driver assistants to detect potential hazards early and maintain safe driving.
[0328] As a concrete example, you can use prompt statements like the following to test the emotion engine's output: "Create a driving assistance message appropriate for when the user is tired and has low concentration," or "Suggest a calming message to display when the user is stressed."
[0329] In summary, this system provides support that combines the driver's real-time emotional state with past accident information, thereby creating a safer driving environment.
[0330] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0331] Step 1:
[0332] The server retrieves numerical and location information related to past accidents and near-miss incidents from information sources. It uses the raw data received from each information source as input. This data is converted into a format suitable for database storage and then saved to the database. This process outputs the basic data that will later be used to generate videos.
[0333] Step 2:
[0334] The server uses video generation technology to create a simulated video based on accident location information and other data. It utilizes detailed accident information stored in a database as input. Applying computer graphics (CG) technology, it outputs a video that recreates the actual accident situation. This video is necessary to visually communicate the accident situation to the driver assistance system.
[0335] Step 3:
[0336] The terminal acquires accident-prone location data provided by the server and uses GPS data as input to determine the vehicle's location in real time. When the current location approaches an accident site, it uses augmented reality technology to display a reconstructed video. This outputs information to alert the driver assistance provider.
[0337] Step 4:
[0338] The device's emotion engine uses camera and voice sensors to analyze the user's emotional state from their facial expressions and voice. This sensor data is used as input to calculate the user's fatigue and stress levels. The analysis results are output and used in later steps to customize alert information.
[0339] Step 5:
[0340] The device optimizes the alert information provided to the driver assistance system based on the analyzed emotional state. Using the emotional analysis results as input, it determines the appropriate frequency of voice guidance and visual displays for the driver assistance system. For example, if fatigue is detected, it conveys only the essential information in a calm voice. The output of this process is an alert message optimized for the driver assistance system.
[0341] Step 6:
[0342] The user receives augmented reality images and audio guidance presented from the device. This allows them to enjoy real-time driving assistance based on past accident information and their own emotional state. As an output, direct visual and auditory alerts are promoted for the user.
[0343] (Application Example 2)
[0344] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0345] There is a need to effectively notify drivers of various dangers they may face while driving, based on past accident information and their real-time emotional state. However, conventional systems have struggled to provide optimal feedback tailored to the emotional state of individual drivers. The present invention aims to provide a system that can reduce accident risk and support safe driving by analyzing the driver's emotional state and adjusting danger notifications and warnings based on this analysis.
[0346] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0347] In this invention, the server includes means for acquiring location information of past dangerous events, means for recreating past events based on said location information using a video generation method, and means for providing the recreated events to the operator using an augmented reality method. This makes it possible to monitor the distance between the current location of a moving vehicle while it is in motion and the location of past dangerous events, issue a warning to the operator when it approaches a dangerous location, and adjust the content of the warning by analyzing the operator's emotional state.
[0348] "Location information of past dangerous events" refers to data on the geographical locations where dangerous incidents such as accidents or near misses occurred in the past.
[0349] "Image generation techniques" is a general term for technologies that use computer graphics or live-action footage to visually recreate past events.
[0350] Augmented reality techniques are technologies that overlay additional digital information onto information from the real world and present it to the user.
[0351] The term "operator" often refers to the person using the system while driving, i.e., the driver of the vehicle, but in a broader sense, it includes all people who operate the system.
[0352] "Mobile entities" refer to means of moving that change position, such as vehicles and automobiles, and include those with engines.
[0353] "Analyzing emotional state" is the process of analyzing a user's facial expressions, voice, and other physiological responses to infer their mental state at that moment.
[0354] "Adjusting warning content" means optimizing the method and amount of information presented in a warning, according to the user's emotions and circumstances.
[0355] This system acquires information on past dangerous events and provides effective warnings to operators based on their emotional state. Its main components are a server, terminals, and users.
[0356] The server collects location information of past hazardous events from traffic management agencies and other sources and stores it in a database. Based on this information, it creates videos that recreate past events using video generation techniques.
[0357] The terminal uses information transmitted from the server to track the vehicle's current location using a GPS module. When the vehicle approaches a dangerous area, it displays a reconstructed image to the operator using augmented reality techniques and issues a warning. The terminal also has an emotion engine that analyzes the user's emotions in real time via its built-in camera and voice sensors, and can adjust the warning content according to the user's state.
[0358] For example, if the device analyzes that the user is feeling fatigued or stressed, it will reduce the amount of warning information and soften the voice guidance. Conversely, if it determines that concentration is needed, it will increase the number of times the reenactment video is displayed and employ techniques to maximize attention.
[0359] An example of a prompt in a generative AI model is, "Please suggest the most appropriate real-time notification method for safe driving assistance based on past accident information and emotional state." By using this prompt, the system can generate and provide more appropriate feedback to the operator.
[0360] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0361] Step 1:
[0362] The server collects location data on past hazardous events from traffic management agencies and other information sources. Based on this input data, the server updates the database of hazardous events, organizing and storing the location information. This process provides the basic data for events that should be reproduced.
[0363] Step 2:
[0364] The server uses collected location data to create videos that recreate past dangerous events using a video generation method. This video generation process uses location data as input, performs 3D modeling and simulation, and outputs visually recreated videos. This operation generates content that can be used as a warning to the operator.
[0365] Step 3:
[0366] The device acquires location data and reconstructed video footage transmitted from the server. It uses GPS to track the vehicle's current location in real time and compares this location information with the data from the server. Based on this comparison, if the vehicle is approaching a dangerous location, a determination is output.
[0367] Step 4:
[0368] The device uses augmented reality techniques to display recreated images within the operator's field of view. By using cameras and sensors, it overlays digital information onto the real world, providing warnings to the operator. This process creates visual information that blends reality and virtuality.
[0369] Step 5:
[0370] The device uses its built-in camera and voice sensors to analyze the operator's facial expressions and voice, and an emotion engine evaluates their emotional state. It uses video and audio data as input, executes an algorithm to identify the emotional state, and outputs an emotion evaluation. This evaluation is used to adjust warnings.
[0371] Step 6:
[0372] The device adjusts the amount and presentation of warning information according to the user's emotional state. For example, if the user is fatigued, the voice guidance becomes gentler and visual information is reduced. Conversely, if it is determined that concentration is needed, the warning is strengthened. This final warning content is delivered to the user in the most optimal way.
[0373] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0374] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0375] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0376] [Third Embodiment]
[0377] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0378] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0379] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0380] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0381] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0383] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0384] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0385] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0386] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0387] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0388] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0389] This invention relates to a system that supports safe driving by providing drivers with information on past accident and near-miss locations when they are actually driving a vehicle. This system mainly consists of a server, terminals, and users.
[0390] First, the server collects information on the locations of past accidents and near misses from traffic management agencies and other information sources. This allows it to store information in a database about locations where accidents frequently occur. The server also uses video generation technology to create simulated videos of past accident situations based on the collected document records and video data. These reconstructed videos are important for visually conveying the actual circumstances of an accident to drivers.
[0391] Next, the terminal's role is to utilize accident-prone location data downloaded from the server while driving. The terminal uses GPS to measure the vehicle's current location in real time and issues a warning when approaching an accident-prone area. Specifically, it uses augmented reality technology to display reconstructed accident footage on the windshield, visually alerting the driver to the danger. At this time, voice guidance and visual alerts can also be used to draw the driver's attention.
[0392] The driver, as the user, then receives alerts from the device and takes precautions to drive safely. By receiving hazard information through sight or sound, they can prevent accidents by taking appropriate deceleration and checking their surroundings. With this system, drivers can drive safely even in unfamiliar territory, as if they were familiar with the area.
[0393] The system of the present invention, configured in this manner, aims to prevent accidents from occurring at accident-prone locations and to comprehensively support safe driving.
[0394] The following describes the processing flow.
[0395] Step 1:
[0396] The server periodically collects data on past traffic accidents and near misses from traffic management agencies and public databases, and stores this information in a database. The collected data includes the date and time of the accident, the location, detailed circumstances of the accident, and its cause.
[0397] Step 2:
[0398] The server uses video generation technology based on collected data to create videos that recreate past accident situations. In this process, it analyzes document records and fragmented video data to perform simulations as detailed as possible. The generated videos reflect vehicle movements, surrounding obstacles, and traffic signal conditions.
[0399] Step 3:
[0400] The device stores information on accident-prone locations downloaded from the server and continuously measures the vehicle's current location in real time using GPS to determine which accident-prone location it is approaching.
[0401] Step 4:
[0402] When the device detects that it is approaching a dangerous location, it retrieves relevant reconstructed video footage from a server. The retrieved footage is displayed on the windshield using augmented reality technology, providing the driver with a visual alert. Furthermore, it uses audio and animation to further draw attention.
[0403] Step 5:
[0404] Users should review the reenactment videos and warnings provided by their devices and strive to drive safely. Specifically, they should take appropriate driving actions, such as slowing down or carefully checking their surroundings, to prevent accidents.
[0405] In this way, each step works in conjunction to realize the overall function of the system, providing a safer driving environment.
[0406] (Example 1)
[0407] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0408] In recent years, traffic accidents have remained a significant social problem, and there are concerns that the risk of accidents increases, especially when driving in unfamiliar areas. Conventional navigation systems primarily provide route guidance to destinations, but they have difficulty considering information about past accident locations. Therefore, there is a need for systems that support safe driving and prevent accidents in accident-prone areas.
[0409] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0410] In this invention, the server includes means for receiving past accident information and having a database for storing such information, means for analyzing the data and detecting frequent accidents at specific locations, and means for generating simulated videos based on the analyzed accident information using generative AI technology. This enables drivers to receive visual and auditory warnings integrated with actual driving conditions when approaching locations with a high risk of accidents.
[0411] "Past accident information" refers to detailed information about traffic accidents that have occurred in the past, including the date, time, location, and cause.
[0412] A "database" refers to an electronic information storage device that accumulates and organizes information according to a specific method, making it easily searchable.
[0413] "Generative AI technology" refers to technology that uses artificial intelligence to generate new data and information based on underlying data.
[0414] A "simulated video" refers to a virtual video generated based on the actual accident situation, which visually recreates the situation.
[0415] Augmented reality technology refers to a technology that adds or supplements information by overlaying computer graphics (CG) onto visual information from the real world.
[0416] "Driver" refers to the person operating a vehicle.
[0417] "Combined audio and visual warnings" refers to the act of using both audio messages and visual information to alert the driver to danger.
[0418] This invention is a system that supports safe driving by providing drivers with information on past accident and near-miss locations while they are operating a vehicle. This system mainly consists of a server, terminals, and users.
[0419] The server receives past accident information from traffic management agencies via API and securely stores this information in a database. It analyzes the data using programming languages such as Python and the Pandas library to identify accident-prone locations. Furthermore, it utilizes generative AI technology to generate simulated videos based on past accident information. These videos visually recreate the situation at dangerous locations, conveying the realistic risks to drivers.
[0420] The terminal receives accident information and simulated video footage transmitted from the server. Using the vehicle's GPS system, it determines the vehicle's current location in real time and warns the driver by comparing it to accident-prone areas. The terminal also utilizes augmented reality technology to display simulated video footage on the windshield. In addition to visual warnings, voice messages using voice guidance software describe the hazards.
[0421] The driver, as the user, accepts the warnings from the terminal and takes appropriate action to ensure safe driving. For example, if the driver approaches a specific intersection in a city they are visiting for the first time, visual and auditory warnings from the terminal allow them to quickly slow down and carefully navigate the intersection. This system aims to significantly improve the safety of vehicle driving.
[0422] A concrete example is that when a driver is driving in a new city, they can drive with confidence even on unfamiliar roads by referring to simulated video and audio warnings displayed on the terminal. This allows the driver to drive safely as if they were familiar with the area. An example of a prompt would be: "Please describe in detail the AI system that generates reconstructed videos of traffic accidents. Please also describe the specific hardware and software usage, and how to alert the driver."
[0423] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0424] Step 1:
[0425] The server receives historical accident information from traffic management agencies via API. This information includes the date, time, location, and cause of the accident. The received data is stored in a database to prepare for future analysis. Data is entered via API in JSON format and stored in the database in an organized format.
[0426] Step 2:
[0427] The server uses the Python Pandas library to analyze past accident data stored in a database. This analysis uses statistical methods to determine whether accidents are frequent in specific areas or intersections, and identifies accident-prone locations. The input is accident data stored in the database, and the output is a list of accident-prone locations.
[0428] Step 3:
[0429] The server uses generational AI technology to generate videos that simulate past accident situations based on identified accident-prone locations. During this generation process, the AI model is prompted with detailed accident information and outputs realistic accident footage. The videos are intended to visually convey the dangers to drivers.
[0430] Step 4:
[0431] The server sends the generated simulated video and a list of accident-prone locations to the terminal. This transmission uses the SSL / TLS encryption protocol to ensure secure data delivery. The input data consists of the generated video and location list, and the output is the secure delivery of data to the terminal.
[0432] Step 5:
[0433] The terminal processes the received accident information and video. It uses the vehicle's GPS to determine its real-time current location and compares it to accident-prone areas. The input is GPS data and a list of accident locations, and the output is the proximity to accident-prone areas.
[0434] Step 6:
[0435] The device utilizes augmented reality technology to display simulated images on the windshield for the driver. During this process, voice guidance is used to alert the driver through both visual and auditory means. The input for the display consists of video and audio data received from a server, while the output consists of a warning display and an audio message for the driver.
[0436] Step 7:
[0437] The user, the driver, receives alerts from the terminal and strives to drive safely. Based on the simulated images and audio warnings displayed by the terminal, they take appropriate action or slow down the vehicle to prevent accidents. In this process, the input is warning data from the terminal, and the output is the driver's safe driving behavior.
[0438] (Application Example 1)
[0439] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0440] To improve safety in autonomous vehicles, it is necessary to proactively identify hazards at locations where past accidents have occurred or where warnings are needed, and to support immediate responses by vehicle users and vehicle control systems. In particular, when driving in unfamiliar locations, there is a need for technology that can prevent accidents and near misses resulting from a lack of prior knowledge, and support safe driving.
[0441] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0442] In this invention, the server includes means for acquiring information on past accidents and locations requiring warnings, means for generating reconstructed images based on said location information, and means for displaying the reconstructed images using augmented reality technology. This makes it possible to provide information that supports safe driving in real time and intuitively through a head-mounted display installed in an autonomous vehicle.
[0443] "Information on past accidents and locations requiring caution" refers to data on places where traffic accidents or other hazards have occurred frequently, and is intended to help identify potential dangers at these locations in advance.
[0444] "Reenactment footage" refers to video data that simulates and visually presents past accident situations or situations requiring caution.
[0445] Augmented reality technology is a technology that overlays virtual information onto images of the real world, providing users with an integrated experience of reality and virtuality.
[0446] "Means for measuring current position and monitoring distance" refers to a technological function that measures the vehicle's current location and continuously monitors its relative position to past accident sites.
[0447] A "head-mounted display" is a display device worn by the user to cover their field of vision, enabling the display of augmented reality content directly within their field of view.
[0448] "Means for issuing warnings to vehicle control devices or users" refers to methods for issuing audible or visual warnings to the vehicle's control system or users to draw their attention when the vehicle approaches a dangerous location.
[0449] The system that implements this application consists of a server, a terminal, and a user.
[0450] First, the server collects information on past accidents and locations requiring warnings from traffic management agencies and other sources, and stores it in a database. Next, the server analyzes the data and uses a generative AI model to generate reenactment videos that simulate past accidents and near misses. These reenactment videos are used to visually demonstrate how the accidents occurred.
[0451] The terminal is installed in the autonomous vehicle and uses GPS to obtain the vehicle's current location in real time. The terminal compares the current location information with accident-prone location data downloaded from a server and monitors whether the vehicle is approaching a dangerous location. If it is approaching, it displays a reconstructed image using augmented reality technology via a head-mounted display and alerts the driver and vehicle control system with an audio warning.
[0452] The user (driver) or vehicle control system takes deceleration or other safety measures as needed based on the information presented. This system allows drivers to operate safely in unfamiliar areas, as if they were familiar with the region, by utilizing past accident data.
[0453] As a concrete example, when an autonomous vehicle approaches a complex intersection in a busy area, a reenactment of a past accident is displayed on the head-mounted display, and the autonomous driving system analyzes this information to adjust its speed and take measures to pass through safely.
[0454] An example of a prompt to input into the generating AI model is: "Based on past traffic accident data, recreate and visualize the dangerous conditions at the specified location."
[0455] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0456] Step 1:
[0457] The server retrieves information on past accidents and locations requiring warnings from traffic management agencies and other sources. It uses data from traffic databases and external sources as input to extract location information. As output, it organizes the retrieved location information into a data format and stores it in an internal database.
[0458] Step 2:
[0459] The server uses a generative AI model to generate video footage that recreates past accident situations. In this process, it uses past accident data as input and prompts the AI model with the message, "Based on past traffic accident data, recreate and visualize the dangerous situation at the specified location," before outputting video data. The generated video data is then converted into a format suitable for AR display.
[0460] Step 3:
[0461] The terminal uses an onboard GPS module to obtain the current location of the autonomous vehicle. It receives GPS signals as input and obtains the coordinate information of the current location as output. This coordinate information is used for comparison with accident-prone area data.
[0462] Step 4:
[0463] The device compares the current location with location information based on past accident data and calculates the distance. It uses the current location coordinates and accident location coordinate data as input and calculates the distance as output. This allows it to determine if the vehicle is approaching an accident-prone area.
[0464] Step 5:
[0465] The device displays augmented reality images via a head-mounted display when the vehicle approaches a dangerous location, providing a visual warning to the user. Using reconstructed video data and the user's gaze information as input, the device presents the images at the optimal position within the user's field of view. The output provides intuitively easy-to-understand visual information.
[0466] Step 6:
[0467] Users perform safe driving actions based on visual information and audio warnings from their devices. They receive visual and audio data as input and select appropriate deceleration and steering actions as output. This entire process enables safe driving even in unfamiliar areas.
[0468] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0469] This invention relates to a system that effectively notifies drivers of various dangers they may face while driving, based on past accident information and real-time emotional states. The system mainly consists of three components: a server, a terminal, and a user, as well as an emotion engine.
[0470] First, the server collects data on past accidents and near misses from traffic management agencies and other information sources, and stores the location information and accident details in a database. Based on this, it creates videos that recreate past accident situations using video generation technology. These videos are important for realistically conveying the actual accident situation to drivers.
[0471] Next, the device uses data on accident-prone locations transmitted from the server to track the vehicle's current location in real time via GPS. When the vehicle approaches an accident site, the device displays a reconstructed video to the driver using augmented reality technology to warn them. Simultaneously, the device is equipped with an emotion engine that uses cameras and sensors to analyze the user's emotions in real time based on their facial expressions and voice. Once the emotional state is read, the warning information is customized according to that emotional state and communicated to the driver in the most optimal way.
[0472] For example, if a user shows signs of fatigue or stress, the device reduces the amount of information displayed and softens the voice guidance to reassure the driver. If the device detects that the driver is prone to distraction, it can increase the frequency of displaying reenactment videos to enhance the driver's concentration. In this way, the emotion engine dynamically adjusts the displayed content and intensity to provide the most effective feedback for the driver.
[0473] This system allows users to receive safe driving assistance tailored to their emotional state. By receiving alerts, drivers can reduce the risk of accidents and maintain safe driving. In this way, the present invention utilizes emotion recognition technology to realize more personalized driving assistance.
[0474] The following describes the processing flow.
[0475] Step 1:
[0476] The server retrieves location information, accident details, and the date and time of past traffic accidents and near misses from traffic management agencies and public databases. This data is stored in the database and used as the basis for accident reconstruction videos using video generation technology.
[0477] Step 2:
[0478] The server uses video generation technology based on the collected information to create videos that recreate past accident situations. These videos realistically reproduce vehicle movements, traffic signals, and surrounding conditions, visually alerting drivers to potential dangers in advance.
[0479] Step 3:
[0480] The terminal downloads accident location and reconstructed video data from the server and installs it in the vehicle. The terminal has a built-in GPS function to constantly measure the vehicle's current location.
[0481] Step 4:
[0482] The device detects in real time when the vehicle is approaching an accident site and displays a reconstructed video using augmented reality technology to the driver. The displayed content includes the circumstances of past accidents and serves to alert the driver.
[0483] Step 5:
[0484] An emotion engine is built into the device, using cameras and sensors to analyze the user's facial expressions and tone of voice, recognizing their emotional state in real time. This allows for real-time monitoring of emotional changes while driving.
[0485] Step 6:
[0486] Based on the output of the emotion engine, the device adjusts the content and display method of the reenacted video according to the driver's emotional state. For example, if the driver is showing signs of stress, the information is simplified and calming voice guidance is used.
[0487] Step 7:
[0488] Users can recognize visual and auditory alerts received from their devices and drive safely. By selecting appropriate driving actions based on the information provided, drivers can reduce the risk of accidents.
[0489] (Example 2)
[0490] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0491] In recent years, driver assistance systems have attracted attention as a means of reducing traffic accidents. However, conventional systems often do not utilize past accident data and do not provide information that is tailored to the driver's real-time emotional state. As a result, the timing and method of warnings may not be optimal for the driver. This invention aims to solve these problems by providing a driver assistance system that utilizes past accident data and is optimized according to the driver's emotional state.
[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0493] In this invention, the server includes means for collecting location information of past hazards and near misses from information sources, means for recreating past situations using video generation technology, and means for analyzing the emotional state of the driver assister and adjusting the content presented. This enables real-time, personalized driving assistance based on past accident information.
[0494] "Information sources" refer to sources from which information on past accidents and near misses is obtained, including traffic management agencies and other data providers.
[0495] "Location information" refers to data about the physical locations where past accidents or near misses occurred.
[0496] "Video generation technology" refers to techniques for recreating past accident situations on a computer, and includes computer graphics and simulation techniques.
[0497] Augmented reality technology refers to the technology of overlaying computer-generated elements onto images of the real world.
[0498] A "driver's assistant" refers to the person operating the vehicle, usually known as the driver.
[0499] A "sensor" refers to a device that captures a user's facial expressions and voice and converts them into digital data.
[0500] "Emotional state" refers to the mental state of the driver assistance provider, and includes factors such as fatigue, stress, and concentration.
[0501] A "warning or advisory" refers to a warning or information provided to a driver assistance provider in specific situations.
[0502] The system of the present invention aims to notify a driver assistance user of hazards they may face while driving, based on past accident information and real-time emotional states. It primarily consists of three components: a server, a terminal, and a user, along with an emotion engine.
[0503] The server collects information on past hazards and near-miss locations from information sources such as traffic management agencies and stores this information in a database. It accumulates detailed information such as the date, time, location, and summary of accidents, and uses video generation technology to recreate past accident situations. The videos generated using CG technology and simulation tools are stored in the server's storage. Through this process, the server can provide basic data for gaining a concrete overview of accident situations.
[0504] The terminal receives data on accident-prone locations from a server and tracks the vehicle's position in real time using GPS. If the vehicle approaches a dangerous location, the terminal uses augmented reality technology to present a reconstructed image to the driver assistance system. In addition, the terminal's built-in camera and voice sensors analyze the user's facial expressions and voice, and an emotion engine evaluates their emotional state. Based on this evaluation, the warning information is customized for the driver assistance system. For example, if the user is showing signs of fatigue, the system will reduce the amount of information presented and provide gentler voice guidance to support the driver assistance system.
[0505] Users can visually receive this augmented reality-based information while driving. As their emotional state is analyzed, the frequency of the recreated images and the intensity of the notifications are adjusted, allowing users to receive information tailored to their emotional state. This system enables driver assistants to detect potential hazards early and maintain safe driving.
[0506] As a concrete example, you can use prompt statements like the following to test the emotion engine's output: "Create a driving assistance message appropriate for when the user is tired and has low concentration," or "Suggest a calming message to display when the user is stressed."
[0507] In summary, this system provides support that combines the driver's real-time emotional state with past accident information, thereby creating a safer driving environment.
[0508] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0509] Step 1:
[0510] The server retrieves numerical and location information related to past accidents and near-miss incidents from information sources. It uses the raw data received from each information source as input. This data is converted into a format suitable for database storage and then saved to the database. This process outputs the basic data that will later be used to generate videos.
[0511] Step 2:
[0512] The server uses video generation technology to create a simulated video based on accident location information and other data. It utilizes detailed accident information stored in a database as input. Applying computer graphics (CG) technology, it outputs a video that recreates the actual accident situation. This video is necessary to visually communicate the accident situation to the driver assistance system.
[0513] Step 3:
[0514] The terminal acquires accident-prone location data provided by the server and uses GPS data as input to determine the vehicle's location in real time. When the current location approaches an accident site, it uses augmented reality technology to display a reconstructed video. This outputs information to alert the driver assistance provider.
[0515] Step 4:
[0516] The device's emotion engine uses camera and voice sensors to analyze the user's emotional state from their facial expressions and voice. This sensor data is used as input to calculate the user's fatigue and stress levels. The analysis results are output and used in later steps to customize alert information.
[0517] Step 5:
[0518] The device optimizes the alert information provided to the driver assistance system based on the analyzed emotional state. Using the emotional analysis results as input, it determines the appropriate frequency of voice guidance and visual displays for the driver assistance system. For example, if fatigue is detected, it conveys only the essential information in a calm voice. The output of this process is an alert message optimized for the driver assistance system.
[0519] Step 6:
[0520] The user receives augmented reality images and audio guidance presented from the device. This allows them to enjoy real-time driving assistance based on past accident information and their own emotional state. As an output, direct visual and auditory alerts are promoted for the user.
[0521] (Application Example 2)
[0522] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0523] There is a need to effectively notify drivers of various dangers they may face while driving, based on past accident information and their real-time emotional state. However, conventional systems have struggled to provide optimal feedback tailored to the emotional state of individual drivers. The present invention aims to provide a system that can reduce accident risk and support safe driving by analyzing the driver's emotional state and adjusting danger notifications and warnings based on this analysis.
[0524] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0525] In this invention, the server includes means for acquiring location information of past dangerous events, means for recreating past events based on said location information using a video generation method, and means for providing the recreated events to the operator using an augmented reality method. This makes it possible to monitor the distance between the current location of a moving vehicle while it is in motion and the location of past dangerous events, issue a warning to the operator when it approaches a dangerous location, and adjust the content of the warning by analyzing the operator's emotional state.
[0526] "Location information of past dangerous events" refers to data on the geographical locations where dangerous incidents such as accidents or near misses occurred in the past.
[0527] "Image generation techniques" is a general term for technologies that use computer graphics or live-action footage to visually recreate past events.
[0528] Augmented reality techniques are technologies that overlay additional digital information onto information from the real world and present it to the user.
[0529] The term "operator" often refers to the person using the system while driving, i.e., the driver of the vehicle, but in a broader sense, it includes all people who operate the system.
[0530] "Mobile entities" refer to means of moving that change position, such as vehicles and automobiles, and include those with engines.
[0531] "Analyzing emotional state" is the process of analyzing a user's facial expressions, voice, and other physiological responses to infer their mental state at that moment.
[0532] "Adjusting warning content" means optimizing the method and amount of information presented in a warning, according to the user's emotions and circumstances.
[0533] This system acquires information on past dangerous events and provides effective warnings to operators based on their emotional state. Its main components are a server, terminals, and users.
[0534] The server collects location information of past hazardous events from traffic management agencies and other sources and stores it in a database. Based on this information, it creates videos that recreate past events using video generation techniques.
[0535] The terminal uses information transmitted from the server to track the vehicle's current location using a GPS module. When the vehicle approaches a dangerous area, it displays a reconstructed image to the operator using augmented reality techniques and issues a warning. The terminal also has an emotion engine that analyzes the user's emotions in real time via its built-in camera and voice sensors, and can adjust the warning content according to the user's state.
[0536] For example, if the device analyzes that the user is feeling fatigued or stressed, it will reduce the amount of warning information and soften the voice guidance. Conversely, if it determines that concentration is needed, it will increase the number of times the reenactment video is displayed and employ techniques to maximize attention.
[0537] An example of a prompt in a generative AI model is, "Please suggest the most appropriate real-time notification method for safe driving assistance based on past accident information and emotional state." By using this prompt, the system can generate and provide more appropriate feedback to the operator.
[0538] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0539] Step 1:
[0540] The server collects location data on past hazardous events from traffic management agencies and other information sources. Based on this input data, the server updates the database of hazardous events, organizing and storing the location information. This process provides the basic data for events that should be reproduced.
[0541] Step 2:
[0542] The server uses collected location data to create videos that recreate past dangerous events using a video generation method. This video generation process uses location data as input, performs 3D modeling and simulation, and outputs visually recreated videos. This operation generates content that can be used as a warning to the operator.
[0543] Step 3:
[0544] The device acquires location data and reconstructed video footage transmitted from the server. It uses GPS to track the vehicle's current location in real time and compares this location information with the data from the server. Based on this comparison, if the vehicle is approaching a dangerous location, a determination is output.
[0545] Step 4:
[0546] The device uses augmented reality techniques to display recreated images within the operator's field of view. By using cameras and sensors, it overlays digital information onto the real world, providing warnings to the operator. This process creates visual information that blends reality and virtuality.
[0547] Step 5:
[0548] The device uses its built-in camera and voice sensors to analyze the operator's facial expressions and voice, and an emotion engine evaluates their emotional state. It uses video and audio data as input, executes an algorithm to identify the emotional state, and outputs an emotion evaluation. This evaluation is used to adjust warnings.
[0549] Step 6:
[0550] The device adjusts the amount and presentation of warning information according to the user's emotional state. For example, if the user is fatigued, the voice guidance becomes gentler and visual information is reduced. Conversely, if it is determined that concentration is needed, the warning is strengthened. This final warning content is delivered to the user in the most optimal way.
[0551] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0552] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0553] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0554] [Fourth Embodiment]
[0555] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0556] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0557] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0558] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0559] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0560] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0561] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0562] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0563] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0564] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0565] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0566] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0567] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0568] This invention relates to a system that supports safe driving by providing drivers with information on past accident and near-miss locations when they are actually driving a vehicle. This system mainly consists of a server, terminals, and users.
[0569] First, the server collects information on the locations of past accidents and near misses from traffic management agencies and other information sources. This allows it to store information in a database about locations where accidents frequently occur. The server also uses video generation technology to create simulated videos of past accident situations based on the collected document records and video data. These reconstructed videos are important for visually conveying the actual circumstances of an accident to drivers.
[0570] Next, the terminal's role is to utilize accident-prone location data downloaded from the server while driving. The terminal uses GPS to measure the vehicle's current location in real time and issues a warning when approaching an accident-prone area. Specifically, it uses augmented reality technology to display reconstructed accident footage on the windshield, visually alerting the driver to the danger. At this time, voice guidance and visual alerts can also be used to draw the driver's attention.
[0571] The driver, as the user, then receives alerts from the device and takes precautions to drive safely. By receiving hazard information through sight or sound, they can prevent accidents by taking appropriate deceleration and checking their surroundings. With this system, drivers can drive safely even in unfamiliar territory, as if they were familiar with the area.
[0572] The system of the present invention, configured in this manner, aims to prevent accidents from occurring at accident-prone locations and to comprehensively support safe driving.
[0573] The following describes the processing flow.
[0574] Step 1:
[0575] The server periodically collects data on past traffic accidents and near misses from traffic management agencies and public databases, and stores this information in a database. The collected data includes the date and time of the accident, the location, detailed circumstances of the accident, and its cause.
[0576] Step 2:
[0577] The server uses video generation technology based on collected data to create videos that recreate past accident situations. In this process, it analyzes document records and fragmented video data to perform simulations as detailed as possible. The generated videos reflect vehicle movements, surrounding obstacles, and traffic signal conditions.
[0578] Step 3:
[0579] The device stores information on accident-prone locations downloaded from the server and continuously measures the vehicle's current location in real time using GPS to determine which accident-prone location it is approaching.
[0580] Step 4:
[0581] When the device detects that it is approaching a dangerous location, it retrieves relevant reconstructed video footage from a server. The retrieved footage is displayed on the windshield using augmented reality technology, providing the driver with a visual alert. Furthermore, it uses audio and animation to further draw attention.
[0582] Step 5:
[0583] Users should review the reenactment videos and warnings provided by their devices and strive to drive safely. Specifically, they should take appropriate driving actions, such as slowing down or carefully checking their surroundings, to prevent accidents.
[0584] In this way, each step works in conjunction to realize the overall function of the system, providing a safer driving environment.
[0585] (Example 1)
[0586] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0587] In recent years, traffic accidents have remained a significant social problem, and there are concerns that the risk of accidents increases, especially when driving in unfamiliar areas. Conventional navigation systems primarily provide route guidance to destinations, but they have difficulty considering information about past accident locations. Therefore, there is a need for systems that support safe driving and prevent accidents in accident-prone areas.
[0588] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0589] In this invention, the server includes means for receiving past accident information and having a database for storing such information, means for analyzing the data and detecting frequent accidents at specific locations, and means for generating simulated videos based on the analyzed accident information using generative AI technology. This enables drivers to receive visual and auditory warnings integrated with actual driving conditions when approaching locations with a high risk of accidents.
[0590] "Past accident information" refers to detailed information about traffic accidents that have occurred in the past, including the date, time, location, and cause.
[0591] A "database" refers to an electronic information storage device that accumulates and organizes information according to a specific method, making it easily searchable.
[0592] "Generative AI technology" refers to technology that uses artificial intelligence to generate new data and information based on underlying data.
[0593] A "simulated video" refers to a virtual video generated based on the actual accident situation, which visually recreates the situation.
[0594] Augmented reality technology refers to a technology that adds or supplements information by overlaying computer graphics (CG) onto visual information from the real world.
[0595] "Driver" refers to the person operating a vehicle.
[0596] "Combined audio and visual warnings" refers to the act of using both audio messages and visual information to alert the driver to danger.
[0597] This invention is a system that supports safe driving by providing drivers with information on past accident and near-miss locations while they are operating a vehicle. This system mainly consists of a server, terminals, and users.
[0598] The server receives past accident information from traffic management agencies via API and securely stores this information in a database. It analyzes the data using programming languages such as Python and the Pandas library to identify accident-prone locations. Furthermore, it utilizes generative AI technology to generate simulated videos based on past accident information. These videos visually recreate the situation at dangerous locations, conveying the realistic risks to drivers.
[0599] The terminal receives accident information and simulated video footage transmitted from the server. Using the vehicle's GPS system, it determines the vehicle's current location in real time and warns the driver by comparing it to accident-prone areas. The terminal also utilizes augmented reality technology to display simulated video footage on the windshield. In addition to visual warnings, voice messages using voice guidance software describe the hazards.
[0600] The driver, as the user, accepts the warnings from the terminal and takes appropriate action to ensure safe driving. For example, if the driver approaches a specific intersection in a city they are visiting for the first time, visual and auditory warnings from the terminal allow them to quickly slow down and carefully navigate the intersection. This system aims to significantly improve the safety of vehicle driving.
[0601] A concrete example is that when a driver is driving in a new city, they can drive with confidence even on unfamiliar roads by referring to simulated video and audio warnings displayed on the terminal. This allows the driver to drive safely as if they were familiar with the area. An example of a prompt would be: "Please describe in detail the AI system that generates reconstructed videos of traffic accidents. Please also describe the specific hardware and software usage, and how to alert the driver."
[0602] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0603] Step 1:
[0604] The server receives historical accident information from traffic management agencies via API. This information includes the date, time, location, and cause of the accident. The received data is stored in a database to prepare for future analysis. Data is entered via API in JSON format and stored in the database in an organized format.
[0605] Step 2:
[0606] The server uses the Python Pandas library to analyze past accident data stored in a database. This analysis uses statistical methods to determine whether accidents are frequent in specific areas or intersections, and identifies accident-prone locations. The input is accident data stored in the database, and the output is a list of accident-prone locations.
[0607] Step 3:
[0608] The server uses generational AI technology to generate videos that simulate past accident situations based on identified accident-prone locations. During this generation process, the AI model is prompted with detailed accident information and outputs realistic accident footage. The videos are intended to visually convey the dangers to drivers.
[0609] Step 4:
[0610] The server sends the generated simulated video and a list of accident-prone locations to the terminal. This transmission uses the SSL / TLS encryption protocol to ensure secure data delivery. The input data consists of the generated video and location list, and the output is the secure delivery of data to the terminal.
[0611] Step 5:
[0612] The terminal processes the received accident information and video. It uses the vehicle's GPS to determine its real-time current location and compares it to accident-prone areas. The input is GPS data and a list of accident locations, and the output is the proximity to accident-prone areas.
[0613] Step 6:
[0614] The device utilizes augmented reality technology to display simulated images on the windshield for the driver. During this process, voice guidance is used to alert the driver through both visual and auditory means. The input for the display consists of video and audio data received from a server, while the output consists of a warning display and an audio message for the driver.
[0615] Step 7:
[0616] The user, the driver, receives alerts from the terminal and strives to drive safely. Based on the simulated images and audio warnings displayed by the terminal, they take appropriate action or slow down the vehicle to prevent accidents. In this process, the input is warning data from the terminal, and the output is the driver's safe driving behavior.
[0617] (Application Example 1)
[0618] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0619] To improve safety in autonomous vehicles, it is necessary to proactively identify hazards at locations where past accidents have occurred or where warnings are needed, and to support immediate responses by vehicle users and vehicle control systems. In particular, when driving in unfamiliar locations, there is a need for technology that can prevent accidents and near misses resulting from a lack of prior knowledge, and support safe driving.
[0620] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0621] In this invention, the server includes means for acquiring information on past accidents and locations requiring warnings, means for generating reconstructed images based on said location information, and means for displaying the reconstructed images using augmented reality technology. This makes it possible to provide information that supports safe driving in real time and intuitively through a head-mounted display installed in an autonomous vehicle.
[0622] "Information on past accidents and locations requiring caution" refers to data on places where traffic accidents or other hazards have occurred frequently, and is intended to help identify potential dangers at these locations in advance.
[0623] "Reenactment footage" refers to video data that simulates and visually presents past accident situations or situations requiring caution.
[0624] Augmented reality technology is a technology that overlays virtual information onto images of the real world, providing users with an integrated experience of reality and virtuality.
[0625] "Means for measuring current position and monitoring distance" refers to a technological function that measures the vehicle's current location and continuously monitors its relative position to past accident sites.
[0626] A "head-mounted display" is a display device worn by the user to cover their field of vision, enabling the display of augmented reality content directly within their field of view.
[0627] "Means for issuing warnings to vehicle control devices or users" refers to methods for issuing audible or visual warnings to the vehicle's control system or users to draw their attention when the vehicle approaches a dangerous location.
[0628] The system that implements this application consists of a server, a terminal, and a user.
[0629] First, the server collects information on past accidents and locations requiring warnings from traffic management agencies and other sources, and stores it in a database. Next, the server analyzes the data and uses a generative AI model to generate reenactment videos that simulate past accidents and near misses. These reenactment videos are used to visually demonstrate how the accidents occurred.
[0630] The terminal is installed in the autonomous vehicle and uses GPS to obtain the vehicle's current location in real time. The terminal compares the current location information with accident-prone location data downloaded from a server and monitors whether the vehicle is approaching a dangerous location. If it is approaching, it displays a reconstructed image using augmented reality technology via a head-mounted display and alerts the driver and vehicle control system with an audio warning.
[0631] The user (driver) or vehicle control system takes deceleration or other safety measures as needed based on the information presented. This system allows drivers to operate safely in unfamiliar areas, as if they were familiar with the region, by utilizing past accident data.
[0632] As a concrete example, when an autonomous vehicle approaches a complex intersection in a busy area, a reenactment of a past accident is displayed on the head-mounted display, and the autonomous driving system analyzes this information to adjust its speed and take measures to pass through safely.
[0633] An example of a prompt to input into the generating AI model is: "Based on past traffic accident data, recreate and visualize the dangerous conditions at the specified location."
[0634] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0635] Step 1:
[0636] The server retrieves information on past accidents and locations requiring warnings from traffic management agencies and other sources. It uses data from traffic databases and external sources as input to extract location information. As output, it organizes the retrieved location information into a data format and stores it in an internal database.
[0637] Step 2:
[0638] The server uses a generative AI model to generate video footage that recreates past accident situations. In this process, it uses past accident data as input and prompts the AI model with the message, "Based on past traffic accident data, recreate and visualize the dangerous situation at the specified location," before outputting video data. The generated video data is then converted into a format suitable for AR display.
[0639] Step 3:
[0640] The terminal uses an onboard GPS module to obtain the current location of the autonomous vehicle. It receives GPS signals as input and obtains the coordinate information of the current location as output. This coordinate information is used for comparison with accident-prone area data.
[0641] Step 4:
[0642] The device compares the current location with location information based on past accident data and calculates the distance. It uses the current location coordinates and accident location coordinate data as input and calculates the distance as output. This allows it to determine if the vehicle is approaching an accident-prone area.
[0643] Step 5:
[0644] The device displays augmented reality images via a head-mounted display when the vehicle approaches a dangerous location, providing a visual warning to the user. Using reconstructed video data and the user's gaze information as input, the device presents the images at the optimal position within the user's field of view. The output provides intuitively easy-to-understand visual information.
[0645] Step 6:
[0646] Users perform safe driving actions based on visual information and audio warnings from their devices. They receive visual and audio data as input and select appropriate deceleration and steering actions as output. This entire process enables safe driving even in unfamiliar areas.
[0647] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0648] This invention relates to a system that effectively notifies drivers of various dangers they may face while driving, based on past accident information and real-time emotional states. The system mainly consists of three components: a server, a terminal, and a user, as well as an emotion engine.
[0649] First, the server collects data on past accidents and near misses from traffic management agencies and other information sources, and stores the location information and accident details in a database. Based on this, it creates videos that recreate past accident situations using video generation technology. These videos are important for realistically conveying the actual accident situation to drivers.
[0650] Next, the device uses data on accident-prone locations transmitted from the server to track the vehicle's current location in real time via GPS. When the vehicle approaches an accident site, the device displays a reconstructed video to the driver using augmented reality technology to warn them. Simultaneously, the device is equipped with an emotion engine that uses cameras and sensors to analyze the user's emotions in real time based on their facial expressions and voice. Once the emotional state is read, the warning information is customized according to that emotional state and communicated to the driver in the most optimal way.
[0651] For example, if a user shows signs of fatigue or stress, the device reduces the amount of information displayed and softens the voice guidance to reassure the driver. If the device detects that the driver is prone to distraction, it can increase the frequency of displaying reenactment videos to enhance the driver's concentration. In this way, the emotion engine dynamically adjusts the displayed content and intensity to provide the most effective feedback for the driver.
[0652] This system allows users to receive safe driving assistance tailored to their emotional state. By receiving alerts, drivers can reduce the risk of accidents and maintain safe driving. In this way, the present invention utilizes emotion recognition technology to realize more personalized driving assistance.
[0653] The following describes the processing flow.
[0654] Step 1:
[0655] The server retrieves location information, accident details, and the date and time of past traffic accidents and near misses from traffic management agencies and public databases. This data is stored in the database and used as the basis for accident reconstruction videos using video generation technology.
[0656] Step 2:
[0657] The server uses video generation technology based on the collected information to create videos that recreate past accident situations. These videos realistically reproduce vehicle movements, traffic signals, and surrounding conditions, visually alerting drivers to potential dangers in advance.
[0658] Step 3:
[0659] The terminal downloads accident location and reconstructed video data from the server and installs it in the vehicle. The terminal has a built-in GPS function to constantly measure the vehicle's current location.
[0660] Step 4:
[0661] The device detects in real time when the vehicle is approaching an accident site and displays a reconstructed video using augmented reality technology to the driver. The displayed content includes the circumstances of past accidents and serves to alert the driver.
[0662] Step 5:
[0663] An emotion engine is built into the device, using cameras and sensors to analyze the user's facial expressions and tone of voice, recognizing their emotional state in real time. This allows for real-time monitoring of emotional changes while driving.
[0664] Step 6:
[0665] Based on the output of the emotion engine, the device adjusts the content and display method of the reenacted video according to the driver's emotional state. For example, if the driver is showing signs of stress, the information is simplified and calming voice guidance is used.
[0666] Step 7:
[0667] Users can recognize visual and auditory alerts received from their devices and drive safely. By selecting appropriate driving actions based on the information provided, drivers can reduce the risk of accidents.
[0668] (Example 2)
[0669] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0670] In recent years, driver assistance systems have attracted attention as a means of reducing traffic accidents. However, conventional systems often do not utilize past accident data and do not provide information that is tailored to the driver's real-time emotional state. As a result, the timing and method of warnings may not be optimal for the driver. This invention aims to solve these problems by providing a driver assistance system that utilizes past accident data and is optimized according to the driver's emotional state.
[0671] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0672] In this invention, the server includes means for collecting location information of past hazards and near misses from information sources, means for recreating past situations using video generation technology, and means for analyzing the emotional state of the driver assister and adjusting the content presented. This enables real-time, personalized driving assistance based on past accident information.
[0673] "Information sources" refer to sources from which information on past accidents and near misses is obtained, including traffic management agencies and other data providers.
[0674] "Location information" refers to data about the physical locations where past accidents or near misses occurred.
[0675] "Video generation technology" refers to techniques for recreating past accident situations on a computer, and includes computer graphics and simulation techniques.
[0676] Augmented reality technology refers to the technology of overlaying computer-generated elements onto images of the real world.
[0677] A "driver's assistant" refers to the person operating the vehicle, usually known as the driver.
[0678] A "sensor" refers to a device that captures a user's facial expressions and voice and converts them into digital data.
[0679] "Emotional state" refers to the mental state of the driver assistance provider, and includes factors such as fatigue, stress, and concentration.
[0680] A "warning or advisory" refers to a warning or information provided to a driver assistance provider in specific situations.
[0681] The system of the present invention aims to notify a driver assistance user of hazards they may face while driving, based on past accident information and real-time emotional states. It primarily consists of three components: a server, a terminal, and a user, along with an emotion engine.
[0682] The server collects information on past hazards and near-miss locations from information sources such as traffic management agencies and stores this information in a database. It accumulates detailed information such as the date, time, location, and summary of accidents, and uses video generation technology to recreate past accident situations. The videos generated using CG technology and simulation tools are stored in the server's storage. Through this process, the server can provide basic data for gaining a concrete overview of accident situations.
[0683] The terminal receives data on accident-prone locations from a server and tracks the vehicle's position in real time using GPS. If the vehicle approaches a dangerous location, the terminal uses augmented reality technology to present a reconstructed image to the driver assistance system. In addition, the terminal's built-in camera and voice sensors analyze the user's facial expressions and voice, and an emotion engine evaluates their emotional state. Based on this evaluation, the warning information is customized for the driver assistance system. For example, if the user is showing signs of fatigue, the system will reduce the amount of information presented and provide gentler voice guidance to support the driver assistance system.
[0684] Users can visually receive this augmented reality-based information while driving. As their emotional state is analyzed, the frequency of the recreated images and the intensity of the notifications are adjusted, allowing users to receive information tailored to their emotional state. This system enables driver assistants to detect potential hazards early and maintain safe driving.
[0685] As a concrete example, you can use prompt statements like the following to test the emotion engine's output: "Create a driving assistance message appropriate for when the user is tired and has low concentration," or "Suggest a calming message to display when the user is stressed."
[0686] In summary, this system provides support that combines the driver's real-time emotional state with past accident information, thereby creating a safer driving environment.
[0687] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0688] Step 1:
[0689] The server retrieves numerical and location information related to past accidents and near-miss incidents from information sources. It uses the raw data received from each information source as input. This data is converted into a format suitable for database storage and then saved to the database. This process outputs the basic data that will later be used to generate videos.
[0690] Step 2:
[0691] The server uses video generation technology to create a simulated video based on accident location information and other data. It utilizes detailed accident information stored in a database as input. Applying computer graphics (CG) technology, it outputs a video that recreates the actual accident situation. This video is necessary to visually communicate the accident situation to the driver assistance system.
[0692] Step 3:
[0693] The terminal acquires accident-prone location data provided by the server and uses GPS data as input to determine the vehicle's location in real time. When the current location approaches an accident site, it uses augmented reality technology to display a reconstructed video. This outputs information to alert the driver assistance provider.
[0694] Step 4:
[0695] The device's emotion engine uses camera and voice sensors to analyze the user's emotional state from their facial expressions and voice. This sensor data is used as input to calculate the user's fatigue and stress levels. The analysis results are output and used in later steps to customize alert information.
[0696] Step 5:
[0697] The device optimizes the alert information provided to the driver assistance system based on the analyzed emotional state. Using the emotional analysis results as input, it determines the appropriate frequency of voice guidance and visual displays for the driver assistance system. For example, if fatigue is detected, it conveys only the essential information in a calm voice. The output of this process is an alert message optimized for the driver assistance system.
[0698] Step 6:
[0699] The user receives augmented reality images and audio guidance presented from the device. This allows them to enjoy real-time driving assistance based on past accident information and their own emotional state. As an output, direct visual and auditory alerts are promoted for the user.
[0700] (Application Example 2)
[0701] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0702] There is a need to effectively notify drivers of various dangers they may face while driving, based on past accident information and their real-time emotional state. However, conventional systems have struggled to provide optimal feedback tailored to the emotional state of individual drivers. The present invention aims to provide a system that can reduce accident risk and support safe driving by analyzing the driver's emotional state and adjusting danger notifications and warnings based on this analysis.
[0703] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0704] In this invention, the server includes means for acquiring location information of past dangerous events, means for recreating past events based on said location information using a video generation method, and means for providing the recreated events to the operator using an augmented reality method. This makes it possible to monitor the distance between the current location of a moving vehicle while it is in motion and the location of past dangerous events, issue a warning to the operator when it approaches a dangerous location, and adjust the content of the warning by analyzing the operator's emotional state.
[0705] "Location information of past dangerous events" refers to data on the geographical locations where dangerous incidents such as accidents or near misses occurred in the past.
[0706] "Image generation techniques" is a general term for technologies that use computer graphics or live-action footage to visually recreate past events.
[0707] Augmented reality techniques are technologies that overlay additional digital information onto information from the real world and present it to the user.
[0708] The term "operator" often refers to the person using the system while driving, i.e., the driver of the vehicle, but in a broader sense, it includes all people who operate the system.
[0709] "Mobile entities" refer to means of moving that change position, such as vehicles and automobiles, and include those with engines.
[0710] "Analyzing emotional state" is the process of analyzing a user's facial expressions, voice, and other physiological responses to infer their mental state at that moment.
[0711] "Adjusting warning content" means optimizing the method and amount of information presented in a warning, according to the user's emotions and circumstances.
[0712] This system acquires information on past dangerous events and provides effective warnings to operators based on their emotional state. Its main components are a server, terminals, and users.
[0713] The server collects location information of past hazardous events from traffic management agencies and other sources and stores it in a database. Based on this information, it creates videos that recreate past events using video generation techniques.
[0714] The terminal uses information transmitted from the server to track the vehicle's current location using a GPS module. When the vehicle approaches a dangerous area, it displays a reconstructed image to the operator using augmented reality techniques and issues a warning. The terminal also has an emotion engine that analyzes the user's emotions in real time via its built-in camera and voice sensors, and can adjust the warning content according to the user's state.
[0715] For example, if the device analyzes that the user is feeling fatigued or stressed, it will reduce the amount of warning information and soften the voice guidance. Conversely, if it determines that concentration is needed, it will increase the number of times the reenactment video is displayed and employ techniques to maximize attention.
[0716] An example of a prompt in a generative AI model is, "Please suggest the most appropriate real-time notification method for safe driving assistance based on past accident information and emotional state." By using this prompt, the system can generate and provide more appropriate feedback to the operator.
[0717] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0718] Step 1:
[0719] The server collects location data on past hazardous events from traffic management agencies and other information sources. Based on this input data, the server updates the database of hazardous events, organizing and storing the location information. This process provides the basic data for events that should be reproduced.
[0720] Step 2:
[0721] The server uses collected location data to create videos that recreate past dangerous events using a video generation method. This video generation process uses location data as input, performs 3D modeling and simulation, and outputs visually recreated videos. This operation generates content that can be used as a warning to the operator.
[0722] Step 3:
[0723] The device acquires location data and reconstructed video footage transmitted from the server. It uses GPS to track the vehicle's current location in real time and compares this location information with the data from the server. Based on this comparison, if the vehicle is approaching a dangerous location, a determination is output.
[0724] Step 4:
[0725] The device uses augmented reality techniques to display recreated images within the operator's field of view. By using cameras and sensors, it overlays digital information onto the real world, providing warnings to the operator. This process creates visual information that blends reality and virtuality.
[0726] Step 5:
[0727] The device uses its built-in camera and voice sensors to analyze the operator's facial expressions and voice, and an emotion engine evaluates their emotional state. It uses video and audio data as input, executes an algorithm to identify the emotional state, and outputs an emotion evaluation. This evaluation is used to adjust warnings.
[0728] Step 6:
[0729] The device adjusts the amount and presentation of warning information according to the user's emotional state. For example, if the user is fatigued, the voice guidance becomes gentler and visual information is reduced. Conversely, if it is determined that concentration is needed, the warning is strengthened. This final warning content is delivered to the user in the most optimal way.
[0730] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0731] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0732] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0733] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0734] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0735] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0736] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0737] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0738] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0739] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0740] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0741] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0742] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0743] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0744] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0745] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0746] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0747] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0748] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0749] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0750] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0751] The following is further disclosed regarding the embodiments described above.
[0752] (Claim 1)
[0753] A means of obtaining location information of past accidents and near misses,
[0754] A means for recreating past accident situations using video generation technology based on location information,
[0755] A means for displaying the recreated accident situation to the driver using augmented reality technology,
[0756] A means for measuring the current location of a vehicle in motion and monitoring the distance between the measured current location and the location of a past accident,
[0757] A system that includes means to alert the driver when approaching a dangerous location.
[0758] (Claim 2)
[0759] The system according to claim 1, characterized in that the recreated accident situation displayed by the augmented reality technology is displayed in a way that is easy for the driver to visually understand while driving.
[0760] (Claim 3)
[0761] The system according to claim 1, characterized in that the driver is alerted in multiple stages, including by voice and visuals.
[0762] "Example 1"
[0763] (Claim 1)
[0764] A means for receiving past accident information and having a database for storing such information,
[0765] A means for analyzing the data and detecting frequent accidents at specific locations,
[0766] A means for generating simulated video based on the analyzed accident information using generation AI technology,
[0767] A means for visually communicating the generated simulated image to the driver using augmented reality technology,
[0768] A means for measuring the current position, monitoring the relative position information of the vehicle's direction of travel and position relative to the accident site, and determining the approach to an unspecified dangerous point,
[0769] A system that includes means of providing drivers with a combination of voice and visual warnings.
[0770] (Claim 2)
[0771] The system according to claim 1, characterized in that the simulated video using the augmented reality technology is presented to the driver in a format that is highly visible and easy to understand.
[0772] (Claim 3)
[0773] The system according to claim 1, characterized in that the warnings to the driver are provided in multiple ways that are progressively strengthened.
[0774] "Application Example 1"
[0775] (Claim 1)
[0776] A means of obtaining information on past accidents and locations where warnings are needed,
[0777] A means for generating a reconstructed video based on the location information,
[0778] A means for displaying the reproduced image using augmented reality technology,
[0779] A means for measuring the current position of an autonomous vehicle while it is in operation and monitoring the distance between the measured current position and the location of a past accident,
[0780] A means of issuing a warning to the vehicle control device or user when approaching a potential hazardous area,
[0781] A display means using a head-mounted display,
[0782] Means of providing information necessary for safe driving visually and audibly,
[0783] A system that includes this.
[0784] (Claim 2)
[0785] The system according to claim 1, characterized in that the information displayed by the head-mounted display is presented in a way that is easily understood intuitively by the driver and the automated driving control system.
[0786] (Claim 3)
[0787] The system according to claim 1, characterized in that the warning includes step-by-step audio instructions and visual displays.
[0788] "Example 2 of combining an emotion engine"
[0789] (Claim 1)
[0790] A means of collecting information on locations where past hazards or near misses occurred, and related information, from information sources.
[0791] A means for recreating past situations using video generation technology based on location information,
[0792] A means of visually presenting the recreated situation to the driver assistance provider using augmented reality technology,
[0793] A means for measuring the vehicle's current position and monitoring the distance between the current position and past incident locations,
[0794] Means for providing a warning when approaching a warning point,
[0795] A means for analyzing the emotional state of the driver assistant using sensors and dynamically adjusting the content presented based on that emotional state,
[0796] A system that includes this.
[0797] (Claim 2)
[0798] The system according to claim 1, characterized in that the recreated situation displayed by the augmented reality technology is provided to the driver assistance provider in a visually easy-to-understand manner, and the presented information is customized based on the emotion analysis results.
[0799] (Claim 3)
[0800] The system according to claim 1, characterized in that warnings to the driver assistance provider are issued in multiple stages, including voice and visual information, based on the analyzed emotional state.
[0801] "Application example 2 of combining emotional engines"
[0802] (Claim 1)
[0803] A means of obtaining location information where past dangerous events occurred,
[0804] A means for reproducing past events using a video generation method based on the location information,
[0805] A means of providing the operator with a reproduced event using augmented reality techniques,
[0806] A means for measuring the current location of a moving object while it is being operated, and monitoring the distance between the measured current location and past locations where danger occurred,
[0807] A means of warning the operator when approaching a dangerous location,
[0808] A system including means for analyzing the emotional state of the operator and adjusting the warning content based on the analysis results.
[0809] (Claim 2)
[0810] The system according to claim 1, characterized in that the reproduced events provided by the augmented reality method are provided in a format that is easily visually understandable to the operator while they are operating the vehicle.
[0811] (Claim 3)
[0812] The system according to claim 1, characterized in that warnings to the operator are provided in multiple stages, including audibly or visually. [Explanation of Symbols]
[0813] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining location information of past accidents and near misses, A means for recreating past accident situations using video generation technology based on location information, A means for displaying the recreated accident situation to the driver using augmented reality technology, A means for measuring the current location of a vehicle in motion and monitoring the distance between the measured current location and the location of a past accident, A system that includes means to alert the driver when approaching a dangerous location.
2. The system according to claim 1, characterized in that the recreated accident situation displayed by the augmented reality technology is displayed in a way that is easy for the driver to visually understand while driving.
3. The system according to claim 1, characterized in that the driver is alerted in multiple stages, including by voice and visuals.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A