system

A wearable device with a generative model analyzes video data to guide visually impaired users safely, addressing navigation challenges by enhancing obstacle and landmark recognition with feedback integration.

JP2026070255APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Existing systems fail to provide visually impaired individuals with sufficient information to navigate complex environments safely and independently, lacking a comprehensive support system that integrates advanced technologies for obstacle and landmark identification.

Method used

A system that uses a user-wearable device to capture video data, transmit it to a central processing unit for analysis with a generative model, and provide guidance via voice or text, with feedback-based model updates for improved accuracy.

Benefits of technology

Enables visually impaired individuals to navigate safely and independently by accurately identifying obstacles and landmarks, providing real-time guidance, and adapting to user feedback for enhanced system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070255000001_ABST
    Figure 2026070255000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for acquiring video data captured by a computer device worn by the user, Means for transmitting video data acquired from the aforementioned computer device to a central processing unit, The central processing unit includes means for analyzing the video data using a generation model and identifying surrounding obstacles and landmarks, A means for generating information to guide the user's movement path based on the aforementioned analysis results, Means for transmitting the generated information to the computer device and presenting it to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern times, the problem is that appropriate support technologies for visually impaired people to move safely and independently are not fully developed. With the conventional methods relying on white canes or guide dogs, it is sometimes difficult to obtain sufficient information, and there is no consistent support system combining new technologies, so the environment in which visually impaired people can move more freely is restricted. In order to solve such problems, it is necessary to provide a system that enables visually impaired people to move safely even in complex environments using advanced technologies.

Means for Solving the Problems

[0005] This invention provides a means for identifying surrounding obstacles and landmarks by acquiring video data captured by a computer device worn by the user, transmitting that data to a central processing unit, and analyzing it using a generative model. Furthermore, it generates information to guide the user's movement path based on the analysis results, transmits that information to the computer device, and presents it to the user via voice or other means, thereby enabling support for visually impaired individuals to move independently. In addition, the invention provides a means for improving the accuracy of the system by receiving feedback from the user and updating the generative model.

[0006] A "computer device" is a device that a user can wear and that has the function of capturing video data and sending and receiving that data.

[0007] "Video data" refers to digital data containing visual information acquired using a computer.

[0008] A "central processing unit" is a computer device that receives data transmitted from a computer and performs analysis using a generative model.

[0009] A "generative model" is an algorithm that identifies obstacles and landmarks from video data and generates the information necessary to assist the user's movement.

[0010] "Analysis" is the process of identifying objects and features within an environment and extracting information based on acquired video data.

[0011] An "obstacle" refers to a physical element that must be avoided along the user's path.

[0012] A "landmark" is a specific geographical or man-made feature used to determine a user's location or guide their travel route.

[0013] "Information for guiding travel routes" refers to information that includes instructions and warnings necessary for users to reach their destination safely and effectively.

[0014] "Feedback" refers to evaluations and information provided to a system based on user experiences and behaviors, and is used to improve the system. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the language used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The system of the present invention is specifically implemented using a user-wearable computing device, namely smart glasses or a smartphone. This device is equipped with a built-in camera to capture the user's surroundings in real time. When the user begins to move, the device acquires video data through its camera. This video data is compressed and encrypted as appropriate with the user's consent and then transmitted to a server via the internet.

[0037] The server analyzes the received video data and uses a generative model to identify obstacles and landmarks. This generative model incorporates the latest machine learning algorithms, enabling highly accurate analysis along with appropriate location information. Based on the analysis results, the server generates guidance messages for the user, including optimal route information and points of caution to ensure safe movement. This information is sent to the user's device as audio guidance or text.

[0038] Within the device, received information is converted into speech via a speech synthesis engine, and guidance is provided to the user in real time. This allows the user to accurately perceive their surroundings and independently avoid obstacles and reach their destination.

[0039] As a concrete example, consider a scenario where a user moves from a train platform to an exit. The terminal continuously sends camera footage to the server, and guidance based on the latest analysis results is provided each time. For example, a voice announcement such as "Please be careful of the next step" allows the user to safely cross the step and exit the station smoothly.

[0040] Furthermore, through user feedback, the server dynamically updates the generative model, improving the accuracy of the guides it provides. This continuous model optimization allows the system to adapt to users and continue to deliver a more efficient and personalized travel experience.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device uses its built-in camera to acquire video data of the user's surroundings in real time. This data includes information about the user's current environment.

[0044] Step 2:

[0045] The device compresses the acquired video data and sends it to the server while ensuring data security using encryption protocols. This reduces the risk of unauthorized access to the data.

[0046] Step 3:

[0047] The server receives video data transmitted from the terminal and begins analysis using a generative model. The model uses machine learning algorithms to identify obstacles and landmarks in the video and, based on this, determines the user's current location.

[0048] Step 4:

[0049] Based on the analysis results, the server generates guidance information to help users move safely and effectively. This information includes instructions for avoiding obstacles and guiding users using landmarks.

[0050] Step 5:

[0051] The server sends the generated guidance information to the terminal. This transmission is also carried out using an ideal protocol that ensures security, and efforts are made to maintain real-time performance.

[0052] Step 6:

[0053] The terminal converts the received guidance information into audio data and provides it to the user using a speech synthesis engine. This allows the user to understand their surroundings through the audio guidance and continue moving safely.

[0054] Step 7:

[0055] Users follow the instructions and enter feedback about their experience and actions into a terminal. This feedback is then sent to the server.

[0056] Step 8:

[0057] The server analyzes the received feedback and uses it to improve the accuracy of the generative model. This increases the overall guidance accuracy of the system, enabling safer and more effective navigation support for users.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] Existing mobility assistance systems have a problem in that they struggle to accurately recognize obstacles and landmarks in the user's surroundings and provide appropriate information to support safe and efficient movement. Furthermore, there is a challenge in that they cannot improve the system's accuracy by utilizing user feedback.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes means for acquiring image information captured by an information processing device wearable by the user; means for transferring the image information acquired from the information processing device to a central processing device; means for analyzing the image information using a generative model and detecting surrounding obstacles and landmarks in the central processing device; means for generating information to instruct the user's movement path based on the analysis results; means for transmitting the generated information to the information processing device and presenting it to the user visually or audibly; and means for acquiring information from the user and dynamically updating the generative model based on the acquired information. This makes it possible for the user to receive safe and optimal route guidance in real time and to provide feedback to improve the accuracy of the system.

[0063] A "user" refers to an individual who wears the system and utilizes mobility assistance services.

[0064] An "information processing device" refers to a device that a user can wear, which acquires image information of the environment and transmits it to a server.

[0065] "Image information" refers to visual data of the user's surroundings acquired by an information processing device, and the material used for analysis.

[0066] A "central processing unit" refers to a computer device that analyzes received image information and generates necessary support information.

[0067] A "generative model" refers to a predictive algorithm used to analyze image information and identify obstacles and landmarks.

[0068] An "obstacle" refers to any object that physically obstructs the user's path.

[0069] A "landmark" refers to a location or object that serves as a reference point to assist the user's movement.

[0070] "Travel route" refers to the appropriate direction of travel for a user to reach their destination.

[0071] "Information for giving instructions" refers to directions provided to the user based on the results of analysis using a generative model.

[0072] "Presenting visually or aurally" means communicating generated information to the user through a display or audio output.

[0073] "Feedback" refers to the user's reaction or opinion to the responses and guidance provided by the system.

[0074] "Dynamic updating" refers to adaptively adjusting the parameters and outputs of a generative model based on user feedback.

[0075] This invention is implemented using a wearable information processing device. Specifically, the user can perceive their surroundings in real time through a device such as smart glasses or a smartphone. This information processing device has a built-in high-resolution camera and is capable of continuously acquiring image information of the environment within the user's field of view.

[0076] The terminal processes the acquired image information, compresses the video for efficient data transfer, and protects the data using encryption algorithms such as AES-256 to ensure security. This compressed and encrypted data is transmitted to the server via a wireless communication network. Wi-Fi and mobile data communication are used, and the system is designed to exchange data with low latency even when the user is on the move.

[0077] The server uses a generative AI model to analyze the received image information. This analysis utilizes the latest machine learning platforms, such as TENSORFLOW® and PyTorch. The generative model accurately identifies obstacles and landmarks around the user and calculates the optimal movement path for the user. Based on the analysis results, the server generates audio and text guidance information to help the user move safely and efficiently.

[0078] The generated guide information is sent back to the terminal and presented to the user in real time using a speech synthesis engine. Utilizing natural language processing technology, the guidance is provided in an intuitive and easy-to-understand format. The user can then move safely by following this information.

[0079] As a concrete example, consider a scenario where a user is searching for an exit within a large public facility. This system supports smooth movement by quickly identifying potential obstacles the user might encounter and providing specific instructions such as, "You are approaching a wall; move to the left to avoid it."

[0080] An example of a prompt message would be, "Identify potential obstacles in the image and guide the user to the optimal detour." This allows the generative AI model to quickly generate and deliver the information the user needs.

[0081] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0082] Step 1:

[0083] The device activates its built-in camera and captures the user's surroundings in real time. The input is video frames from the camera, which are acquired and processed as continuous data. Specifically, high-resolution image data is acquired and prepared to be passed to the next processing step.

[0084] Step 2:

[0085] The terminal compresses and encrypts the acquired video data. The input is raw video frame data. This data is compressed using compression techniques such as H.264, and then encrypted using algorithms such as AES-256. As a result, video data is output that has been converted into a secure format while reducing the amount of data.

[0086] Step 3:

[0087] The terminal sends compressed and encrypted video data to the server. The input is compressed and encrypted data, and the output is transmitted over the internet via Wi-Fi or a mobile network. The goal of this step is to transfer the data efficiently and with low latency.

[0088] Step 4:

[0089] The server decrypts and decompresses the received data. The input is encrypted and compressed video data. By decrypting the data using AES-256 or similar encryption methods, and then decompressing the compressed data, the server obtains analyzable raw data.

[0090] Step 5:

[0091] The server analyzes the decoded data using a generative AI model. The input is decoded video data. The generative AI model (using TensorFlow or PyTorch) analyzes the data to identify obstacles and landmarks. As output, location information and type information are obtained as analysis results.

[0092] Step 6:

[0093] The server generates guidance messages based on the analysis results. The input is the analysis results from the AI ​​model. Using natural language processing technology, it generates guidance messages that include necessary travel routes and points to note for the user, and provides the information as output in the form of voice or text.

[0094] Step 7:

[0095] The server sends the generated guidance message to the terminal. The input consists of the generated audio and text data. This data is then forwarded back to the terminal and output so that it can be received by the user's device.

[0096] Step 8:

[0097] The terminal outputs received guidance messages through a speech synthesis engine. The input consists of voice and text data received from the server. Speech synthesis converts this data into voice messages in real time, directly providing instructions to the user audibly.

[0098] Step 9:

[0099] Users provide feedback based on their experiences while actually moving around. The input consists of their real-world travel experiences and impressions. This data is sent to a server via the terminal and used as output for further optimization of the AI ​​model and system improvement.

[0100] (Application Example 1)

[0101] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0102] There is a need for technology that enables autonomous vehicles to move safely and efficiently in complex environments. Conventional technologies have problems such as difficulty in accurately understanding the surrounding environment using only the sensors on the vehicle, and the inability to provide the necessary information to improve safety in real time. Therefore, there is a need to develop a dynamic guidance system that allows the vehicle to accurately perceive its surroundings and avoid obstacles.

[0103] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0104] In this invention, the server includes means for acquiring video data captured by a computer device worn by the user, means for transmitting the video data acquired from the computer device to a central processing unit, and means for analyzing the video data using a generative model to identify surrounding obstacles and landmarks. This makes it possible to accurately recognize the surrounding conditions of a transported object and provide an appropriate travel path.

[0105] A "user-worn computing device" is an information processing device that functions in a form that can be worn by an individual, and has the ability to acquire video data and communicate with other devices.

[0106] "Means of acquiring video data" refers to a system that collects visual information as digital data using a camera or other imaging device.

[0107] A "central processing unit" is a primary data processing unit that analyzes received data and generates information for decision-making and control.

[0108] A "generative model" is a computational model designed to analyze and predict data using algorithms based on machine learning or artificial intelligence.

[0109] "Transportation equipment" refers to vehicles and mechanical devices used as means of moving people or goods.

[0110] "Information for guiding travel routes" refers to data intended to guide travelers, including instructions and precautions regarding the optimal route to their destination.

[0111] "User feedback" refers to opinions and evaluations regarding the experience and performance of the system, provided by individuals or organizations that use the system.

[0112] The system that realizes this invention basically consists of a computer device worn by the user (such as smart glasses or a smartphone) and a server. The user's computer device is equipped with a camera that acquires real-time video data of the surroundings while the transported object is moving. The acquired video data is compressed and encrypted via the internet and transmitted to the server.

[0113] Upon receiving video data, the server uses a generative AI model to perform analysis. This analysis employs state-of-the-art machine learning algorithms, enabling high-precision identification of surrounding obstacles and landmarks. Based on the analysis results, the server generates information to guide the user's transported vehicle safely and efficiently along its route. This includes information on how to avoid obstacles and points of caution along the route.

[0114] The generated guidance information is transmitted to the user's computer and presented as audio guides or text messages. This enables the autonomous movement of the transported object. Furthermore, the server receives feedback from the user, dynamically updating the generated AI model and continuously improving the accuracy of the guidance.

[0115] As a concrete example, consider an autonomous driving scenario in an urban area. When a moving vehicle approaches an intersection, the server uses video analysis to recognize traffic lights and pedestrians and provides voice guidance such as, "The light has turned red, please stop." This ensures safe driving.

[0116] An example of a prompt message might be, "Analyze the current road conditions and provide a safe route." By entering this prompt message, the user can begin the system's analysis and guidance process.

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] The device uses its built-in camera to capture the user's surroundings as video data in real time. The captured video data is compressed and encrypted and sent to a server via the internet. In this step, the input is visual information of the surroundings, and the output is compressed and encrypted video data. The device uses a data processing algorithm to convert the video into an appropriate format.

[0120] Step 2:

[0121] The server receives video data transmitted from the terminal and performs decompression and decryption of the data. It receives compressed and encrypted data as input and obtains restored video data as output. The server decrypts the data using security protocols and prepares it in an analyzable format.

[0122] Step 3:

[0123] The server uses a generative AI model to analyze the received video data and identify obstacles and landmarks. It takes the reconstructed video data as input and generates analyzed environmental information as output. In this process, machine learning algorithms utilize image recognition technology to extract specific features.

[0124] Step 4:

[0125] The server generates information to guide the transporter's movement path based on the analysis results. It uses the analyzed environmental information as input and generates instructions, including route guidance and points of caution, as output. Based on the environmental information, the server calculates a safe and efficient route and constructs a guidance message.

[0126] Step 5:

[0127] The server sends the generated guidance information to the terminal. It takes guidance information as input and distributes data to the terminal as output. The server provides information in real time using a communication protocol.

[0128] Step 6:

[0129] The terminal presents the received guidance information to the user as an audio guide or text. It receives guidance information transmitted as input and displays it as audio or text as output. The terminal uses a speech synthesis engine to generate clear audio guides and deliver instructions to the user.

[0130] Step 7:

[0131] The user provides feedback on the operation and environmental changes during the movement of the transported object. The system receives user feedback as input and obtains updated data for the generated AI model as output. This feedback is used by the server as crucial data for further system optimization.

[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0133] This invention relates to a system that incorporates an emotion engine to further personalize user assistance. This system is implemented using a user-wearable computing device, such as smart glasses or a smartphone. These devices are equipped with a built-in camera and emotion recognition sensor, which can analyze the user's emotions from their facial expressions and voice.

[0134] The device acquires real-time video data from the user and transmits it to the central processing unit. During the transmission of video data, the data is appropriately encrypted to protect the user's privacy. The device also analyzes emotional data using an emotion engine and transmits that information to the server.

[0135] The server analyzes video data transmitted from the terminal using a generative model to identify obstacles and landmarks in the environment. Furthermore, it uses data from the emotion engine to understand the user's emotional state and generate appropriate navigation information. For example, if the user is feeling stressed, the server will provide navigation information in a calmer tone and avoid more direct instructions.

[0136] The generated guidance information is transmitted to the terminal via a speech synthesis engine. The terminal presents the transmitted information to the user verbally and visually, providing real-time guidance. This personalized guidance allows the user to have a better travel experience while ensuring security.

[0137] As a concrete example, consider a situation where a user is walking through a busy area and a car suddenly approaches. In this case, if the emotion engine detects the user's anxiety, the server will send swift and calm instructions to inform the user of how to safely avoid the car. In this way, instructions that take the user's emotional state into consideration significantly improve the user's sense of security and safety.

[0138] Feedback is also collected based on emotional states, and the server uses this to optimize its generative model and emotion engine. As a result, the system can continuously adapt to the user's changing emotions and provide optimal guidance.

[0139] The following describes the processing flow.

[0140] Step 1:

[0141] The device uses a built-in camera to acquire video data of the user's surroundings in real time. Simultaneously, it uses an emotion recognition sensor to collect emotional data from the user's facial expressions and voice.

[0142] Step 2:

[0143] The device encrypts the acquired video and emotional data to protect privacy and transmits it to the server using a secure communication protocol.

[0144] Step 3:

[0145] The server activates a generative model to analyze the received video data and identify obstacles and landmarks in the environment. Machine learning algorithms are used in this process.

[0146] Step 4:

[0147] The server uses an emotion engine to analyze the received emotion data and evaluate the user's emotional state. This determines the user's current emotional state.

[0148] Step 5:

[0149] The server integrates the analysis results and generates navigation information tailored to the user's current emotional state. For example, if the user is stressed, the guidance will be delivered in a gentle tone, and directions will be simplified in crowded areas.

[0150] Step 6:

[0151] The server sends the generated guidance information to the terminal. This information is then sent to the speech synthesis engine and provided to the user as an easy-to-understand voice message.

[0152] Step 7:

[0153] The terminal provides users with guidance information via speech synthesis, offering real-time navigation assistance. This enables users to travel to their destination safely and effectively.

[0154] Step 8:

[0155] Users follow the provided instructions and enter feedback about their experience and feelings into a terminal. This feedback is later sent to the server.

[0156] Step 9:

[0157] The server receives user feedback and uses it as data to optimize the generative model and sentiment engine. This allows the system to adapt more closely to the user and provide more accurate support.

[0158] (Example 2)

[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0160] Conventional mobility assistance systems struggle to provide guidance that takes into account the individual emotional state of the user, resulting in many situations where users experience anxiety and stress. To solve this problem and improve the user's mobility experience, it is necessary to analyze the user's emotions in real time and provide personalized guidance based on that analysis.

[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0162] In this invention, the server includes means for analyzing observational data using a generative AI model to identify elements within the environment, means for analyzing the user's emotional state, and means for generating and providing navigation guidance information based on this information. This makes it possible to provide guidance tailored to the user's individual emotional state in real time.

[0163] A "user" is an entity that wears a computer device and receives guidance information provided by the system.

[0164] A "terminal device" is a computer device worn by the user to acquire observation data and transmit it to the central control system.

[0165] "Observational data" refers to real-time data including the user's surrounding environment, facial expressions, and voice.

[0166] A "server device" is a central computing device that analyzes data sent from terminal devices and generates guidance information for users.

[0167] "Emotional analysis" is the process of determining a user's emotional state from their facial expressions and tone of voice.

[0168] A "generative AI model" is an artificial intelligence model that analyzes observational data and recognizes the surrounding environment.

[0169] "Travel guidance" refers to information provided to help users travel safely and comfortably.

[0170] "Feedback" refers to information obtained from users regarding their usage and emotional responses.

[0171] This invention relates to an individualized navigation system comprising a user-wearable terminal device and a server device. The terminal device is a portable information device such as smart glasses or a smartphone, and acquires user observation data using a camera and microphone. This observation data includes real-time information such as the user's facial expressions and tone of voice. This data is processed by an emotion engine to analyze the user's emotions.

[0172] The terminal device transmits the acquired observation data to the server device. The server device is a powerful computing system that analyzes the data using a generative AI model. This analysis identifies surrounding obstacles and landmarks and generates information to determine a safe travel path for the user.

[0173] The emotion engine analyzes the user's emotional state in real time and influences the generated guidance information. For example, if the user is feeling anxious, guidance will be provided in a calmer tone. Furthermore, the feedback function allows the generating AI model and emotion engine to continuously optimize based on information gained from the user's experience.

[0174] As a concrete example, consider a scenario where a user is walking through a crowded city and encounters an unexpected obstacle. In this case, the server quickly suggests a safe alternative route, and the device provides this information verbally through a speech synthesis engine. Simultaneously, visual guidelines can be displayed on the AR display.

[0175] An example of a prompt message is as follows:

[0176] "Analyze the user's current emotional state and generate a safe and reassuring travel route based on that analysis. This route should include information on landmarks and relaxation points to reduce the user's stress. Additionally, provide guidance with simple instructions to help the user cope in emergencies."

[0177] In this way, this system enhances user confidence and safety by providing dynamic and personalized guidance that responds to the user's unique state and surrounding circumstances.

[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0179] Step 1:

[0180] The device acquires observational data from the camera and microphone using smart glasses or a smartphone. Input includes the user's facial expressions and voice data. This data is captured in real time and forms the basis for capturing the user's dynamic emotional changes. This observational data is then encrypted and sent to a server.

[0181] Step 2:

[0182] The device sends observational data to its built-in emotion engine, which then analyzes the user's emotional state. Here, observational data is used as input, and the data is processed using machine learning algorithms. The emotion engine outputs an emotion label based on facial features and voice tone, and sends this information to the server.

[0183] Step 3:

[0184] The server receives observational data transmitted from the terminal and analyzes it using a generative AI model. The input includes observational data and emotion labels. Through this analysis, the server processes the video data and identifies obstacles and landmarks in the surrounding environment. The output provides basic information for a safe travel route for the user.

[0185] Step 4:

[0186] The server generates user-directed navigation information based on the results of emotion analysis and environmental analysis. Input includes the user's emotional state and environmental details. A generative AI model is used to create guidance tailored to the user's emotions using prompts. These prompts may include instructions such as, "Please provide guidance in a slow voice to help the user relax." The final output is personalized guidance information, including voice guidance and visual guidelines.

[0187] Step 5:

[0188] The terminal receives guidance information sent from the server and presents it to the user. The input includes guidance information, which is then delivered to the user in natural-sounding voice using a speech synthesis engine. This also includes specific actions such as displaying visual guidelines via an AR display. Users can navigate with confidence based on the guidance information they receive. Feedback is also collected simultaneously and used for optimization in future applications.

[0189] (Application Example 2)

[0190] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0191] Conventional navigation systems struggle to provide personalized assistance based on the user's emotional state. Therefore, ensuring both emotional comfort and safety for passengers is a challenge, especially in autonomous vehicles. A travel experience that doesn't reflect the user's emotional state can cause unnecessary stress and anxiety.

[0192] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0193] In this invention, the server includes means for acquiring visual data obtained by a computing device worn by the user, means for transmitting the visual data from the computing device to a processing device, and means for analyzing the visual data using a generative model in the processing device to identify surrounding objects and reference points. This makes it possible to adjust the navigation guidance information based on the user's emotional state. Specifically, by flexibly adjusting the navigation route and output information according to the user's emotions, it is possible to provide passengers with a safer and more comfortable travel experience.

[0194] A "user" refers to an individual or group that uses this system, and is the subject whose emotional state is analyzed.

[0195] A "computational device" refers to a device worn by a user that has the function of acquiring and transmitting visual data.

[0196] "Visual data" refers to image and video information captured by the computing device, and is fundamental data for analyzing the user's emotional state.

[0197] A "processing unit" refers to a central computer that receives visual data transmitted from the computing unit and performs analysis using a generative model.

[0198] A "generative model" is a set of algorithms used in data analysis that perform intelligent processing to identify objects and reference points from visual data.

[0199] An "object" is a specific object that exists in the user's surrounding environment and is identified by the generative model.

[0200] A "reference point" refers to a landmark or marker used to determine the user's location in their surrounding environment, and is utilized in generating navigation information.

[0201] "Navigation guidance information" refers to guidelines and advice generated to assist users in their travels, and is adjusted based on the user's emotional state.

[0202] "Emotional state" refers to the emotional reactions and situations exhibited by the user, and is captured through emotion recognition sensors and visual data analysis.

[0203] To implement this invention, the user wears smart glasses or a smartphone as a computing device, thereby acquiring visual data. The device incorporates a camera and emotion recognition sensor to understand the user's emotions, and these are used to analyze video and audio data in real time.

[0204] The server receives encrypted visual data transmitted from the terminal and uses a generative model to identify surrounding objects and reference points. Based on this, it optimizes the user's movement path and generates appropriate navigation information. This navigation information is adjusted according to the user's emotional state; for example, if the user is feeling anxious, the server will explain the situation in a calm tone to support a more reassuring journey.

[0205] For example, if a passenger in an autonomous vehicle shows signs of anxiety, the vehicle could choose a more comfortable route and play calming music. In this way, support that takes user emotions into consideration can significantly improve the quality of the travel experience.

[0206] An example of a prompt would be, "What environmental conditions should you set based on the user's emotional state?" This serves as a guide for the generative AI model to learn the optimal guidance based on emotional states.

[0207] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0208] Step 1:

[0209] The device collects the user's visual and audio data. The device's built-in camera and emotion recognition sensor are activated, acquiring the user's facial expressions and audio information in real time. Through this process, the device obtains basic data as input to determine the user's emotional state.

[0210] Step 2:

[0211] The device securely and efficiently transmits the collected visual and audio data to the server. The data is encrypted and sent to the server via wireless communication. This allows the server to receive the data necessary for analysis while protecting privacy.

[0212] Step 3:

[0213] The server uses a generative AI model to analyze the received visual data and identify surrounding objects and reference points. It extracts features from the input data and recognizes obstacles and landmarks based on the generative model. As output, it provides detailed information about the user's environment.

[0214] Step 4:

[0215] The server integrates the user's emotional state with surrounding information to generate optimal navigation guidance. Based on the emotion recognition results, it determines the instructions that correspond to the user's emotions. For example, if the user is feeling anxious, it generates a more relaxing route or voice guidance. In this step, the generated guidance information is output.

[0216] Step 5:

[0217] The terminal displays guidance information received from the server to the user. This information is conveyed to the user via audio or visual means. The terminal continuously monitors the user's reactions and provides guidance in real time. This process allows the user to intuitively understand the situation and take the optimal action.

[0218] Step 6:

[0219] The user sends feedback to the server while on the move. By sending changes in emotion and responses to guidance as input to the server, the system's feedback loop is completed. This step allows the server to optimize the entire system, including the generative AI model, to improve the accuracy of future assistance.

[0220] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0221] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0222] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0223] [Second Embodiment]

[0224] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0225] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0226] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0227] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0228] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0229] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0230] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0231] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0232] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0233] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0234] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0235] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0236] The system of the present invention is specifically implemented using a user-wearable computing device, namely smart glasses or a smartphone. This device is equipped with a built-in camera to capture the user's surroundings in real time. When the user begins to move, the device acquires video data through its camera. This video data is compressed and encrypted as appropriate with the user's consent and then transmitted to a server via the internet.

[0237] The server analyzes the received video data and uses a generative model to identify obstacles and landmarks. This generative model incorporates the latest machine learning algorithms, enabling highly accurate analysis along with appropriate location information. Based on the analysis results, the server generates guidance messages for the user, including optimal route information and points of caution to ensure safe movement. This information is sent to the user's device as audio guidance or text.

[0238] Within the device, received information is converted into speech via a speech synthesis engine, and guidance is provided to the user in real time. This allows the user to accurately perceive their surroundings and independently avoid obstacles and reach their destination.

[0239] As a concrete example, consider a scenario where a user moves from a train platform to an exit. The terminal continuously sends camera footage to the server, and guidance based on the latest analysis results is provided each time. For example, a voice announcement such as "Please be careful of the next step" allows the user to safely cross the step and exit the station smoothly.

[0240] Furthermore, through user feedback, the server dynamically updates the generative model, improving the accuracy of the guides it provides. This continuous model optimization allows the system to adapt to users and continue to deliver a more efficient and personalized travel experience.

[0241] The following describes the processing flow.

[0242] Step 1:

[0243] The device uses its built-in camera to acquire video data of the user's surroundings in real time. This data includes information about the user's current environment.

[0244] Step 2:

[0245] The device compresses the acquired video data and sends it to the server while ensuring data security using encryption protocols. This reduces the risk of unauthorized access to the data.

[0246] Step 3:

[0247] The server receives video data transmitted from the terminal and begins analysis using a generative model. The model uses machine learning algorithms to identify obstacles and landmarks in the video and, based on this, determines the user's current location.

[0248] Step 4:

[0249] Based on the analysis results, the server generates guidance information to help users move safely and effectively. This information includes instructions for avoiding obstacles and guiding users using landmarks.

[0250] Step 5:

[0251] The server sends the generated guidance information to the terminal. This transmission is also carried out using an ideal protocol that ensures security, and efforts are made to maintain real-time performance.

[0252] Step 6:

[0253] The terminal converts the received guidance information into audio data and provides it to the user using a speech synthesis engine. This allows the user to understand their surroundings through the audio guidance and continue moving safely.

[0254] Step 7:

[0255] Users follow the instructions and enter feedback about their experience and actions into a terminal. This feedback is then sent to the server.

[0256] Step 8:

[0257] The server analyzes the received feedback and uses it to improve the accuracy of the generative model. This increases the overall guidance accuracy of the system, enabling safer and more effective navigation support for users.

[0258] (Example 1)

[0259] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0260] Existing mobility assistance systems have a problem in that they struggle to accurately recognize obstacles and landmarks in the user's surroundings and provide appropriate information to support safe and efficient movement. Furthermore, there is a challenge in that they cannot improve the system's accuracy by utilizing user feedback.

[0261] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0262] In this invention, the server includes means for acquiring image information captured by an information processing device wearable by the user; means for transferring the image information acquired from the information processing device to a central processing device; means for analyzing the image information using a generative model and detecting surrounding obstacles and landmarks in the central processing device; means for generating information to instruct the user's movement path based on the analysis results; means for transmitting the generated information to the information processing device and presenting it to the user visually or audibly; and means for acquiring information from the user and dynamically updating the generative model based on the acquired information. This makes it possible for the user to receive safe and optimal route guidance in real time and to provide feedback to improve the accuracy of the system.

[0263] A "user" refers to an individual who wears the system and utilizes mobility assistance services.

[0264] An "information processing device" refers to a device that a user can wear, which acquires image information of the environment and transmits it to a server.

[0265] "Image information" refers to visual data of the user's surroundings acquired by an information processing device, and the material used for analysis.

[0266] A "central processing unit" refers to a computer device that analyzes received image information and generates necessary support information.

[0267] A "generative model" refers to a predictive algorithm used to analyze image information and identify obstacles and landmarks.

[0268] An "obstacle" refers to any object that physically obstructs the user's path.

[0269] A "landmark" refers to a location or object that serves as a reference point to assist the user's movement.

[0270] "Travel route" refers to the appropriate direction of travel for a user to reach their destination.

[0271] "Information for giving instructions" refers to directions provided to the user based on the results of analysis using a generative model.

[0272] "Presenting visually or aurally" means communicating generated information to the user through a display or audio output.

[0273] "Feedback" refers to the user's reaction or opinion to the responses and guidance provided by the system.

[0274] "Dynamic updating" refers to adaptively adjusting the parameters and outputs of a generative model based on user feedback.

[0275] This invention is implemented using a wearable information processing device. Specifically, the user can perceive their surroundings in real time through a device such as smart glasses or a smartphone. This information processing device has a built-in high-resolution camera and is capable of continuously acquiring image information of the environment within the user's field of view.

[0276] The terminal processes the acquired image information, compresses the video for efficient data transfer, and protects the data using encryption algorithms such as AES-256 to ensure security. This compressed and encrypted data is transmitted to the server via a wireless communication network. Wi-Fi and mobile data communication are used, and the system is designed to exchange data with low latency even when the user is on the move.

[0277] The server uses a generative AI model to analyze the received image information. This analysis utilizes the latest machine learning platforms, TensorFlow and PyTorch. The generative model accurately identifies obstacles and landmarks in the user's surroundings and calculates the optimal movement path for the user. Based on the analysis results, the server generates audio and text guidance information to help the user move safely and efficiently.

[0278] The generated guide information is sent back to the terminal and presented to the user in real time using a speech synthesis engine. Utilizing natural language processing technology, the guidance is provided in an intuitive and easy-to-understand format. The user can then move safely by following this information.

[0279] As a concrete example, consider a scenario where a user is searching for an exit within a large public facility. This system supports smooth movement by quickly identifying potential obstacles the user might encounter and providing specific instructions such as, "You are approaching a wall; move to the left to avoid it."

[0280] Examples of prompt texts include instructions such as "Identify potential obstacles in the image and guide the optimal detour route." As a result, the generative AI model can quickly generate and deliver the information required by the user.

[0281] The flow of the specific process in Example 1 will be described using FIG. 11.

[0282] Step 1:

[0283] The terminal activates the built-in camera and captures the user's surrounding environment in real time. The input is the video frame from the camera, which is acquired and processed as continuous data. Specifically, high-resolution image data is acquired and prepared to be passed to the next processing step.

[0284] Step 2:

[0285] The terminal compresses and encrypts the acquired video data. The input is the raw video frame data. This data is compressed using a compression technology such as H.264 and then encrypted using an algorithm such as AES-256. As a result, video data that has been converted into a secure format while reducing the data volume is output.

[0286] Step 3:

[0287] The terminal sends the compressed and encrypted video data to the server. The input is the compressed and encrypted data, which is output via the Internet through Wi-Fi or a mobile network. The goal in this step is to transfer the data efficiently and with low latency.

[0288] Step 4:

[0289] The server decrypts and decompresses the received data. The input is the encrypted and compressed video data. By decrypting with AES-256 etc. and then decompressing the compressed data, raw data that can be analyzed is obtained.

[0290] Step 5:

[0291] The server analyzes the decoded data using a generative AI model. The input is decoded video data. The generative AI model (using TensorFlow or PyTorch) analyzes the data to identify obstacles and landmarks. As output, location information and type information are obtained as analysis results.

[0292] Step 6:

[0293] The server generates guidance messages based on the analysis results. The input is the analysis results from the AI ​​model. Using natural language processing technology, it generates guidance messages that include necessary travel routes and points to note for the user, and provides the information as output in the form of voice or text.

[0294] Step 7:

[0295] The server sends the generated guidance message to the terminal. The input consists of the generated audio and text data. This data is then forwarded back to the terminal and output so that it can be received by the user's device.

[0296] Step 8:

[0297] The terminal outputs received guidance messages through a speech synthesis engine. The input consists of voice and text data received from the server. Speech synthesis converts this data into voice messages in real time, directly providing instructions to the user audibly.

[0298] Step 9:

[0299] Users provide feedback based on their experiences while actually moving around. The input consists of their real-world travel experiences and impressions. This data is sent to a server via the terminal and used as output for further optimization of the AI ​​model and system improvement.

[0300] (Application Example 1)

[0301] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0302] There is a need for a technology that enables a carrier equipped with an autonomous driving system to move safely and efficiently in a complex environment. In the conventional technology, it may be difficult to accurately grasp the surrounding situation only with the sensors of the carrier, and there is also a problem that the information necessary to improve safety cannot be provided in real time. Therefore, the development of a dynamic guidance system that enables the carrier to accurately recognize the surrounding situation and avoid obstacles is required.

[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0304] In this invention, the server includes means for acquiring video data captured by a computer device worn by a user, means for transmitting the video data acquired from the computer device to a central processing unit, and means for analyzing the video data using a generation model to identify surrounding obstacles and landmarks. As a result, it becomes possible to accurately recognize the surrounding situation of the carrier and provide an appropriate movement route.

[0305] The "computer device worn by a user" is an information processing device that functions in a form worn by an individual and has the ability to acquire video data and communicate with other devices.

[0306] The "means for acquiring video data" is a mechanism for collecting visual information as digital data using a photographing device such as a camera.

[0307] The "central processing unit" is a main data processing unit for analyzing the received data and generating information for decision-making and control.

[0308] A "generative model" is a computational model designed to analyze and predict data using algorithms based on machine learning or artificial intelligence.

[0309] "Transportation equipment" refers to vehicles and mechanical devices used as means of moving people or goods.

[0310] "Information for guiding travel routes" refers to data intended to guide travelers, including instructions and precautions regarding the optimal route to their destination.

[0311] "User feedback" refers to opinions and evaluations regarding the experience and performance of the system, provided by individuals or organizations that use the system.

[0312] The system that realizes this invention basically consists of a computer device worn by the user (such as smart glasses or a smartphone) and a server. The user's computer device is equipped with a camera that acquires real-time video data of the surroundings while the transported object is moving. The acquired video data is compressed and encrypted via the internet and transmitted to the server.

[0313] Upon receiving video data, the server uses a generative AI model to perform analysis. This analysis employs state-of-the-art machine learning algorithms, enabling high-precision identification of surrounding obstacles and landmarks. Based on the analysis results, the server generates information to guide the user's transported vehicle safely and efficiently along its route. This includes information on how to avoid obstacles and points of caution along the route.

[0314] The generated guidance information is transmitted to the user's computer and presented as audio guides or text messages. This enables the autonomous movement of the transported object. Furthermore, the server receives feedback from the user, dynamically updating the generated AI model and continuously improving the accuracy of the guidance.

[0315] As a concrete example, consider an autonomous driving scenario in an urban area. When a moving vehicle approaches an intersection, the server uses video analysis to recognize traffic lights and pedestrians and provides voice guidance such as, "The light has turned red, please stop." This ensures safe driving.

[0316] An example of a prompt message might be, "Analyze the current road conditions and provide a safe route." By entering this prompt message, the user can begin the system's analysis and guidance process.

[0317] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0318] Step 1:

[0319] The device uses its built-in camera to capture the user's surroundings as video data in real time. The captured video data is compressed and encrypted and sent to a server via the internet. In this step, the input is visual information of the surroundings, and the output is compressed and encrypted video data. The device uses a data processing algorithm to convert the video into an appropriate format.

[0320] Step 2:

[0321] The server receives video data transmitted from the terminal and performs decompression and decryption of the data. It receives compressed and encrypted data as input and obtains restored video data as output. The server decrypts the data using security protocols and prepares it in an analyzable format.

[0322] Step 3:

[0323] The server uses a generative AI model to analyze the received video data and identify obstacles and landmarks. It takes the reconstructed video data as input and generates analyzed environmental information as output. In this process, machine learning algorithms utilize image recognition technology to extract specific features.

[0324] Step 4:

[0325] The server generates information to guide the transporter's movement path based on the analysis results. It uses the analyzed environmental information as input and generates instructions, including route guidance and points of caution, as output. Based on the environmental information, the server calculates a safe and efficient route and constructs a guidance message.

[0326] Step 5:

[0327] The server sends the generated guidance information to the terminal. It takes guidance information as input and distributes data to the terminal as output. The server provides information in real time using a communication protocol.

[0328] Step 6:

[0329] The terminal presents the received guidance information to the user as an audio guide or text. It receives guidance information transmitted as input and displays it as audio or text as output. The terminal uses a speech synthesis engine to generate clear audio guides and deliver instructions to the user.

[0330] Step 7:

[0331] The user provides feedback on the operation and environmental changes during the movement of the transported object. The system receives user feedback as input and obtains updated data for the generated AI model as output. This feedback is used by the server as crucial data for further system optimization.

[0332] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0333] This invention relates to a system that incorporates an emotion engine to further personalize user assistance. This system is implemented using a user-wearable computing device, such as smart glasses or a smartphone. These devices are equipped with a built-in camera and emotion recognition sensor, which can analyze the user's emotions from their facial expressions and voice.

[0334] The device acquires real-time video data from the user and transmits it to the central processing unit. During the transmission of video data, the data is appropriately encrypted to protect the user's privacy. The device also analyzes emotional data using an emotion engine and transmits that information to the server.

[0335] The server analyzes video data transmitted from the terminal using a generative model to identify obstacles and landmarks in the environment. Furthermore, it uses data from the emotion engine to understand the user's emotional state and generate appropriate navigation information. For example, if the user is feeling stressed, the server will provide navigation information in a calmer tone and avoid more direct instructions.

[0336] The generated guidance information is transmitted to the terminal via a speech synthesis engine. The terminal presents the transmitted information to the user verbally and visually, providing real-time guidance. This personalized guidance allows the user to have a better travel experience while ensuring security.

[0337] As a concrete example, consider a situation where a user is walking through a busy area and a car suddenly approaches. In this case, if the emotion engine detects the user's anxiety, the server will send swift and calm instructions to inform the user of how to safely avoid the car. In this way, instructions that take the user's emotional state into consideration significantly improve the user's sense of security and safety.

[0338] Feedback is also collected based on emotional states, and the server uses this to optimize its generative model and emotion engine. As a result, the system can continuously adapt to the user's changing emotions and provide optimal guidance.

[0339] The following describes the processing flow.

[0340] Step 1:

[0341] The device uses a built-in camera to acquire video data of the user's surroundings in real time. Simultaneously, it uses an emotion recognition sensor to collect emotional data from the user's facial expressions and voice.

[0342] Step 2:

[0343] The device encrypts the acquired video and emotional data to protect privacy and transmits it to the server using a secure communication protocol.

[0344] Step 3:

[0345] The server activates a generative model to analyze the received video data and identify obstacles and landmarks in the environment. Machine learning algorithms are used in this process.

[0346] Step 4:

[0347] The server uses an emotion engine to analyze the received emotion data and evaluate the user's emotional state. This determines the user's current emotional state.

[0348] Step 5:

[0349] The server integrates the analysis results and generates navigation information tailored to the user's current emotional state. For example, if the user is stressed, the guidance will be delivered in a gentle tone, and directions will be simplified in crowded areas.

[0350] Step 6:

[0351] The server sends the generated guidance information to the terminal. This information is then sent to the speech synthesis engine and provided to the user as an easy-to-understand voice message.

[0352] Step 7:

[0353] The terminal provides users with guidance information via speech synthesis, offering real-time navigation assistance. This enables users to travel to their destination safely and effectively.

[0354] Step 8:

[0355] Users follow the provided instructions and enter feedback about their experience and feelings into a terminal. This feedback is later sent to the server.

[0356] Step 9:

[0357] The server receives user feedback and uses it as data to optimize the generative model and sentiment engine. This allows the system to adapt more closely to the user and provide more accurate support.

[0358] (Example 2)

[0359] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0360] Conventional mobility assistance systems struggle to provide guidance that takes into account the individual emotional state of the user, resulting in many situations where users experience anxiety and stress. To solve this problem and improve the user's mobility experience, it is necessary to analyze the user's emotions in real time and provide personalized guidance based on that analysis.

[0361] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0362] In this invention, the server includes means for analyzing observational data using a generative AI model to identify elements within the environment, means for analyzing the user's emotional state, and means for generating and providing navigation guidance information based on this information. This makes it possible to provide guidance tailored to the user's individual emotional state in real time.

[0363] A "user" is an entity that wears a computer device and receives guidance information provided by the system.

[0364] A "terminal device" is a computer device worn by the user to acquire observation data and transmit it to the central control system.

[0365] "Observational data" refers to real-time data including the user's surrounding environment, facial expressions, and voice.

[0366] A "server device" is a central computing device that analyzes data sent from terminal devices and generates guidance information for users.

[0367] "Emotional analysis" is the process of determining a user's emotional state from their facial expressions and tone of voice.

[0368] A "generative AI model" is an artificial intelligence model that analyzes observational data and recognizes the surrounding environment.

[0369] "Travel guidance" refers to information provided to help users travel safely and comfortably.

[0370] "Feedback" refers to information obtained from users regarding their usage and emotional responses.

[0371] This invention relates to an individualized navigation system comprising a user-wearable terminal device and a server device. The terminal device is a portable information device such as smart glasses or a smartphone, and acquires user observation data using a camera and microphone. This observation data includes real-time information such as the user's facial expressions and tone of voice. This data is processed by an emotion engine to analyze the user's emotions.

[0372] The terminal device transmits the acquired observation data to the server device. The server device is a powerful computing system that analyzes the data using a generative AI model. This analysis identifies surrounding obstacles and landmarks and generates information to determine a safe travel path for the user.

[0373] The emotion engine analyzes the user's emotional state in real time and influences the generated guidance information. For example, if the user is feeling anxious, guidance will be provided in a calmer tone. Furthermore, the feedback function allows the generating AI model and emotion engine to continuously optimize based on information gained from the user's experience.

[0374] As a concrete example, consider a scenario where a user is walking through a crowded city and encounters an unexpected obstacle. In this case, the server quickly suggests a safe alternative route, and the device provides this information verbally through a speech synthesis engine. Simultaneously, visual guidelines can be displayed on the AR display.

[0375] An example of a prompt message is as follows:

[0376] "Analyze the user's current emotional state and generate a safe and reassuring travel route based on that analysis. This route should include information on landmarks and relaxation points to reduce the user's stress. Additionally, provide guidance with simple instructions to help the user cope in emergencies."

[0377] In this way, this system enhances user confidence and safety by providing dynamic and personalized guidance that responds to the user's unique state and surrounding circumstances.

[0378] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0379] Step 1:

[0380] The device acquires observational data from the camera and microphone using smart glasses or a smartphone. Input includes the user's facial expressions and voice data. This data is captured in real time and forms the basis for capturing the user's dynamic emotional changes. This observational data is then encrypted and sent to a server.

[0381] Step 2:

[0382] The device sends observational data to its built-in emotion engine, which then analyzes the user's emotional state. Here, observational data is used as input, and the data is processed using machine learning algorithms. The emotion engine outputs an emotion label based on facial features and voice tone, and sends this information to the server.

[0383] Step 3:

[0384] The server receives observational data transmitted from the terminal and analyzes it using a generative AI model. The input includes observational data and emotion labels. Through this analysis, the server processes the video data and identifies obstacles and landmarks in the surrounding environment. The output provides basic information for a safe travel route for the user.

[0385] Step 4:

[0386] The server generates user-directed navigation information based on the results of emotion analysis and environmental analysis. Input includes the user's emotional state and environmental details. A generative AI model is used to create guidance tailored to the user's emotions using prompts. These prompts may include instructions such as, "Please provide guidance in a slow voice to help the user relax." The final output is personalized guidance information, including voice guidance and visual guidelines.

[0387] Step 5:

[0388] The terminal receives guidance information sent from the server and presents it to the user. The input includes guidance information, which is then delivered to the user in natural-sounding voice using a speech synthesis engine. This also includes specific actions such as displaying visual guidelines via an AR display. Users can navigate with confidence based on the guidance information they receive. Feedback is also collected simultaneously and used for optimization in future applications.

[0389] (Application Example 2)

[0390] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0391] Conventional navigation systems struggle to provide personalized assistance based on the user's emotional state. Therefore, ensuring both emotional comfort and safety for passengers is a challenge, especially in autonomous vehicles. A travel experience that doesn't reflect the user's emotional state can cause unnecessary stress and anxiety.

[0392] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0393] In this invention, the server includes means for acquiring visual data obtained by a computing device worn by the user, means for transmitting the visual data from the computing device to a processing device, and means for analyzing the visual data using a generative model in the processing device to identify surrounding objects and reference points. This makes it possible to adjust the navigation guidance information based on the user's emotional state. Specifically, by flexibly adjusting the navigation route and output information according to the user's emotions, it is possible to provide passengers with a safer and more comfortable travel experience.

[0394] A "user" refers to an individual or group that uses this system, and is the subject whose emotional state is analyzed.

[0395] A "computational device" refers to a device worn by a user that has the function of acquiring and transmitting visual data.

[0396] "Visual data" refers to image and video information captured by the computing device, and is fundamental data for analyzing the user's emotional state.

[0397] A "processing unit" refers to a central computer that receives visual data transmitted from the computing unit and performs analysis using a generative model.

[0398] A "generative model" is a set of algorithms used in data analysis that perform intelligent processing to identify objects and reference points from visual data.

[0399] An "object" is a specific object that exists in the user's surrounding environment and is identified by the generative model.

[0400] A "reference point" refers to a landmark or marker used to determine the user's location in their surrounding environment, and is utilized in generating navigation information.

[0401] "Navigation guidance information" refers to guidelines and advice generated to assist users in their travels, and is adjusted based on the user's emotional state.

[0402] "Emotional state" refers to the emotional reactions and situations exhibited by the user, and is captured through emotion recognition sensors and visual data analysis.

[0403] To implement this invention, the user wears smart glasses or a smartphone as a computing device, thereby acquiring visual data. The device incorporates a camera and emotion recognition sensor to understand the user's emotions, and these are used to analyze video and audio data in real time.

[0404] The server receives encrypted visual data transmitted from the terminal and uses a generative model to identify surrounding objects and reference points. Based on this, it optimizes the user's movement path and generates appropriate navigation information. This navigation information is adjusted according to the user's emotional state; for example, if the user is feeling anxious, the server will explain the situation in a calm tone to support a more reassuring journey.

[0405] For example, if a passenger in an autonomous vehicle shows signs of anxiety, the vehicle could choose a more comfortable route and play calming music. In this way, support that takes user emotions into consideration can significantly improve the quality of the travel experience.

[0406] An example of a prompt would be, "What environmental conditions should you set based on the user's emotional state?" This serves as a guide for the generative AI model to learn the optimal guidance based on emotional states.

[0407] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0408] Step 1:

[0409] The device collects the user's visual and audio data. The device's built-in camera and emotion recognition sensor are activated, acquiring the user's facial expressions and audio information in real time. Through this process, the device obtains basic data as input to determine the user's emotional state.

[0410] Step 2:

[0411] The device securely and efficiently transmits the collected visual and audio data to the server. The data is encrypted and sent to the server via wireless communication. This allows the server to receive the data necessary for analysis while protecting privacy.

[0412] Step 3:

[0413] The server uses a generative AI model to analyze the received visual data and identify surrounding objects and reference points. It extracts features from the input data and recognizes obstacles and landmarks based on the generative model. As output, it provides detailed information about the user's environment.

[0414] Step 4:

[0415] The server integrates the user's emotional state with surrounding information to generate optimal navigation guidance. Based on the emotion recognition results, it determines the instructions that correspond to the user's emotions. For example, if the user is feeling anxious, it generates a more relaxing route or voice guidance. In this step, the generated guidance information is output.

[0416] Step 5:

[0417] The terminal displays guidance information received from the server to the user. This information is conveyed to the user via audio or visual means. The terminal continuously monitors the user's reactions and provides guidance in real time. This process allows the user to intuitively understand the situation and take the optimal action.

[0418] Step 6:

[0419] The user sends feedback to the server while on the move. By sending changes in emotion and responses to guidance as input to the server, the system's feedback loop is completed. This step allows the server to optimize the entire system, including the generative AI model, to improve the accuracy of future assistance.

[0420] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0421] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0422] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0423] [Third Embodiment]

[0424] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0425] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0426] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0427] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0428] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0429] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0430] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0431] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0432] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0433] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0434] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0435] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0436] The system of the present invention is specifically implemented using a user-wearable computing device, namely smart glasses or a smartphone. This device is equipped with a built-in camera to capture the user's surroundings in real time. When the user begins to move, the device acquires video data through its camera. This video data is compressed and encrypted as appropriate with the user's consent and then transmitted to a server via the internet.

[0437] The server analyzes the received video data and uses a generative model to identify obstacles and landmarks. This generative model incorporates the latest machine learning algorithms, enabling highly accurate analysis along with appropriate location information. Based on the analysis results, the server generates guidance messages for the user, including optimal route information and points of caution to ensure safe movement. This information is sent to the user's device as audio guidance or text.

[0438] Within the device, received information is converted into speech via a speech synthesis engine, and guidance is provided to the user in real time. This allows the user to accurately perceive their surroundings and independently avoid obstacles and reach their destination.

[0439] As a concrete example, consider a scenario where a user moves from a train platform to an exit. The terminal continuously sends camera footage to the server, and guidance based on the latest analysis results is provided each time. For example, a voice announcement such as "Please be careful of the next step" allows the user to safely cross the step and exit the station smoothly.

[0440] Furthermore, through user feedback, the server dynamically updates the generative model, improving the accuracy of the guides it provides. This continuous model optimization allows the system to adapt to users and continue to deliver a more efficient and personalized travel experience.

[0441] The following describes the processing flow.

[0442] Step 1:

[0443] The device uses its built-in camera to acquire video data of the user's surroundings in real time. This data includes information about the user's current environment.

[0444] Step 2:

[0445] The device compresses the acquired video data and sends it to the server while ensuring data security using encryption protocols. This reduces the risk of unauthorized access to the data.

[0446] Step 3:

[0447] The server receives video data transmitted from the terminal and begins analysis using a generative model. The model uses machine learning algorithms to identify obstacles and landmarks in the video and, based on this, determines the user's current location.

[0448] Step 4:

[0449] Based on the analysis results, the server generates guidance information to help users move safely and effectively. This information includes instructions for avoiding obstacles and guiding users using landmarks.

[0450] Step 5:

[0451] The server sends the generated guidance information to the terminal. This transmission is also carried out using an ideal protocol that ensures security, and efforts are made to maintain real-time performance.

[0452] Step 6:

[0453] The terminal converts the received guidance information into audio data and provides it to the user using a speech synthesis engine. This allows the user to understand their surroundings through the audio guidance and continue moving safely.

[0454] Step 7:

[0455] Users follow the instructions and enter feedback about their experience and actions into a terminal. This feedback is then sent to the server.

[0456] Step 8:

[0457] The server analyzes the received feedback and uses it to improve the accuracy of the generative model. This increases the overall guidance accuracy of the system, enabling safer and more effective navigation support for users.

[0458] (Example 1)

[0459] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0460] Existing mobility assistance systems have a problem in that they struggle to accurately recognize obstacles and landmarks in the user's surroundings and provide appropriate information to support safe and efficient movement. Furthermore, there is a challenge in that they cannot improve the system's accuracy by utilizing user feedback.

[0461] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0462] In this invention, the server includes means for acquiring image information captured by an information processing device wearable by the user; means for transferring the image information acquired from the information processing device to a central processing device; means for analyzing the image information using a generative model and detecting surrounding obstacles and landmarks in the central processing device; means for generating information to instruct the user's movement path based on the analysis results; means for transmitting the generated information to the information processing device and presenting it to the user visually or audibly; and means for acquiring information from the user and dynamically updating the generative model based on the acquired information. This makes it possible for the user to receive safe and optimal route guidance in real time and to provide feedback to improve the accuracy of the system.

[0463] A "user" refers to an individual who wears the system and utilizes mobility assistance services.

[0464] An "information processing device" refers to a device that a user can wear, which acquires image information of the environment and transmits it to a server.

[0465] "Image information" refers to visual data of the user's surroundings acquired by an information processing device, and the material used for analysis.

[0466] A "central processing unit" refers to a computer device that analyzes received image information and generates necessary support information.

[0467] A "generative model" refers to a predictive algorithm used to analyze image information and identify obstacles and landmarks.

[0468] An "obstacle" refers to any object that physically obstructs the user's path.

[0469] A "landmark" refers to a location or object that serves as a reference point to assist the user's movement.

[0470] "Travel route" refers to the appropriate direction of travel for a user to reach their destination.

[0471] "Information for giving instructions" refers to directions provided to the user based on the results of analysis using a generative model.

[0472] "Presenting visually or aurally" means communicating generated information to the user through a display or audio output.

[0473] "Feedback" refers to the user's reaction or opinion to the responses and guidance provided by the system.

[0474] "Dynamic updating" refers to adaptively adjusting the parameters and outputs of a generative model based on user feedback.

[0475] This invention is implemented using a wearable information processing device. Specifically, the user can perceive their surroundings in real time through a device such as smart glasses or a smartphone. This information processing device has a built-in high-resolution camera and is capable of continuously acquiring image information of the environment within the user's field of view.

[0476] The terminal processes the acquired image information, compresses the video for efficient data transfer, and protects the data using encryption algorithms such as AES-256 to ensure security. This compressed and encrypted data is transmitted to the server via a wireless communication network. Wi-Fi and mobile data communication are used, and the system is designed to exchange data with low latency even when the user is on the move.

[0477] The server uses a generative AI model to analyze the received image information. This analysis utilizes the latest machine learning platforms, TensorFlow and PyTorch. The generative model accurately identifies obstacles and landmarks in the user's surroundings and calculates the optimal movement path for the user. Based on the analysis results, the server generates audio and text guidance information to help the user move safely and efficiently.

[0478] The generated guide information is sent back to the terminal and presented to the user in real time using a speech synthesis engine. Utilizing natural language processing technology, the guidance is provided in an intuitive and easy-to-understand format. The user can then move safely by following this information.

[0479] As a concrete example, consider a scenario where a user is searching for an exit within a large public facility. This system supports smooth movement by quickly identifying potential obstacles the user might encounter and providing specific instructions such as, "You are approaching a wall; move to the left to avoid it."

[0480] An example of a prompt message would be, "Identify potential obstacles in the image and guide the user to the optimal detour." This allows the generative AI model to quickly generate and deliver the information the user needs.

[0481] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0482] Step 1:

[0483] The device activates its built-in camera and captures the user's surroundings in real time. The input is video frames from the camera, which are acquired and processed as continuous data. Specifically, high-resolution image data is acquired and prepared to be passed to the next processing step.

[0484] Step 2:

[0485] The terminal compresses and encrypts the acquired video data. The input is raw video frame data. This data is compressed using compression techniques such as H.264, and then encrypted using algorithms such as AES-256. As a result, video data is output that has been converted into a secure format while reducing the amount of data.

[0486] Step 3:

[0487] The terminal sends compressed and encrypted video data to the server. The input is compressed and encrypted data, and the output is transmitted over the internet via Wi-Fi or a mobile network. The goal of this step is to transfer the data efficiently and with low latency.

[0488] Step 4:

[0489] The server decrypts and decompresses the received data. The input is encrypted and compressed video data. By decrypting the data using AES-256 or similar encryption methods, and then decompressing the compressed data, the server obtains analyzable raw data.

[0490] Step 5:

[0491] The server analyzes the decoded data using a generative AI model. The input is decoded video data. The generative AI model (using TensorFlow or PyTorch) analyzes the data to identify obstacles and landmarks. As output, location information and type information are obtained as analysis results.

[0492] Step 6:

[0493] The server generates guidance messages based on the analysis results. The input is the analysis results from the AI ​​model. Using natural language processing technology, it generates guidance messages that include necessary travel routes and points to note for the user, and provides the information as output in the form of voice or text.

[0494] Step 7:

[0495] The server sends the generated guidance message to the terminal. The input consists of the generated audio and text data. This data is then forwarded back to the terminal and output so that it can be received by the user's device.

[0496] Step 8:

[0497] The terminal outputs received guidance messages through a speech synthesis engine. The input consists of voice and text data received from the server. Speech synthesis converts this data into voice messages in real time, directly providing instructions to the user audibly.

[0498] Step 9:

[0499] Users provide feedback based on their experiences while actually moving around. The input consists of their real-world travel experiences and impressions. This data is sent to a server via the terminal and used as output for further optimization of the AI ​​model and system improvement.

[0500] (Application Example 1)

[0501] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] There is a need for technology that enables autonomous vehicles to move safely and efficiently in complex environments. Conventional technologies have problems such as difficulty in accurately understanding the surrounding environment using only the sensors on the vehicle, and the inability to provide the necessary information to improve safety in real time. Therefore, there is a need to develop a dynamic guidance system that allows the vehicle to accurately perceive its surroundings and avoid obstacles.

[0503] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0504] In this invention, the server includes means for acquiring video data captured by a computer device worn by the user, means for transmitting the video data acquired from the computer device to a central processing unit, and means for analyzing the video data using a generative model to identify surrounding obstacles and landmarks. This makes it possible to accurately recognize the surrounding conditions of a transported object and provide an appropriate travel path.

[0505] A "user-worn computing device" is an information processing device that functions in a form that can be worn by an individual, and has the ability to acquire video data and communicate with other devices.

[0506] "Means of acquiring video data" refers to a system that collects visual information as digital data using a camera or other imaging device.

[0507] A "central processing unit" is a primary data processing unit that analyzes received data and generates information for decision-making and control.

[0508] A "generative model" is a computational model designed to analyze and predict data using algorithms based on machine learning or artificial intelligence.

[0509] "Transportation equipment" refers to vehicles and mechanical devices used as means of moving people or goods.

[0510] "Information for guiding travel routes" refers to data intended to guide travelers, including instructions and precautions regarding the optimal route to their destination.

[0511] "User feedback" refers to opinions and evaluations regarding the experience and performance of the system, provided by individuals or organizations that use the system.

[0512] The system that realizes this invention basically consists of a computer device worn by the user (such as smart glasses or a smartphone) and a server. The user's computer device is equipped with a camera that acquires real-time video data of the surroundings while the transported object is moving. The acquired video data is compressed and encrypted via the internet and transmitted to the server.

[0513] Upon receiving video data, the server uses a generative AI model to perform analysis. This analysis employs state-of-the-art machine learning algorithms, enabling high-precision identification of surrounding obstacles and landmarks. Based on the analysis results, the server generates information to guide the user's transported vehicle safely and efficiently along its route. This includes information on how to avoid obstacles and points of caution along the route.

[0514] The generated guidance information is transmitted to the user's computer and presented as audio guides or text messages. This enables the autonomous movement of the transported object. Furthermore, the server receives feedback from the user, dynamically updating the generated AI model and continuously improving the accuracy of the guidance.

[0515] As a concrete example, consider an autonomous driving scenario in an urban area. When a moving vehicle approaches an intersection, the server uses video analysis to recognize traffic lights and pedestrians and provides voice guidance such as, "The light has turned red, please stop." This ensures safe driving.

[0516] An example of a prompt message might be, "Analyze the current road conditions and provide a safe route." By entering this prompt message, the user can begin the system's analysis and guidance process.

[0517] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0518] Step 1:

[0519] The device uses its built-in camera to capture the user's surroundings as video data in real time. The captured video data is compressed and encrypted and sent to a server via the internet. In this step, the input is visual information of the surroundings, and the output is compressed and encrypted video data. The device uses a data processing algorithm to convert the video into an appropriate format.

[0520] Step 2:

[0521] The server receives video data transmitted from the terminal and performs decompression and decryption of the data. It receives compressed and encrypted data as input and obtains restored video data as output. The server decrypts the data using security protocols and prepares it in an analyzable format.

[0522] Step 3:

[0523] The server uses a generative AI model to analyze the received video data and identify obstacles and landmarks. It takes the reconstructed video data as input and generates analyzed environmental information as output. In this process, machine learning algorithms utilize image recognition technology to extract specific features.

[0524] Step 4:

[0525] The server generates information to guide the transporter's movement path based on the analysis results. It uses the analyzed environmental information as input and generates instructions, including route guidance and points of caution, as output. Based on the environmental information, the server calculates a safe and efficient route and constructs a guidance message.

[0526] Step 5:

[0527] The server sends the generated guidance information to the terminal. It takes guidance information as input and distributes data to the terminal as output. The server provides information in real time using a communication protocol.

[0528] Step 6:

[0529] The terminal presents the received guidance information to the user as an audio guide or text. It receives guidance information transmitted as input and displays it as audio or text as output. The terminal uses a speech synthesis engine to generate clear audio guides and deliver instructions to the user.

[0530] Step 7:

[0531] The user provides feedback on the operation and environmental changes during the movement of the transported object. The system receives user feedback as input and obtains updated data for the generated AI model as output. This feedback is used by the server as crucial data for further system optimization.

[0532] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0533] This invention relates to a system that incorporates an emotion engine to further personalize user assistance. This system is implemented using a user-wearable computing device, such as smart glasses or a smartphone. These devices are equipped with a built-in camera and emotion recognition sensor, which can analyze the user's emotions from their facial expressions and voice.

[0534] The device acquires real-time video data from the user and transmits it to the central processing unit. During the transmission of video data, the data is appropriately encrypted to protect the user's privacy. The device also analyzes emotional data using an emotion engine and transmits that information to the server.

[0535] The server analyzes video data transmitted from the terminal using a generative model to identify obstacles and landmarks in the environment. Furthermore, it uses data from the emotion engine to understand the user's emotional state and generate appropriate navigation information. For example, if the user is feeling stressed, the server will provide navigation information in a calmer tone and avoid more direct instructions.

[0536] The generated guidance information is transmitted to the terminal via a speech synthesis engine. The terminal presents the transmitted information to the user verbally and visually, providing real-time guidance. This personalized guidance allows the user to have a better travel experience while ensuring security.

[0537] As a concrete example, consider a situation where a user is walking through a busy area and a car suddenly approaches. In this case, if the emotion engine detects the user's anxiety, the server will send swift and calm instructions to inform the user of how to safely avoid the car. In this way, instructions that take the user's emotional state into consideration significantly improve the user's sense of security and safety.

[0538] Feedback is also collected based on emotional states, and the server uses this to optimize its generative model and emotion engine. As a result, the system can continuously adapt to the user's changing emotions and provide optimal guidance.

[0539] The following describes the processing flow.

[0540] Step 1:

[0541] The device uses a built-in camera to acquire video data of the user's surroundings in real time. Simultaneously, it uses an emotion recognition sensor to collect emotional data from the user's facial expressions and voice.

[0542] Step 2:

[0543] The device encrypts the acquired video and emotional data to protect privacy and transmits it to the server using a secure communication protocol.

[0544] Step 3:

[0545] The server activates a generative model to analyze the received video data and identify obstacles and landmarks in the environment. Machine learning algorithms are used in this process.

[0546] Step 4:

[0547] The server uses an emotion engine to analyze the received emotion data and evaluate the user's emotional state. This determines the user's current emotional state.

[0548] Step 5:

[0549] The server integrates the analysis results and generates navigation information tailored to the user's current emotional state. For example, if the user is stressed, the guidance will be delivered in a gentle tone, and directions will be simplified in crowded areas.

[0550] Step 6:

[0551] The server sends the generated guidance information to the terminal. This information is then sent to the speech synthesis engine and provided to the user as an easy-to-understand voice message.

[0552] Step 7:

[0553] The terminal provides users with guidance information via speech synthesis, offering real-time navigation assistance. This enables users to travel to their destination safely and effectively.

[0554] Step 8:

[0555] Users follow the provided instructions and enter feedback about their experience and feelings into a terminal. This feedback is later sent to the server.

[0556] Step 9:

[0557] The server receives user feedback and uses it as data to optimize the generative model and sentiment engine. This allows the system to adapt more closely to the user and provide more accurate support.

[0558] (Example 2)

[0559] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0560] Conventional mobility assistance systems struggle to provide guidance that takes into account the individual emotional state of the user, resulting in many situations where users experience anxiety and stress. To solve this problem and improve the user's mobility experience, it is necessary to analyze the user's emotions in real time and provide personalized guidance based on that analysis.

[0561] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0562] In this invention, the server includes means for analyzing observational data using a generative AI model to identify elements within the environment, means for analyzing the user's emotional state, and means for generating and providing navigation guidance information based on this information. This makes it possible to provide guidance tailored to the user's individual emotional state in real time.

[0563] A "user" is an entity that wears a computer device and receives guidance information provided by the system.

[0564] A "terminal device" is a computer device worn by the user to acquire observation data and transmit it to the central control system.

[0565] "Observational data" refers to real-time data including the user's surrounding environment, facial expressions, and voice.

[0566] A "server device" is a central computing device that analyzes data sent from terminal devices and generates guidance information for users.

[0567] "Emotional analysis" is the process of determining a user's emotional state from their facial expressions and tone of voice.

[0568] A "generative AI model" is an artificial intelligence model that analyzes observational data and recognizes the surrounding environment.

[0569] "Travel guidance" refers to information provided to help users travel safely and comfortably.

[0570] "Feedback" refers to information obtained from users regarding their usage and emotional responses.

[0571] This invention relates to an individualized navigation system comprising a user-wearable terminal device and a server device. The terminal device is a portable information device such as smart glasses or a smartphone, and acquires user observation data using a camera and microphone. This observation data includes real-time information such as the user's facial expressions and tone of voice. This data is processed by an emotion engine to analyze the user's emotions.

[0572] The terminal device transmits the acquired observation data to the server device. The server device is a powerful computing system that analyzes the data using a generative AI model. This analysis identifies surrounding obstacles and landmarks and generates information to determine a safe travel path for the user.

[0573] The emotion engine analyzes the user's emotional state in real time and influences the generated guidance information. For example, if the user is feeling anxious, guidance will be provided in a calmer tone. Furthermore, the feedback function allows the generating AI model and emotion engine to continuously optimize based on information gained from the user's experience.

[0574] As a concrete example, consider a scenario where a user is walking through a crowded city and encounters an unexpected obstacle. In this case, the server quickly suggests a safe alternative route, and the device provides this information verbally through a speech synthesis engine. Simultaneously, visual guidelines can be displayed on the AR display.

[0575] An example of a prompt message is as follows:

[0576] "Analyze the user's current emotional state and generate a safe and reassuring travel route based on that analysis. This route should include information on landmarks and relaxation points to reduce the user's stress. Additionally, provide guidance with simple instructions to help the user cope in emergencies."

[0577] In this way, this system enhances user confidence and safety by providing dynamic and personalized guidance that responds to the user's unique state and surrounding circumstances.

[0578] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0579] Step 1:

[0580] The device acquires observational data from the camera and microphone using smart glasses or a smartphone. Input includes the user's facial expressions and voice data. This data is captured in real time and forms the basis for capturing the user's dynamic emotional changes. This observational data is then encrypted and sent to a server.

[0581] Step 2:

[0582] The device sends observational data to its built-in emotion engine, which then analyzes the user's emotional state. Here, observational data is used as input, and the data is processed using machine learning algorithms. The emotion engine outputs an emotion label based on facial features and voice tone, and sends this information to the server.

[0583] Step 3:

[0584] The server receives observational data transmitted from the terminal and analyzes it using a generative AI model. The input includes observational data and emotion labels. Through this analysis, the server processes the video data and identifies obstacles and landmarks in the surrounding environment. The output provides basic information for a safe travel route for the user.

[0585] Step 4:

[0586] The server generates user-directed navigation information based on the results of emotion analysis and environmental analysis. Input includes the user's emotional state and environmental details. A generative AI model is used to create guidance tailored to the user's emotions using prompts. These prompts may include instructions such as, "Please provide guidance in a slow voice to help the user relax." The final output is personalized guidance information, including voice guidance and visual guidelines.

[0587] Step 5:

[0588] The terminal receives guidance information sent from the server and presents it to the user. The input includes guidance information, which is then delivered to the user in natural-sounding voice using a speech synthesis engine. This also includes specific actions such as displaying visual guidelines via an AR display. Users can navigate with confidence based on the guidance information they receive. Feedback is also collected simultaneously and used for optimization in future applications.

[0589] (Application Example 2)

[0590] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0591] Conventional navigation systems struggle to provide personalized assistance based on the user's emotional state. Therefore, ensuring both emotional comfort and safety for passengers is a challenge, especially in autonomous vehicles. A travel experience that doesn't reflect the user's emotional state can cause unnecessary stress and anxiety.

[0592] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0593] In this invention, the server includes means for acquiring visual data obtained by a computing device worn by the user, means for transmitting the visual data from the computing device to a processing device, and means for analyzing the visual data using a generative model in the processing device to identify surrounding objects and reference points. This makes it possible to adjust the navigation guidance information based on the user's emotional state. Specifically, by flexibly adjusting the navigation route and output information according to the user's emotions, it is possible to provide passengers with a safer and more comfortable travel experience.

[0594] A "user" refers to an individual or group that uses this system, and is the subject whose emotional state is analyzed.

[0595] A "computational device" refers to a device worn by a user that has the function of acquiring and transmitting visual data.

[0596] "Visual data" refers to image and video information captured by the computing device, and is fundamental data for analyzing the user's emotional state.

[0597] A "processing unit" refers to a central computer that receives visual data transmitted from the computing unit and performs analysis using a generative model.

[0598] A "generative model" is a set of algorithms used in data analysis that perform intelligent processing to identify objects and reference points from visual data.

[0599] An "object" is a specific object that exists in the user's surrounding environment and is identified by the generative model.

[0600] A "reference point" refers to a landmark or marker used to determine the user's location in their surrounding environment, and is utilized in generating navigation information.

[0601] "Navigation guidance information" refers to guidelines and advice generated to assist users in their travels, and is adjusted based on the user's emotional state.

[0602] "Emotional state" refers to the emotional reactions and situations exhibited by the user, and is captured through emotion recognition sensors and visual data analysis.

[0603] To implement this invention, the user wears smart glasses or a smartphone as a computing device, thereby acquiring visual data. The device incorporates a camera and emotion recognition sensor to understand the user's emotions, and these are used to analyze video and audio data in real time.

[0604] The server receives encrypted visual data transmitted from the terminal and uses a generative model to identify surrounding objects and reference points. Based on this, it optimizes the user's movement path and generates appropriate navigation information. This navigation information is adjusted according to the user's emotional state; for example, if the user is feeling anxious, the server will explain the situation in a calm tone to support a more reassuring journey.

[0605] For example, if a passenger in an autonomous vehicle shows signs of anxiety, the vehicle could choose a more comfortable route and play calming music. In this way, support that takes user emotions into consideration can significantly improve the quality of the travel experience.

[0606] An example of a prompt would be, "What environmental conditions should you set based on the user's emotional state?" This serves as a guide for the generative AI model to learn the optimal guidance based on emotional states.

[0607] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0608] Step 1:

[0609] The device collects the user's visual and audio data. The device's built-in camera and emotion recognition sensor are activated, acquiring the user's facial expressions and audio information in real time. Through this process, the device obtains basic data as input to determine the user's emotional state.

[0610] Step 2:

[0611] The device securely and efficiently transmits the collected visual and audio data to the server. The data is encrypted and sent to the server via wireless communication. This allows the server to receive the data necessary for analysis while protecting privacy.

[0612] Step 3:

[0613] The server uses a generative AI model to analyze the received visual data and identify surrounding objects and reference points. It extracts features from the input data and recognizes obstacles and landmarks based on the generative model. As output, it provides detailed information about the user's environment.

[0614] Step 4:

[0615] The server integrates the user's emotional state with surrounding information to generate optimal navigation guidance. Based on the emotion recognition results, it determines the instructions that correspond to the user's emotions. For example, if the user is feeling anxious, it generates a more relaxing route or voice guidance. In this step, the generated guidance information is output.

[0616] Step 5:

[0617] The terminal displays guidance information received from the server to the user. This information is conveyed to the user via audio or visual means. The terminal continuously monitors the user's reactions and provides guidance in real time. This process allows the user to intuitively understand the situation and take the optimal action.

[0618] Step 6:

[0619] The user sends feedback to the server while on the move. By sending changes in emotion and responses to guidance as input to the server, the system's feedback loop is completed. This step allows the server to optimize the entire system, including the generative AI model, to improve the accuracy of future assistance.

[0620] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0621] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0622] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0623] [Fourth Embodiment]

[0624] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0625] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0626] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0627] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0628] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0629] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0630] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0631] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0632] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0633] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0634] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0635] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0636] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0637] The system of the present invention is specifically implemented using a user-wearable computing device, namely smart glasses or a smartphone. This device is equipped with a built-in camera to capture the user's surroundings in real time. When the user begins to move, the device acquires video data through its camera. This video data is compressed and encrypted as appropriate with the user's consent and then transmitted to a server via the internet.

[0638] The server analyzes the received video data and uses a generative model to identify obstacles and landmarks. This generative model incorporates the latest machine learning algorithms, enabling highly accurate analysis along with appropriate location information. Based on the analysis results, the server generates guidance messages for the user, including optimal route information and points of caution to ensure safe movement. This information is sent to the user's device as audio guidance or text.

[0639] Within the device, received information is converted into speech via a speech synthesis engine, and guidance is provided to the user in real time. This allows the user to accurately perceive their surroundings and independently avoid obstacles and reach their destination.

[0640] As a concrete example, consider a scenario where a user moves from a train platform to an exit. The terminal continuously sends camera footage to the server, and guidance based on the latest analysis results is provided each time. For example, a voice announcement such as "Please be careful of the next step" allows the user to safely cross the step and exit the station smoothly.

[0641] Furthermore, through user feedback, the server dynamically updates the generative model, improving the accuracy of the guides it provides. This continuous model optimization allows the system to adapt to users and continue to deliver a more efficient and personalized travel experience.

[0642] The following describes the processing flow.

[0643] Step 1:

[0644] The device uses its built-in camera to acquire video data of the user's surroundings in real time. This data includes information about the user's current environment.

[0645] Step 2:

[0646] The device compresses the acquired video data and sends it to the server while ensuring data security using encryption protocols. This reduces the risk of unauthorized access to the data.

[0647] Step 3:

[0648] The server receives video data transmitted from the terminal and begins analysis using a generative model. The model uses machine learning algorithms to identify obstacles and landmarks in the video and, based on this, determines the user's current location.

[0649] Step 4:

[0650] Based on the analysis results, the server generates guidance information to help users move safely and effectively. This information includes instructions for avoiding obstacles and guiding users using landmarks.

[0651] Step 5:

[0652] The server sends the generated guidance information to the terminal. This transmission is also carried out using an ideal protocol that ensures security, and efforts are made to maintain real-time performance.

[0653] Step 6:

[0654] The terminal converts the received guidance information into audio data and provides it to the user using a speech synthesis engine. This allows the user to understand their surroundings through the audio guidance and continue moving safely.

[0655] Step 7:

[0656] Users follow the instructions and enter feedback about their experience and actions into a terminal. This feedback is then sent to the server.

[0657] Step 8:

[0658] The server analyzes the received feedback and uses it to improve the accuracy of the generative model. This increases the overall guidance accuracy of the system, enabling safer and more effective navigation support for users.

[0659] (Example 1)

[0660] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0661] Existing mobility assistance systems have a problem in that they struggle to accurately recognize obstacles and landmarks in the user's surroundings and provide appropriate information to support safe and efficient movement. Furthermore, there is a challenge in that they cannot improve the system's accuracy by utilizing user feedback.

[0662] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0663] In this invention, the server includes means for acquiring image information captured by an information processing device wearable by the user; means for transferring the image information acquired from the information processing device to a central processing device; means for analyzing the image information using a generative model and detecting surrounding obstacles and landmarks in the central processing device; means for generating information to instruct the user's movement path based on the analysis results; means for transmitting the generated information to the information processing device and presenting it to the user visually or audibly; and means for acquiring information from the user and dynamically updating the generative model based on the acquired information. This makes it possible for the user to receive safe and optimal route guidance in real time and to provide feedback to improve the accuracy of the system.

[0664] A "user" refers to an individual who wears the system and utilizes mobility assistance services.

[0665] An "information processing device" refers to a device that a user can wear, which acquires image information of the environment and transmits it to a server.

[0666] "Image information" refers to visual data of the user's surroundings acquired by an information processing device, and the material used for analysis.

[0667] A "central processing unit" refers to a computer device that analyzes received image information and generates necessary support information.

[0668] A "generative model" refers to a predictive algorithm used to analyze image information and identify obstacles and landmarks.

[0669] An "obstacle" refers to any object that physically obstructs the user's path.

[0670] A "landmark" refers to a location or object that serves as a reference point to assist the user's movement.

[0671] "Travel route" refers to the appropriate direction of travel for a user to reach their destination.

[0672] "Information for giving instructions" refers to directions provided to the user based on the results of analysis using a generative model.

[0673] "Presenting visually or aurally" means communicating generated information to the user through a display or audio output.

[0674] "Feedback" refers to the user's reaction or opinion to the responses and guidance provided by the system.

[0675] "Dynamic updating" refers to adaptively adjusting the parameters and outputs of a generative model based on user feedback.

[0676] This invention is implemented using a wearable information processing device. Specifically, the user can perceive their surroundings in real time through a device such as smart glasses or a smartphone. This information processing device has a built-in high-resolution camera and is capable of continuously acquiring image information of the environment within the user's field of view.

[0677] The terminal processes the acquired image information, compresses the video for efficient data transfer, and protects the data using encryption algorithms such as AES-256 to ensure security. This compressed and encrypted data is transmitted to the server via a wireless communication network. Wi-Fi and mobile data communication are used, and the system is designed to exchange data with low latency even when the user is on the move.

[0678] The server uses a generative AI model to analyze the received image information. This analysis utilizes the latest machine learning platforms, TensorFlow and PyTorch. The generative model accurately identifies obstacles and landmarks in the user's surroundings and calculates the optimal movement path for the user. Based on the analysis results, the server generates audio and text guidance information to help the user move safely and efficiently.

[0679] The generated guide information is sent back to the terminal and presented to the user in real time using a speech synthesis engine. Utilizing natural language processing technology, the guidance is provided in an intuitive and easy-to-understand format. The user can then move safely by following this information.

[0680] As a concrete example, consider a scenario where a user is searching for an exit within a large public facility. This system supports smooth movement by quickly identifying potential obstacles the user might encounter and providing specific instructions such as, "You are approaching a wall; move to the left to avoid it."

[0681] An example of a prompt message would be, "Identify potential obstacles in the image and guide the user to the optimal detour." This allows the generative AI model to quickly generate and deliver the information the user needs.

[0682] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0683] Step 1:

[0684] The device activates its built-in camera and captures the user's surroundings in real time. The input is video frames from the camera, which are acquired and processed as continuous data. Specifically, high-resolution image data is acquired and prepared to be passed to the next processing step.

[0685] Step 2:

[0686] The terminal compresses and encrypts the acquired video data. The input is raw video frame data. This data is compressed using compression techniques such as H.264, and then encrypted using algorithms such as AES-256. As a result, video data is output that has been converted into a secure format while reducing the amount of data.

[0687] Step 3:

[0688] The terminal sends compressed and encrypted video data to the server. The input is compressed and encrypted data, and the output is transmitted over the internet via Wi-Fi or a mobile network. The goal of this step is to transfer the data efficiently and with low latency.

[0689] Step 4:

[0690] The server decrypts and decompresses the received data. The input is encrypted and compressed video data. By decrypting the data using AES-256 or similar encryption methods, and then decompressing the compressed data, the server obtains analyzable raw data.

[0691] Step 5:

[0692] The server analyzes the decoded data using a generative AI model. The input is decoded video data. The generative AI model (using TensorFlow or PyTorch) analyzes the data to identify obstacles and landmarks. As output, location information and type information are obtained as analysis results.

[0693] Step 6:

[0694] The server generates guidance messages based on the analysis results. The input is the analysis results from the AI ​​model. Using natural language processing technology, it generates guidance messages that include necessary travel routes and points to note for the user, and provides the information as output in the form of voice or text.

[0695] Step 7:

[0696] The server sends the generated guidance message to the terminal. The input consists of the generated audio and text data. This data is then forwarded back to the terminal and output so that it can be received by the user's device.

[0697] Step 8:

[0698] The terminal outputs received guidance messages through a speech synthesis engine. The input consists of voice and text data received from the server. Speech synthesis converts this data into voice messages in real time, directly providing instructions to the user audibly.

[0699] Step 9:

[0700] Users provide feedback based on their experiences while actually moving around. The input consists of their real-world travel experiences and impressions. This data is sent to a server via the terminal and used as output for further optimization of the AI ​​model and system improvement.

[0701] (Application Example 1)

[0702] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0703] There is a need for technology that enables autonomous vehicles to move safely and efficiently in complex environments. Conventional technologies have problems such as difficulty in accurately understanding the surrounding environment using only the sensors on the vehicle, and the inability to provide the necessary information to improve safety in real time. Therefore, there is a need to develop a dynamic guidance system that allows the vehicle to accurately perceive its surroundings and avoid obstacles.

[0704] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0705] In this invention, the server includes means for acquiring video data captured by a computer device worn by the user, means for transmitting the video data acquired from the computer device to a central processing unit, and means for analyzing the video data using a generative model to identify surrounding obstacles and landmarks. This makes it possible to accurately recognize the surrounding conditions of a transported object and provide an appropriate travel path.

[0706] A "user-worn computing device" is an information processing device that functions in a form that can be worn by an individual, and has the ability to acquire video data and communicate with other devices.

[0707] "Means of acquiring video data" refers to a system that collects visual information as digital data using a camera or other imaging device.

[0708] A "central processing unit" is a primary data processing unit that analyzes received data and generates information for decision-making and control.

[0709] A "generative model" is a computational model designed to analyze and predict data using algorithms based on machine learning or artificial intelligence.

[0710] "Transportation equipment" refers to vehicles and mechanical devices used as means of moving people or goods.

[0711] "Information for guiding travel routes" refers to data intended to guide travelers, including instructions and precautions regarding the optimal route to their destination.

[0712] "User feedback" refers to opinions and evaluations regarding the experience and performance of the system, provided by individuals or organizations that use the system.

[0713] The system that realizes this invention basically consists of a computer device worn by the user (such as smart glasses or a smartphone) and a server. The user's computer device is equipped with a camera that acquires real-time video data of the surroundings while the transported object is moving. The acquired video data is compressed and encrypted via the internet and transmitted to the server.

[0714] Upon receiving video data, the server uses a generative AI model to perform analysis. This analysis employs state-of-the-art machine learning algorithms, enabling high-precision identification of surrounding obstacles and landmarks. Based on the analysis results, the server generates information to guide the user's transported vehicle safely and efficiently along its route. This includes information on how to avoid obstacles and points of caution along the route.

[0715] The generated guidance information is transmitted to the user's computer and presented as audio guides or text messages. This enables the autonomous movement of the transported object. Furthermore, the server receives feedback from the user, dynamically updating the generated AI model and continuously improving the accuracy of the guidance.

[0716] As a concrete example, consider an autonomous driving scenario in an urban area. When a moving vehicle approaches an intersection, the server uses video analysis to recognize traffic lights and pedestrians and provides voice guidance such as, "The light has turned red, please stop." This ensures safe driving.

[0717] An example of a prompt message might be, "Analyze the current road conditions and provide a safe route." By entering this prompt message, the user can begin the system's analysis and guidance process.

[0718] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0719] Step 1:

[0720] The device uses its built-in camera to capture the user's surroundings as video data in real time. The captured video data is compressed and encrypted and sent to a server via the internet. In this step, the input is visual information of the surroundings, and the output is compressed and encrypted video data. The device uses a data processing algorithm to convert the video into an appropriate format.

[0721] Step 2:

[0722] The server receives video data transmitted from the terminal and performs decompression and decryption of the data. It receives compressed and encrypted data as input and obtains restored video data as output. The server decrypts the data using security protocols and prepares it in an analyzable format.

[0723] Step 3:

[0724] The server uses a generative AI model to analyze the received video data and identify obstacles and landmarks. It takes the reconstructed video data as input and generates analyzed environmental information as output. In this process, machine learning algorithms utilize image recognition technology to extract specific features.

[0725] Step 4:

[0726] The server generates information to guide the transporter's movement path based on the analysis results. It uses the analyzed environmental information as input and generates instructions, including route guidance and points of caution, as output. Based on the environmental information, the server calculates a safe and efficient route and constructs a guidance message.

[0727] Step 5:

[0728] The server sends the generated guidance information to the terminal. It takes guidance information as input and distributes data to the terminal as output. The server provides information in real time using a communication protocol.

[0729] Step 6:

[0730] The terminal presents the received guidance information to the user as an audio guide or text. It receives guidance information transmitted as input and displays it as audio or text as output. The terminal uses a speech synthesis engine to generate clear audio guides and deliver instructions to the user.

[0731] Step 7:

[0732] The user provides feedback on the operation and environmental changes during the movement of the transported object. The system receives user feedback as input and obtains updated data for the generated AI model as output. This feedback is used by the server as crucial data for further system optimization.

[0733] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0734] This invention relates to a system that incorporates an emotion engine to further personalize user assistance. This system is implemented using a user-wearable computing device, such as smart glasses or a smartphone. These devices are equipped with a built-in camera and emotion recognition sensor, which can analyze the user's emotions from their facial expressions and voice.

[0735] The device acquires real-time video data from the user and transmits it to the central processing unit. During the transmission of video data, the data is appropriately encrypted to protect the user's privacy. The device also analyzes emotional data using an emotion engine and transmits that information to the server.

[0736] The server analyzes video data transmitted from the terminal using a generative model to identify obstacles and landmarks in the environment. Furthermore, it uses data from the emotion engine to understand the user's emotional state and generate appropriate navigation information. For example, if the user is feeling stressed, the server will provide navigation information in a calmer tone and avoid more direct instructions.

[0737] The generated guidance information is transmitted to the terminal via a speech synthesis engine. The terminal presents the transmitted information to the user verbally and visually, providing real-time guidance. This personalized guidance allows the user to have a better travel experience while ensuring security.

[0738] As a concrete example, consider a situation where a user is walking through a busy area and a car suddenly approaches. In this case, if the emotion engine detects the user's anxiety, the server will send swift and calm instructions to inform the user of how to safely avoid the car. In this way, instructions that take the user's emotional state into consideration significantly improve the user's sense of security and safety.

[0739] Feedback is also collected based on emotional states, and the server uses this to optimize its generative model and emotion engine. As a result, the system can continuously adapt to the user's changing emotions and provide optimal guidance.

[0740] The following describes the processing flow.

[0741] Step 1:

[0742] The device uses a built-in camera to acquire video data of the user's surroundings in real time. Simultaneously, it uses an emotion recognition sensor to collect emotional data from the user's facial expressions and voice.

[0743] Step 2:

[0744] The device encrypts the acquired video and emotional data to protect privacy and transmits it to the server using a secure communication protocol.

[0745] Step 3:

[0746] The server activates a generative model to analyze the received video data and identify obstacles and landmarks in the environment. Machine learning algorithms are used in this process.

[0747] Step 4:

[0748] The server uses an emotion engine to analyze the received emotion data and evaluate the user's emotional state. This determines the user's current emotional state.

[0749] Step 5:

[0750] The server integrates the analysis results and generates navigation information tailored to the user's current emotional state. For example, if the user is stressed, the guidance will be delivered in a gentle tone, and directions will be simplified in crowded areas.

[0751] Step 6:

[0752] The server sends the generated guidance information to the terminal. This information is then sent to the speech synthesis engine and provided to the user as an easy-to-understand voice message.

[0753] Step 7:

[0754] The terminal provides users with guidance information via speech synthesis, offering real-time navigation assistance. This enables users to travel to their destination safely and effectively.

[0755] Step 8:

[0756] Users follow the provided instructions and enter feedback about their experience and feelings into a terminal. This feedback is later sent to the server.

[0757] Step 9:

[0758] The server receives user feedback and uses it as data to optimize the generative model and sentiment engine. This allows the system to adapt more closely to the user and provide more accurate support.

[0759] (Example 2)

[0760] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0761] Conventional mobility assistance systems struggle to provide guidance that takes into account the individual emotional state of the user, resulting in many situations where users experience anxiety and stress. To solve this problem and improve the user's mobility experience, it is necessary to analyze the user's emotions in real time and provide personalized guidance based on that analysis.

[0762] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0763] In this invention, the server includes means for analyzing observational data using a generative AI model to identify elements within the environment, means for analyzing the user's emotional state, and means for generating and providing navigation guidance information based on this information. This makes it possible to provide guidance tailored to the user's individual emotional state in real time.

[0764] A "user" is an entity that wears a computer device and receives guidance information provided by the system.

[0765] A "terminal device" is a computer device worn by the user to acquire observation data and transmit it to the central control system.

[0766] "Observational data" refers to real-time data including the user's surrounding environment, facial expressions, and voice.

[0767] A "server device" is a central computing device that analyzes data sent from terminal devices and generates guidance information for users.

[0768] "Emotional analysis" is the process of determining a user's emotional state from their facial expressions and tone of voice.

[0769] A "generative AI model" is an artificial intelligence model that analyzes observational data and recognizes the surrounding environment.

[0770] "Travel guidance" refers to information provided to help users travel safely and comfortably.

[0771] "Feedback" refers to information obtained from users regarding their usage and emotional responses.

[0772] This invention relates to an individualized navigation system comprising a user-wearable terminal device and a server device. The terminal device is a portable information device such as smart glasses or a smartphone, and acquires user observation data using a camera and microphone. This observation data includes real-time information such as the user's facial expressions and tone of voice. This data is processed by an emotion engine to analyze the user's emotions.

[0773] The terminal device transmits the acquired observation data to the server device. The server device is a powerful computing system that analyzes the data using a generative AI model. This analysis identifies surrounding obstacles and landmarks and generates information to determine a safe travel path for the user.

[0774] The emotion engine analyzes the user's emotional state in real time and influences the generated guidance information. For example, if the user is feeling anxious, guidance will be provided in a calmer tone. Furthermore, the feedback function allows the generating AI model and emotion engine to continuously optimize based on information gained from the user's experience.

[0775] As a concrete example, consider a scenario where a user is walking through a crowded city and encounters an unexpected obstacle. In this case, the server quickly suggests a safe alternative route, and the device provides this information verbally through a speech synthesis engine. Simultaneously, visual guidelines can be displayed on the AR display.

[0776] An example of a prompt message is as follows:

[0777] "Analyze the user's current emotional state and generate a safe and reassuring travel route based on that analysis. This route should include information on landmarks and relaxation points to reduce the user's stress. Additionally, provide guidance with simple instructions to help the user cope in emergencies."

[0778] In this way, this system enhances user confidence and safety by providing dynamic and personalized guidance that responds to the user's unique state and surrounding circumstances.

[0779] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0780] Step 1:

[0781] The device acquires observational data from the camera and microphone using smart glasses or a smartphone. Input includes the user's facial expressions and voice data. This data is captured in real time and forms the basis for capturing the user's dynamic emotional changes. This observational data is then encrypted and sent to a server.

[0782] Step 2:

[0783] The device sends observational data to its built-in emotion engine, which then analyzes the user's emotional state. Here, observational data is used as input, and the data is processed using machine learning algorithms. The emotion engine outputs an emotion label based on facial features and voice tone, and sends this information to the server.

[0784] Step 3:

[0785] The server receives observational data transmitted from the terminal and analyzes it using a generative AI model. The input includes observational data and emotion labels. Through this analysis, the server processes the video data and identifies obstacles and landmarks in the surrounding environment. The output provides basic information for a safe travel route for the user.

[0786] Step 4:

[0787] The server generates user-directed navigation information based on the results of emotion analysis and environmental analysis. Input includes the user's emotional state and environmental details. A generative AI model is used to create guidance tailored to the user's emotions using prompts. These prompts may include instructions such as, "Please provide guidance in a slow voice to help the user relax." The final output is personalized guidance information, including voice guidance and visual guidelines.

[0788] Step 5:

[0789] The terminal receives guidance information sent from the server and presents it to the user. The input includes guidance information, which is then delivered to the user in natural-sounding voice using a speech synthesis engine. This also includes specific actions such as displaying visual guidelines via an AR display. Users can navigate with confidence based on the guidance information they receive. Feedback is also collected simultaneously and used for optimization in future applications.

[0790] (Application Example 2)

[0791] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0792] Conventional navigation systems struggle to provide personalized assistance based on the user's emotional state. Therefore, ensuring both emotional comfort and safety for passengers is a challenge, especially in autonomous vehicles. A travel experience that doesn't reflect the user's emotional state can cause unnecessary stress and anxiety.

[0793] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0794] In this invention, the server includes means for acquiring visual data obtained by a computing device worn by the user, means for transmitting the visual data from the computing device to a processing device, and means for analyzing the visual data using a generative model in the processing device to identify surrounding objects and reference points. This makes it possible to adjust the navigation guidance information based on the user's emotional state. Specifically, by flexibly adjusting the navigation route and output information according to the user's emotions, it is possible to provide passengers with a safer and more comfortable travel experience.

[0795] A "user" refers to an individual or group that uses this system, and is the subject whose emotional state is analyzed.

[0796] A "computational device" refers to a device worn by a user that has the function of acquiring and transmitting visual data.

[0797] "Visual data" refers to image and video information captured by the computing device, and is fundamental data for analyzing the user's emotional state.

[0798] A "processing unit" refers to a central computer that receives visual data transmitted from the computing unit and performs analysis using a generative model.

[0799] A "generative model" is a set of algorithms used in data analysis that perform intelligent processing to identify objects and reference points from visual data.

[0800] An "object" is a specific object that exists in the user's surrounding environment and is identified by the generative model.

[0801] A "reference point" refers to a landmark or marker used to determine the user's location in their surrounding environment, and is utilized in generating navigation information.

[0802] "Navigation guidance information" refers to guidelines and advice generated to assist users in their travels, and is adjusted based on the user's emotional state.

[0803] "Emotional state" refers to the emotional reactions and situations exhibited by the user, and is captured through emotion recognition sensors and visual data analysis.

[0804] To implement this invention, the user wears smart glasses or a smartphone as a computing device, thereby acquiring visual data. The device incorporates a camera and emotion recognition sensor to understand the user's emotions, and these are used to analyze video and audio data in real time.

[0805] The server receives encrypted visual data transmitted from the terminal and uses a generative model to identify surrounding objects and reference points. Based on this, it optimizes the user's movement path and generates appropriate navigation information. This navigation information is adjusted according to the user's emotional state; for example, if the user is feeling anxious, the server will explain the situation in a calm tone to support a more reassuring journey.

[0806] For example, if a passenger in an autonomous vehicle shows signs of anxiety, the vehicle could choose a more comfortable route and play calming music. In this way, support that takes user emotions into consideration can significantly improve the quality of the travel experience.

[0807] An example of a prompt would be, "What environmental conditions should you set based on the user's emotional state?" This serves as a guide for the generative AI model to learn the optimal guidance based on emotional states.

[0808] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0809] Step 1:

[0810] The device collects the user's visual and audio data. The device's built-in camera and emotion recognition sensor are activated, acquiring the user's facial expressions and audio information in real time. Through this process, the device obtains basic data as input to determine the user's emotional state.

[0811] Step 2:

[0812] The device securely and efficiently transmits the collected visual and audio data to the server. The data is encrypted and sent to the server via wireless communication. This allows the server to receive the data necessary for analysis while protecting privacy.

[0813] Step 3:

[0814] The server uses a generative AI model to analyze the received visual data and identify surrounding objects and reference points. It extracts features from the input data and recognizes obstacles and landmarks based on the generative model. As output, it provides detailed information about the user's environment.

[0815] Step 4:

[0816] The server integrates the user's emotional state with surrounding information to generate optimal navigation guidance. Based on the emotion recognition results, it determines the instructions that correspond to the user's emotions. For example, if the user is feeling anxious, it generates a more relaxing route or voice guidance. In this step, the generated guidance information is output.

[0817] Step 5:

[0818] The terminal displays guidance information received from the server to the user. This information is conveyed to the user via audio or visual means. The terminal continuously monitors the user's reactions and provides guidance in real time. This process allows the user to intuitively understand the situation and take the optimal action.

[0819] Step 6:

[0820] The user sends feedback to the server while on the move. By sending changes in emotion and responses to guidance as input to the server, the system's feedback loop is completed. This step allows the server to optimize the entire system, including the generative AI model, to improve the accuracy of future assistance.

[0821] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0822] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0823] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0824] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0825] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0826] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0827] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0828] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0829] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0830] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0831] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0832] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0833] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0834] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0835] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0836] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0837] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0838] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0839] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0840] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0841] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0842] The following is further disclosed regarding the embodiments described above.

[0843] (Claim 1)

[0844] A means for acquiring video data captured by a computer device worn by a user,

[0845] Means for transmitting video data acquired from the aforementioned computer device to a central processing unit,

[0846] The central processing unit includes means for analyzing the video data using a generation model and identifying surrounding obstacles and landmarks,

[0847] A means for generating information to guide the user's movement path based on the aforementioned analysis results,

[0848] Means for transmitting the generated information to the computer device and presenting it to the user,

[0849] A system that includes this.

[0850] (Claim 2)

[0851] The system according to claim 1, wherein the information to be presented to the user is provided by voice.

[0852] (Claim 3)

[0853] The system according to claim 1, wherein the central processing unit receives feedback from a user and updates the generation model based on that feedback.

[0854] "Example 1"

[0855] (Claim 1)

[0856] A means for acquiring image information captured by an information processing device wearable by a user,

[0857] means for transferring image information acquired from the aforementioned information processing device to a central processing device,

[0858] The central processing unit includes means for analyzing the image information using a generation model and detecting surrounding obstacles and landmarks,

[0859] A means for generating information to instruct the user's movement path based on the aforementioned analysis results,

[0860] Means for transmitting the generated information to the information processing device and presenting it to the user visually or audibly,

[0861] A means for obtaining information from the user and dynamically updating the generative model based on the obtained information,

[0862] A system that includes this.

[0863] (Claim 2)

[0864] The system according to claim 1, wherein the information presented to the user is provided by synthesized speech.

[0865] (Claim 3)

[0866] The system according to claim 1, wherein the information processing device performs a process to efficiently compress and encrypt the image information to be transmitted.

[0867] "Application Example 1"

[0868] (Claim 1)

[0869] A means for acquiring video data captured by a computer device worn by a user,

[0870] Means for transmitting video data acquired from the aforementioned computer device to a central processing unit,

[0871] The central processing unit includes means for analyzing the video data using a generation model and identifying surrounding obstacles and landmarks,

[0872] A means for generating information to guide the movement path of the transporter based on the analysis results,

[0873] Means for transmitting the generated information to the computer device and presenting it to the transporter,

[0874] A system that includes this.

[0875] (Claim 2)

[0876] The system according to claim 1, wherein the information to be presented to the carrier is provided by voice.

[0877] (Claim 3)

[0878] The system according to claim 1, wherein the central processing unit receives feedback from the user and updates the generation model based on that feedback.

[0879] "Example 2 of combining an emotion engine"

[0880] (Claim 1)

[0881] A means for acquiring observation data captured by a terminal device worn by the user,

[0882] Means for transmitting observation data acquired from the terminal device to a server device,

[0883] The aforementioned terminal device includes means for performing emotion analysis and analyzing the acquired data,

[0884] The server device includes means for analyzing the observation data using a generated AI model and identifying elements within the environment,

[0885] Means for generating information to provide user navigation guidance based on the results of the aforementioned analysis and interpretation,

[0886] Means for transmitting the generated information to the terminal device and presenting it to the user,

[0887] A system that includes this.

[0888] (Claim 2)

[0889] The system according to claim 1, wherein the information to be presented to the user is provided both audibly and visually using a speech synthesis engine.

[0890] (Claim 3)

[0891] The system according to claim 1, wherein the server device receives feedback from the user and updates and optimizes the generated AI model and sentiment analysis based on that feedback.

[0892] "Application example 2 when combining with an emotional engine"

[0893] (Claim 1)

[0894] A means for acquiring visual data obtained by a computing device worn by the user,

[0895] Means for transmitting visual data acquired from the aforementioned computing device to a processing device,

[0896] The processing apparatus includes means for analyzing the visual data using a generation model and identifying surrounding objects and reference points,

[0897] A means for generating information to guide the user's movement path based on the aforementioned analysis results,

[0898] Means for transmitting the generated information to the computing device and presenting it to the user,

[0899] A means for analyzing the user's emotional state and adjusting the output information accordingly,

[0900] A system that includes this.

[0901] (Claim 2)

[0902] The system according to claim 1, wherein the information to be presented to the user is provided by sound.

[0903] (Claim 3)

[0904] The system according to claim 1, wherein the processing device receives feedback from the user, updates the generation model based on that feedback, and optimizes the individualization of the output information. [Explanation of symbols]

[0905] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring video data captured by a computer device worn by a user, Means for transmitting video data acquired from the aforementioned computer device to a central processing unit, The central processing unit includes means for analyzing the video data using a generation model and identifying surrounding obstacles and landmarks, A means for generating information to guide the user's movement path based on the aforementioned analysis results, Means for transmitting the generated information to the computer device and presenting it to the user, A system that includes this.

2. The system according to claim 1, wherein the information to be presented to the user is provided by voice.

3. The system according to claim 1, wherein the central processing unit receives feedback from the user and updates the generation model based on that feedback.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A