system
The system addresses the challenge of inefficient travel planning by using image recognition to identify destinations and generate customized travel plans, enhancing user experience through automated and personalized route generation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Users face difficulties in identifying travel destinations from images and creating efficient travel plans that align with their preferences, leading to time-consuming and cumbersome processes.
A system that utilizes image recognition technology to identify geographical locations from user-uploaded images, generates optimal travel routes and schedules based on user instructions, and provides customized suggestions based on past preferences.
Enables efficient and customized travel planning by automatically generating travel routes and schedules from images, allowing users to easily identify destinations and make adjustments, while incorporating personalized recommendations.
Smart Images

Figure 2026047890000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When deciding on a travel destination, although the user wants to go to the places seen in the images on the Internet or on SNS, it is difficult to identify the places, and it also takes time to create a schedule. In addition, since there is a lack of a mechanism to automatically propose a travel plan that suits the user's preferences, there is a problem that the efficiency of travel planning is reduced. In such a situation, there is a need for a system for the user to identify a travel destination based on an image and create an efficient travel plan.
Means for Solving the Problems
[0005] This invention is a system that generates an optimal travel route based on the user's instructions and presents a detailed schedule by having the user input an image of a place they want to go and using an analysis means to identify the geographical location from that image. Furthermore, it is characterized by the use of image recognition technology such as Google's "Gemini" to accurately identify the geographical location from the image. In addition, it includes means for the user to set the duration of stay at each location and means for making additional suggestions based on the user's preferences, thereby providing the user with an optimal travel plan.
[0006] An "image" is visual information expressed in digital data or analog format, and includes visual representations of landscapes, buildings, people, and so on.
[0007] "Input means" refers to an interface for users to provide images and text to the system, and includes terminals such as smartphones and personal computers.
[0008] "Analysis means" refers to the technologies and functions used to identify geographical location from input images, and is a component of a system that includes image recognition technology.
[0009] "Plan generation means" refers to a function of a system that creates the optimal travel route based on analyzed geographical location information and user instructions.
[0010] A "presentation method" is an interface that presents the generated travel route and detailed schedule to the user in visual or text format.
[0011] "Suggestion method" refers to a system function that suggests additional travel plans and destinations based on the user's preferences.
[0012] "Geographic location" refers to information that indicates a specific place on Earth, and is expressed in forms such as latitude, longitude, or address.
[0013] A "detailed schedule" is a specific timetable that includes departure times, arrival times, length of stay, and mode of transportation at each stage of the trip.
[0014] A "travel route" refers to the sequence of points of passage and the path taken from the starting point to the destination.
[0015] "Preferences" refer to the types of places, landscapes, and facilities that users are particularly interested in and would like to visit. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. The system of this invention includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[0038] System Overview
[0039] 1. Image Input and Analysis
[0040] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[0041] 2. Plan generation and user specification
[0042] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[0043] 3. Generating the optimal route and detailed schedule
[0044] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[0045] 4. Customize your plan
[0046] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[0047] 5. Suggestions based on user preferences
[0048] The server analyzes the user's preferences based on their past usage history and current travel plan, and incorporates this into future recommendations. This enables customized suggestions tailored to the user's preferences.
[0049] Explain the program's processing in natural language.
[0050] 1. Image input and analysis
[0051] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[0052] 2. The device sends the uploaded image data to the server.
[0053] 3. The server uses image recognition technology to analyze specific locations within the image.
[0054] 4. The server saves the analysis results to the database.
[0055] 2. Creating a travel plan
[0056] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[0057] 2. The terminal sends instructions from the user to the server.
[0058] 3. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and destination XX, and generates a detailed schedule.
[0059] 4. The server sends the generated plan to the user's terminal.
[0060] 3. Plan presentation and customization
[0061] 1. The terminal displays the detailed travel plan received from the server to the user.
[0062] 2. The user reviews the proposed plan and instructs the user to make changes as needed (e.g., add a route through a local products fair).
[0063] 3. The terminal sends a change instruction to the server.
[0064] 4. The server updates the plan, taking the changes into account, and sends it back to the user's device.
[0065] 4. Additional proposals and historical analysis
[0066] 1. The server analyzes the user's preferences based on the database and incorporates them into suggestions for the next travel plan.
[0067] 2. When users receive new suggestions, it becomes easier for them to explore further preferred locations and routes.
[0068] Specific example
[0069] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[0070] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0071] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[0072] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[0073] 5. The server sends the generated plan to the user, who then confirms it.
[0074] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[0075] 7. The server is updated, the plan is regenerated, and sent to the user's terminal.
[0076] The above describes the specific embodiments and program processing details for implementing the present invention.
[0077] The following describes the processing flow.
[0078] Step 1:
[0079] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[0080] Step 2:
[0081] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[0082] Step 3:
[0083] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[0084] Step 4:
[0085] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[0086] Step 5:
[0087] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[0088] Step 6:
[0089] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[0090] Step 7:
[0091] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[0092] Step 8:
[0093] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[0094] Step 9:
[0095] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[0096] Step 10:
[0097] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[0098] Step 11:
[0099] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[0100] Step 12:
[0101] The server stores the user's past usage history and current plan data in a database and analyzes the user's preferences. This allows for the creation of customized travel suggestions that are tailored to the user's tastes.
[0102] (Example 1)
[0103] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0104] Traditional travel planning systems had the problem of requiring users to manually input detailed information about the places they wanted to visit, which was time-consuming. Furthermore, the generation of travel routes and schedules relied heavily on user input, lacking automation. As a result, users were unable to create travel plans efficiently, resulting in a time-consuming and cumbersome process. Additionally, the lack of customized suggestions based on user preferences made it difficult to improve user satisfaction.
[0105] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0106] In this invention, the server includes means for inputting an image, means for analyzing the input image to identify a geographical location, means for generating a plan that generates a travel route based on the analyzed geographical location information and user instructions, means for presenting the generated travel route and detailed schedule to the user, and means for making additional suggestions based on the user's preferences. This makes it possible for users to easily upload an image and have a travel route and detailed schedule automatically generated, enabling the provision of efficient and customized travel plans.
[0107] "Means of inputting images" refers to the function that allows users to upload image data to the system via their device.
[0108] "Means for analyzing input images to determine geographical location" refers to technologies that analyze uploaded image data to identify geographical information (e.g., longitude and latitude) of the location shown in the image.
[0109] The "means for generating plans" refer to a function that automatically creates the optimal travel route and detailed schedule for visiting multiple locations, based on analyzed geographical location information and user instructions.
[0110] "Means of presentation" refers to functions that visually display the generated travel route and detailed schedule to the user.
[0111] The "suggestion method" refers to a function that provides additional suggestions tailored to the user's preferences, based on the user's past travel history and current travel plan.
[0112] "Image recognition technology" is a technique that uses machine learning and algorithms to identify specific objects or locations within an image.
[0113] "A means for users to set the duration of their stay at each location" refers to a function that allows users to set the number of days or hours they will stay at each designated location.
[0114] This invention relates to a system that identifies a desired destination based on images provided by a user and generates a detailed travel route and schedule. The system includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[0115] First, the user uploads an image of the place they want to go to their device (such as a smartphone or PC). The device then sends this image data to the server. The server uses image recognition technology (such as "Gemini") to analyze the geographical location of the uploaded image and stores the analysis results in a database.
[0116] Next, the user uses text input to specify their desired destination, indicating that they want to travel via designated locations, and to specify the duration of their stay at each point. Upon receiving these instructions, the terminal sends the user's information to the server. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and then to the final destination XX, generating a detailed schedule that takes into account transportation methods and travel time. The generated plan is immediately presented to the user's terminal via the internet.
[0117] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can enter that information. The terminal sends this information to the server, which then updates the plan.
[0118] Furthermore, the server can analyze the user's preferences based on their past usage history and current travel plans, and reflect this in future travel suggestions. This enables customized suggestions tailored to the user's preferences.
[0119] Specific example
[0120] 1. The user uploads a landscape photo of the foot of Mt. Fuji and a night view photo of Tokyo Tower to the system.
[0121] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0122] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[0123] 4. The terminal sends instruction information to the server, which generates the optimal route and detailed schedule.
[0124] 5. The server sends the generated plan to the user's terminal, and the user confirms it.
[0125] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[0126] 7. The server regenerates the updated plan and sends it to the user's terminal.
[0127] Examples of prompts to input into a generative AI model
[0128] A user has uploaded landscape photos of the foot of Mt. Fuji and nighttime photos of Tokyo Tower to the system. Please use these photos to create a travel plan to Kyoto.
[0129] I stayed for one day at the foot of Mt. Fuji.
[0130] I stayed at Tokyo Tower for two days.
[0131] The final destination of the trip is Kyoto
[0132] I'd like to stop by the Osaka product fair on my way home.
[0133] Generate the optimal route and detailed schedule, and present them to the user.
[0134] Thus, this invention enables users to easily upload images, and the server automatically generates travel routes and detailed schedules. Furthermore, by providing customized suggestions based on user preferences, it realizes a system that offers more satisfying travel plans.
[0135] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0136] Step 1:
[0137] The user accesses the system and uploads an image of the place they want to go to their device.
[0138] Input: Image of the place you want to go (e.g., a landscape photo of Mt. Fuji)
[0139] Specific operation: The user accesses the system's web page or application and clicks the image upload button. The upload screen appears, and the user selects an image file from their device and sends it to the server.
[0140] Step 2:
[0141] The device sends the uploaded image data to the server.
[0142] Input: Image data uploaded by the user
[0143] Specific operation: The terminal divides the image data into packets and sends them to the server over the network.
[0144] Step 3:
[0145] The server uses image recognition technology to analyze the geographical location of the input image.
[0146] Input: Image data sent to the server
[0147] Specific operation: The server uses image recognition technology (e.g., "Gemini") to analyze images and identify specific landmarks or geographical features within them. The identified geographical locations (e.g., longitude and latitude) are recorded in a database.
[0148] Step 4:
[0149] The server saves the analysis results to the database.
[0150] Input: Geographic location information obtained using image recognition technology
[0151] Specific operation: The server connects to the database and saves the identified geographic coordinates and related information to the "Location" table.
[0152] Step 5:
[0153] The user provides text input indicating that they want to travel to their destination via specific locations and specifies the length of stay at each point.
[0154] Input: User instructions in text format (Example: "I want to spend one day at Mt. Fuji, two days at Tokyo Tower, and then go to my final destination, Kyoto.")
[0155] Specific operation: The user enters the desired location and length of stay into the system's text input field and clicks the submit button.
[0156] Step 6:
[0157] The terminal sends instructions from the user to the server.
[0158] Input: User instructions
[0159] Specific operation: The terminal sends user instructions in text format to the server.
[0160] Step 7:
[0161] The server searches for the optimal route from the starting point to each identified point and the final destination, and generates a detailed schedule.
[0162] Input: User-specified geographical location and length of stay
[0163] Specific operation: The server uses map databases and traffic information to calculate the optimal route from the starting point (current location) to each specified point and the final destination. It then creates a planned schedule, taking into account the length of stay.
[0164] Step 8:
[0165] The server sends the generated plan to the user's device.
[0166] Input: Detailed travel plan (route and schedule)
[0167] Specific operation: Format the generated plan and send it to the user's terminal via the internet.
[0168] Step 9:
[0169] The terminal displays the detailed travel plan received from the server to the user.
[0170] Input: Travel plan sent from the server
[0171] Specific operation: Visually display the travel route and schedule on the device screen (web page or application).
[0172] Step 10:
[0173] The user reviews the proposed plan and is instructed to make changes as needed.
[0174] Input: Instructions for the change (Example: I want to stop by the Osaka product exhibition on my way home)
[0175] Specific actions: The user reviews the presented plan, enters any necessary changes in text, and clicks the submit button.
[0176] Step 11:
[0177] The terminal sends a change instruction to the server.
[0178] Input: Instructions for change
[0179] Specific action: The terminal sends the user's change instruction to the server.
[0180] Step 12:
[0181] The server updates the plan, taking the changes into account, and sends it back to the user's device.
[0182] Input: Changed instructions
[0183] Specific operation: The server recalculates the travel plan based on the new instructions and sends the updated plan to the user's terminal.
[0184] Step 13:
[0185] The server analyzes user preferences based on a database and incorporates them into suggestions for the next travel plan.
[0186] Input: User's past usage history and current travel plan
[0187] Specific operation: The server analyzes historical data in the database, extracts user preferences and trends, and performs calculations to reflect these in future travel suggestions.
[0188] Step 14:
[0189] When a user receives new suggestions, they will explore further preferred locations and routes.
[0190] Input: New travel suggestions from the server
[0191] Specific operation: The user receives suggestions and considers new travel destinations and routes that suit their preferences.
[0192] (Application Example 1)
[0193] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0194] Traditional food delivery services struggle to identify specific dishes users want, lacking ways to improve the user experience. They also lack efficient support for users searching for specific cuisines or restaurants while traveling or out and about. This results in users having to spend a lot of time choosing meals and deciding on delivery plans, leading to reduced convenience.
[0195] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0196] In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated route and detailed schedule to the user, a suggestion means for making additional suggestions based on the user's preferences, a means for identifying a specific dish and serving location from a food image, and a means for generating an optimal route and delivery plan from the serving location. This makes it possible for the user to identify the restaurant they want to go to or the dish they want to eat using an image, and for the server to efficiently suggest the location and the optimal delivery plan.
[0197] "Image input means" refers to hardware or software functions that allow a user to upload images to a system.
[0198] "Analysis means" refers to technologies for identifying geographical location information and specific dishes from input images.
[0199] A "plan generation means" is a system function for generating routes and delivery plans based on analyzed information and user instructions.
[0200] "Presentation means" refers to methods or devices for displaying the generated plan or detailed schedule to the user.
[0201] The "suggestion method" is a function that makes additional suggestions based on the user's past usage history and preferences.
[0202] "Specific dish" refers to the type or name of food that can be identified from images uploaded by users.
[0203] "Place of service" refers to the location of the restaurant or delivery service where the specific dish is served.
[0204] A "delivery plan" is a plan that includes the optimal delivery route and time from the user's current location to the delivery destination.
[0205] This invention relates to a system that identifies specific locations or dishes based on images input by a user, and efficiently generates and presents plans and delivery plans based on those locations. This system consists of an image input means, an analysis means, a plan generation means, a presentation means, and a suggestion means.
[0206] System Overview
[0207] 1. Image input and analysis
[0208] Users access the system using a smartphone or PC and upload images of restaurants they want to visit or dishes they want to eat. The device sends this image data to the server. The server uses image recognition technology, such as Google's "Gemini," to analyze the specific locations and dishes in the input images and stores that information in a database.
[0209] 2. Plan generation and user specification
[0210] The user instructs the server via text input, "I want the food in the uploaded photo delivered." They also specify the length of stay at each location and the desired delivery time. The device then sends this information to the server.
[0211] 3. Generating the optimal route and detailed schedule
[0212] The server generates the optimal route from the user's current location or starting point to a restaurant serving a specific dish. Considering travel time and transportation methods, the server calculates a detailed delivery plan. The generated plan is immediately presented to the user's device.
[0213] 4. Customize your plan
[0214] Users can review the presented delivery plan and make changes as needed. For example, they can select a different restaurant or add information to their route home if they want to stop at a specific location. The device sends this information to the server, which then updates the plan.
[0215] 5. Suggestions based on user preferences
[0216] The server analyzes the user's past usage history and current plan to understand their preferences and incorporates these into future recommendations. This allows for customized recommendations tailored to the user's tastes.
[0217] Specific processing instructions
[0218] 1. Image input and analysis:
[0219] Users access the system from their smartphones or PCs and upload images of food or restaurants. The device sends the uploaded image data to the server, which analyzes the images using tools such as Google's "Gemini" and stores the information in a database.
[0220] 2. Plan generation and user specification:
[0221] The user instructs the server via text input, "I want the food in the picture delivered," and this information is sent from the terminal to the server. The server then generates the optimal plan based on the restaurants and delivery services that offer the specific dish.
[0222] 3. Generating the optimal route and detailed schedule:
[0223] The server calculates the optimal route from the user's current location to the restaurant and generates a detailed schedule. The generated plan is then displayed on the terminal.
[0224] 4. Customize your plan:
[0225] The user reviews the presented plan and makes changes as needed. The device sends the change information to the server, which then generates and presents the updated plan again.
[0226] 5. Additional proposals and historical analysis:
[0227] The server analyzes the user's past usage history and preferences based on a database to optimize future recommendations.
[0228] Specific example
[0229] For example, if a user uploads a photo of fried rice, the server analyzes the photo and identifies it as fried rice. The user then enters "I want fried rice delivered," and the server generates the optimal delivery plan from a specific restaurant. The user can review this plan and request changes if necessary. The server also generates a plan that reflects any stops along the return journey.
[0230] Example of a prompt
[0231] "Upload a photo of fried rice and suggest restaurants that offer delivery of that dish."
[0232] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0233] Step 1:
[0234] Uploading and sending images
[0235] Users log in to the system using their smartphones or PCs and upload images of the food they want to eat or the restaurants they want to visit. The device receives the image data sent by the user and sends it to the server. The input here is the image taken by the user, and the output is the image data sent to the server.
[0236] Step 2:
[0237] Image analysis
[0238] The server uses image recognition technologies such as Google's "Gemini" to analyze the received image data. In this process, the image data serves as input, and features and content within the image are recognized during the analysis process, identifying specific dishes or locations. The analysis results are stored in a database as geographical location information and information about the specific dish.
[0239] Step 3:
[0240] Text input and sending instructions
[0241] The user enters a text instruction via the system interface, such as "I want the food in the uploaded image delivered." They can also enter additional information, such as their desired delivery time. The terminal then sends these instructions to the server. The input is the user's text instruction, and the output is the data sent to the server.
[0242] Step 4:
[0243] Generating a delivery plan
[0244] The server generates the optimal delivery route and schedule based on the user's current location and information on the nearest restaurants that serve the specified dish. Analyzed information and user instructions serve as input, and the optimal route and detailed schedule are output. This calculation also utilizes data such as transportation methods and travel time.
[0245] Step 5:
[0246] Presentation of the plan
[0247] The generated delivery plan and detailed schedule are sent from the server to the terminal, which then displays them to the user. The input here is the plan data generated on the server side, and the output is the plan information displayed on the user's terminal.
[0248] Step 6:
[0249] Customize your plan
[0250] The user reviews the presented delivery plan and makes changes as needed, such as selecting a different restaurant or specifying additional transit points. The terminal sends the user's change instructions to the server, which then regenerates the plan based on this information. The input is the user's change instructions, and the output is the updated plan.
[0251] Step 7:
[0252] Presentation and approval of the final plan
[0253] The server generates the updated plan and sends it back to the user's device. The user approves the final plan, and the plan is finalized. The input here is the new plan data, and the output is the final plan display and approval on the user's device.
[0254] Step 8:
[0255] Additional proposals and historical analysis
[0256] The server makes delivery suggestions for future deliveries based on the user's past usage history and preferences. In this process, past usage data is used as input, and suggestion data based on the user's preferences is output. The suggested content will be presented to the user the next time they use the service.
[0257] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0258] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. Furthermore, this invention incorporates an emotion engine that recognizes the user's emotions and includes a function to adjust the plan according to the user's emotional state. The system of this invention includes image input means, analysis means, plan generation means, presentation means, suggestion means, and an emotion engine.
[0259] System Overview
[0260] 1. Image Input and Analysis
[0261] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[0262] 2. Plan generation and user specification
[0263] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[0264] 3. Generating the optimal route and detailed schedule
[0265] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[0266] 4. User state recognition using an emotion engine
[0267] This invention incorporates an emotion engine that analyzes the user's facial expressions and voice, and incorporates this into the travel plan. The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits this data to a server. The server analyzes the collected data and recognizes the user's emotional state.
[0268] 5. Adjusting plans based on emotions
[0269] Based on the perceived emotions, the server adjusts travel routes and itineraries in real time. For example, if a user appears tired, the plan can be changed to include more rest periods.
[0270] 6. Plan presentation and customization
[0271] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[0272] 7. Use of long-term emotional data
[0273] The server records the user's emotional state over the long term and uses this history to optimize future travel suggestions. This historical data is used to further refine the user's preferred types of travel plans.
[0274] Explain the program's processing in natural language.
[0275] 1. Image input and analysis
[0276] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[0277] 2. The device sends the uploaded image data to the server.
[0278] 3. The server uses image recognition technology to analyze specific positions in the image.
[0279] 4. The server identifies the geographical location based on the analysis result and saves this information in the database.
[0280] 2. Generation of travel plan
[0281] 1. The user gives an instruction in text input saying "I want to go to ○○ via Photo A and Photo B" and specifies the number of days of stay at each location.
[0282] 2. The terminal sends the instruction from the user to the server.
[0283] 3. The server organizes information about the departure point, each specific location, and the destination, and generates an optimal route.
[0284] 4. The server calculates the means of transportation and the required time, and determines a detailed schedule.
[0285] 5. The server sends the generated travel plan to the user's terminal.
[0286] 3. Recognition of user state by the emotion engine
[0287] 1. The terminal uses the built-in camera and microphone to collect the user's facial expressions and voice data.
[0288] 2. The terminal sends the collected data to the server. <0 2. For example, if a user is feeling stressed, the system can regenerate a plan to visit places or facilities where they can relax.
[0293] 3. The server sends the adjusted plan to the user's terminal.
[0294] 5. Plan presentation and customization
[0295] 1. The user reviews the travel plan presented via their device and is instructed to make changes as needed.
[0296] 2. The terminal sends the user's change instruction to the server.
[0297] 3. The server updates the plan to reflect the changed information and sends it to the user's device again.
[0298] 6. Use of long-term emotional data
[0299] 1. The server records the user's emotional state over the long term and stores it in a database.
[0300] 2. The server optimizes future travel suggestions based on accumulated historical data and analyzes user preferences in a more precise manner.
[0301] Specific example
[0302] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[0303] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0304] 3. The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[0305] 4. The terminal sends the instruction information to the server, and the server generates an optimal route and a detailed schedule.
[0306] 5. The server sends the plan generated to the user, and the user checks it.
[0307] 6. The terminal uses the built-in camera and microphone to collect the user's expressions and voices when checking the plan, and sends them to the server.
[0308] 7. The server analyzes the user's emotions with an emotion engine and adjusts the plan. For example, if the user is judged to be tired, change the plan to increase the rest time.
[0309] 8. The user inputs that they want to pass by the Osaka物产展 on the way back, and the terminal sends it to the server.
[0310] 9. The server regenerates the changed plan and sends it to the user's terminal.
[0311] 10. The server records the user's emotional state in the long term and uses it for the improvement of the next travel proposal.
[0312] The above are the specific forms and the processing contents of the program for implementing the present invention.
[0313] The following explains the processing flow.
[0314] Step 1:
[0315] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images from a terminal such as a smartphone or a PC through the interface of the system and presses the upload button.
[0316] Step 2:
[0317] Note: The "物产展" in the original text seems to be a specific name that might need to be accurately translated according to the actual context. Here it is left as it is for lack of more information.The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[0318] Step 3:
[0319] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[0320] Step 4:
[0321] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[0322] Step 5:
[0323] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[0324] Step 6:
[0325] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[0326] Step 7:
[0327] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[0328] Step 8:
[0329] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[0330] Step 9:
[0331] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[0332] Step 10:
[0333] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[0334] Step 11:
[0335] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[0336] Step 12:
[0337] The device uses its built-in camera and microphone to collect facial expressions and audio data as the user reviews the plan. This data is collected only with the user's permission.
[0338] Step 13:
[0339] The device sends collected facial and audio data to the server. The server uses an emotion engine to analyze the user's emotional state.
[0340] Step 14:
[0341] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. For example, if the user looks tired, it might increase the rest time.
[0342] Step 15:
[0343] The server sends the adjusted plan to the user's device. The user can review the new plan and confirm it again before starting their trip.
[0344] Step 16:
[0345] The server records the user's emotional state over the long term and stores it in a database. When planning the next trip, the server takes the user's past emotional data into consideration to provide more customized suggestions.
[0346] The above outlines the specific processing steps of the invention that combines an emotion engine.
[0347] (Example 2)
[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0349] Traditional travel planning systems have limitations in identifying desired destinations based on user-provided images and generating optimal travel routes that include those locations. Furthermore, they lack the functionality to adjust plans based on the user's emotional state during the trip, making it difficult to mitigate stress and dissatisfaction. Additionally, they haven't adequately provided mechanisms for optimizing future travel plans using long-term user emotional data.
[0350] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated travel route and detailed schedule to the user, an emotion engine means for analyzing the user's emotional state and adjusting the travel plan, and a suggestion means for making additional suggestions based on the user's preferences. This makes it possible to identify desired destinations and generate optimal travel routes based on images provided by the user, and furthermore, to adjust the travel plan in real time according to the user's emotional state and optimize future travel plans using long-term emotional data.
[0351] The "image input method" is a function that allows users to upload images of places they want to visit to the system.
[0352] "Analysis means" refers to a function for determining geographical location from input image data.
[0353] The "plan generation method" is a function that generates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[0354] "Presentation means" refers to a function for displaying the generated travel route and detailed schedule to the user.
[0355] The "emotional engine" is a function that analyzes the user's facial expressions and voice to adjust the travel plan in real time.
[0356] A "suggestion mechanism" is a function that provides additional suggestions based on the user's preferences.
[0357] "Image recognition technology" is a technology that analyzes the content of an input image and extracts specific information from it.
[0358] "Means for setting the length of stay" refers to a function that allows users to specify the number of days or hours they will stay at each visited location.
[0359] This invention is a system that identifies desired destinations based on images provided by the user and generates the optimal travel route and detailed schedule that passes through those locations. Furthermore, it incorporates a function that recognizes the user's emotions and adjusts the travel plan according to their emotional state. This invention includes the following main functions:
[0360] 1. Image input and analysis
[0361] Users upload images of places they want to visit to the system from their smartphones, PCs, or other devices. The device then sends this image data to the server. The server uses "image recognition technology" to analyze and identify the geographical location of the input image. For example, if a user uploads a photo of the foot of Mount Fuji, the server analyzes the photo to identify the foot of Mount Fuji.
[0362] 2. Creating a travel plan
[0363] The user specifies their desired destination, such as "I want to go to XX via photo A and photo B," through text input, and also specifies the length of stay at each location. The terminal receives this information and sends it to the server. The server generates the optimal route based on the information about the starting point, specific locations, and destination. For example, if the user wants to go from the foot of Mt. Fuji to Kyoto via Tokyo Tower, the server will calculate a detailed schedule considering the means of transportation and travel time at each location.
[0364] 3. Presentation of the generated plan
[0365] The server sends the generated travel plan to the user's device, which then displays it. The user can review the plan and customize it as needed.
[0366] 4. User state recognition using an emotion engine
[0367] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and sends this data to a server. The server uses an "emotion engine" to analyze the data and recognize the user's emotional state. For example, if the user is smiling or looks tired while reviewing a plan, the emotion engine will analyze that data.
[0368] 5. Adjusting plans based on emotions
[0369] Based on the perceived emotions, the server may adjust the travel route and itinerary in real time. For example, if the user appears tired, the plan may be changed to include more rest periods or visits to relaxing locations. The adjusted plan is then sent back to the user's device.
[0370] 6. Accumulation and utilization of long-term emotional data
[0371] The server records the user's emotional state over the long term and stores this data in a database. This accumulated emotional data can be used to optimize future travel suggestions and propose plans that are more tailored to the user's preferences.
[0372] Specific example
[0373] For example, if a user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system, the server analyzes these photos to identify Mt. Fuji and Tokyo Tower. The user then inputs, "I want to go to Kyoto via Photo A and Photo B," and specifies the length of stay at each location. The server generates the optimal route and creates a detailed schedule considering transportation methods and travel time. The generated plan is displayed on the terminal for the user to review. At this time, the terminal uses the camera and microphone to collect the user's facial expressions and voice, which the server analyzes with an emotion engine to adjust the plan. For example, if the user inputs that they want to stop by an Osaka product exhibition on the way back, the plan will be regenerated.
[0374] Example of a prompt
[0375] "Please generate the optimal travel plan based on the following instructions. I would like to travel to Kyoto, passing through the foot of Mt. Fuji and Tokyo Tower along the way. I will stay at the foot of Mt. Fuji for one day and at Tokyo Tower for two days."
[0376] "Based on user sentiment data, please suggest a day trip plan that includes plenty of breaks."
[0377] The above describes specific embodiments for carrying out the present invention.
[0378] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0379] Program processing flow and detailed explanation
[0380] Image input and analysis
[0381] Step 1:
[0382] Users upload images of places they want to visit to the system from their smartphones or PCs. Input is image files (e.g., JPEG or PNG format). Output is the image data being saved on the device. Specifically, the user uses the system's upload function, selects an image from a browser or dedicated app, and clicks the upload button.
[0383] Step 2:
[0384] The device sends uploaded image data to the server. The input is the image file data, and the output is the image data received by the server. The device uses an internet connection to send the image data to the server as packets. Specifically, it sends the image data using an HTTP request.
[0385] Step 3:
[0386] The server uses image recognition technology to analyze the input image. The input is the received image data, and the output is the geographical location information of the analysis result. Specifically, the server uses Google's "image recognition technology" to execute an algorithm that extracts image features and determines the geographical location.
[0387] Step 4:
[0388] The server identifies the geographic location based on the analysis results and stores this information in a database. The input is geographic location information, and the output is the geographic location information stored in the database. The server performs an insert operation into the database to persist the identified location data.
[0389] Travel plan generation
[0390] Step 1:
[0391] The user enters text such as "I want to go to XX via photo A and photo B" and specifies the length of stay at each location. The input is text data (intermediate points and length of stay). The output is the instruction information entered into the terminal. Specifically, the user enters the desired intermediate points and length of stay in the text input field and clicks the submit button.
[0392] Step 2:
[0393] The terminal sends user instructions to the server. The input is text data (intermediate locations and length of stay), and the output is the instructions received by the server. The terminal sends the instructions to the server using an HTTP request.
[0394] Step 3:
[0395] The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route. The input is the user's instruction information (waypoints, starting point, destination, length of stay), and the output is the generated optimal route information. Specifically, the server uses a Geographic Information System (GIS) to calculate the optimal route connecting each point using an algorithm.
[0396] Step 4:
[0397] The server calculates the mode of transport and travel time to determine a detailed schedule. The input is the generated optimal route information, and the output is a detailed schedule (mode of transport, departure time, arrival time, etc.). Specifically, the server uses a traffic information API to calculate the travel time for each segment of the journey.
[0398] Step 5:
[0399] The server generates a travel plan and sends it to the user's terminal. The input is a detailed schedule, and the output is the travel plan displayed on the terminal. The server sends the detailed schedule to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[0400] User state recognition using an emotion engine
[0401] Step 1:
[0402] The device uses its built-in camera and microphone to collect user facial expressions and audio data. The input is the user's facial expressions and voice, and the output is the collected data. Specifically, the device's app activates the camera, takes a picture of the user's face, and records audio using the microphone.
[0403] Step 2:
[0404] The terminal sends the collected data to the server. The input consists of facial expressions and voice data, and the output is the data received by the server. The terminal encrypts the data and sends it to the server via an HTTP request.
[0405] Step 3:
[0406] The server uses an emotion engine to analyze data and recognize the user's emotional state. The input is facial expression and voice data, and the output is recognized emotional state information. Specifically, the server applies an emotion recognition algorithm to analyze the user's emotions (e.g., joy, sadness, fatigue, etc.).
[0407] Adjusting plans based on emotions
[0408] Step 1:
[0409] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. The input is the recognized emotional state information, and the output is the adjusted travel plan. Specifically, the server makes adjustments such as changing the schedule to include more rest.
[0410] Step 2:
[0411] The server sends the adjusted plan to the user's terminal. The input is the adjusted travel plan, and the output is the adjusted plan displayed on the terminal. The server sends the adjusted plan to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[0412] Plan presentation and customization
[0413] Step 1:
[0414] The user reviews the presented travel plan. The input is the travel plan displayed on the device, and the output is the user's review process. Specifically, the user views the page displaying the travel plan.
[0415] Step 2:
[0416] If a user wishes to change their plan, they enter that information. The input is the instruction for the desired change, and the output is the change instruction entered into the terminal. For example, they might enter, "I would like to stop by the Osaka product exhibition on my way home."
[0417] Step 3:
[0418] The terminal sends the user's change instructions to the server. The input is the information of the desired change, and the output is the change instructions received by the server. The terminal sends the change instructions to the server using an HTTP request.
[0419] Step 4:
[0420] The server regenerates the plan to reflect the changed information. The input is the change instruction information, and the output is the regenerated plan. The server recalculates the optimal route and detailed schedule and generates a new plan.
[0421] Step 5:
[0422] The server sends the updated plan to the user's terminal. The input is the regenerated plan, and the output is the updated plan displayed on the terminal. The server sends the updated plan to the terminal via an HTTP response, which the terminal receives and displays.
[0423] Use of long-term emotional data
[0424] Step 1:
[0425] The server records the user's emotional state over the long term. The input is recognized emotional state information, and the output is emotional data stored in a database. The server performs insert operations into the database to save the emotional data.
[0426] Step 2:
[0427] The server optimizes the next travel suggestion based on accumulated historical data. The input is long-term sentiment data, and the output is an optimized travel suggestion. The server analyzes the sentiment data and proposes the best travel plan based on the user's priorities and preferences.
[0428] The above outlines the specific processing steps of the system.
[0429] (Application Example 2)
[0430] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0431] Traditional travel planning systems have been somewhat effective in creating travel plans that meet users' preferences, but they have a weakness in their inability to adjust plans to reflect real-time emotional states during the trip. Furthermore, means for users to smartly review their plans during their trip, and visually supportive devices, are still not fully utilized. Therefore, there is a need to reduce stress during travel and improve the user experience.
[0432] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0433] In this invention, the server includes means for inputting an image, analysis means for identifying a geographical location from the input image, plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, presentation means for presenting the generated travel route and detailed schedule to the user, suggestion means for making additional suggestions based on the user's preferences, means including an emotion engine that analyzes the user's facial expressions and voice data and adjusts the travel plan based on their emotional state, and means for presenting the travel plan to the user using a real-time adaptable mobile or visual device. This makes it possible to adjust the plan in real time to reflect the user's emotional state during the trip and to improve the user's travel experience through visual support from a smart device.
[0434] "Image input means" refers to a device or software that receives image data input by a user and transmits it to a server.
[0435] "Analysis means" refers to a device or software that identifies a specific geographical location from input image data and provides that information to a server.
[0436] A "plan generation means" refers to a device or software that creates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[0437] "Presentation means" refers to a device or software that provides the user with the generated travel route and detailed schedule visually or audibly.
[0438] "Suggestion means" refers to a device or software that provides additional information or recommended plans regarding travel routes based on the user's preferences.
[0439] "Means including an emotion engine" refers to a device or software that analyzes the user's facial expressions and voice data to recognize their emotional state and adjust the travel plan accordingly.
[0440] "Means for presenting a travel plan to a user using a mobile or visual device" refers to a device or software that visually presents a travel plan while reflecting the user's location information and emotional state in real time.
[0441] "Real-time" refers to a state where the user's current status and location information are reflected in real time, allowing for immediate adaptation or modification.
[0442] This invention is a system for efficiently creating travel plans and adjusting them in real time according to the user's emotional state. The system includes means for image input, analysis, plan generation, presentation, suggestion, an emotion engine, and means for presenting the travel plan to the user using a mobile device or visual device.
[0443] Description of the program's overall operation
[0444] 1. Image input and analysis
[0445] Users upload images of places they want to visit from their smartphones or PCs to the system in order to create travel plans. In this case, the image input mechanism is activated. The device sends these images to the server, which uses image recognition technology (e.g., a common image recognition API) to determine the geographical location of the images. The results of this analysis are stored in a database.
[0446] 2. Plan generation and user specification
[0447] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location. Based on this, the terminal sends the instructions to the server. The server generates the optimal travel route and detailed schedule based on the analyzed geographic location information and the user's instructions. This generation process uses a route calculation engine and transportation information.
[0448] 3. Collection and analysis of facial expressions and voice.
[0449] The device uses its built-in camera and microphone to collect user facial expressions and voice data. This data is sent to a server in real time, where the server uses an emotion engine to analyze the user's emotional state. Specifically, facial expressions are analyzed using facial recognition software (e.g., OpenCV or dlib), and the data is sent to an emotion analysis API.
[0450] 4. Adjusting plans based on emotions
[0451] Based on the analysis results from the emotion engine, the server adjusts the travel plan to match the user's emotional state. For example, if the user is tired, the plan will be changed to include relaxing facilities. Conversely, if the user is excited, the plan can be adjusted to include more active spots.
[0452] 5. Real-time plan presentation
[0453] The travel plan, generated in real time, is presented to the user through visual devices such as smart glasses. This allows the user to enjoy their trip while checking a plan that suits their emotional state in real time.
[0454] Specific example
[0455] The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a photo of an urban landmark (Photo B) to the system.
[0456] The server identifies photo A as the foot of Mount Fuji and photo B as an urban area, and saves them in the database.
[0457] The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[0458] The terminal sends instruction information to the server, which then generates the optimal route and a detailed schedule.
[0459] The device uses its built-in camera and microphone to collect the user's facial expressions and voice in real time as they review the plan, and transmits this information to the server.
[0460] The server uses an emotion engine to analyze the user's emotions and adjust the plan accordingly. For example, if it determines that the user is tired, it will change the plan to include more rest time.
[0461] Example of a prompt
[0462] "Since you seem tired this time, please suggest a travel plan that includes places where you can relax."
[0463] This system allows for real-time adjustments to travel plans that reflect the user's emotional state during their trip, and enhances the user's travel experience through visual support via smart devices.
[0464] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0465] Step 1:
[0466] The user uses their device to upload images of the place they want to go to the system.
[0467] The specific input is the user's image data, which the device sends to the server. The server uses image recognition technology to analyze the uploaded image. For example, it uses an image recognition API to determine the geographical location and stores this information in a database. The output is the analyzed geographical location information.
[0468] Step 2:
[0469] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location.
[0470] The input is user text instructions, which the terminal sends to the server. The server generates the optimal travel route and detailed schedule based on the starting point, analyzed geographic location information, and user instructions. It uses a route calculation engine to determine the best route, taking into account transportation methods and travel time. The output is the generated travel plan and detailed schedule.
[0471] Step 3:
[0472] The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[0473] The input consists of the user's real-time facial expressions and voice data, which the terminal collects and sends to the server. The server uses an emotion engine to analyze the data and recognize the user's emotional state. Specifically, it analyzes facial expressions using facial recognition software (e.g., OpenCV or dlib) and sends the data to an emotion analysis API to identify the emotional state. The output is the user's emotional state as determined by the emotion engine.
[0474] Step 4:
[0475] The server adjusts the travel plan in real time according to the user's emotional state, based on the analysis results from the emotion engine.
[0476] The input consists of the user's emotional state, as determined by the emotion engine, and the existing travel plan. Based on this data, the server regenerates the plan, including new suggestions and modifications. For example, if the user is tired, the server might add rest areas and modify the plan to include relaxing facilities. The output is the adjusted travel plan.
[0477] Step 5:
[0478] The device presents the user with travel plans that have been generated or adjusted in real time.
[0479] The input is a pre-configured travel plan, and the device visually presents this information to the user via smart glasses or a smartphone. The user can continue their trip while checking the plan in real time to match their emotional state. The output is the pre-configured travel plan displayed on the user's device.
[0480] Step 6:
[0481] The server records the user's emotional state over the long term to optimize future travel suggestions.
[0482] The input consists of past travel plans and user emotional state data, which the server stores in a database. This data is used to generate optimal travel plans based on the user's preferences for future trips. The output is a more personalized suggestion for the next trip.
[0483] Through these steps, users can enjoy travel plans tailored to their emotional state in real time, enhancing their travel experience.
[0484] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0485] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0486] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0487] [Second Embodiment]
[0488] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0489] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0490] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0491] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0492] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0493] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0494] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0495] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0496] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0497] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0498] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0499] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0500] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. The system of this invention includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[0501] System Overview
[0502] 1. Image Input and Analysis
[0503] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[0504] 2. Plan generation and user specification
[0505] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[0506] 3. Generating the optimal route and detailed schedule
[0507] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[0508] 4. Customize your plan
[0509] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[0510] 5. Suggestions based on user preferences
[0511] The server analyzes the user's preferences based on their past usage history and current travel plan, and incorporates this into future recommendations. This enables customized suggestions tailored to the user's preferences.
[0512] Explain the program's processing in natural language.
[0513] 1. Image input and analysis
[0514] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[0515] 2. The device sends the uploaded image data to the server.
[0516] 3. The server uses image recognition technology to analyze specific locations within the image.
[0517] 4. The server saves the analysis results to the database.
[0518] 2. Creating a travel plan
[0519] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[0520] 2. The terminal sends instructions from the user to the server.
[0521] 3. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and destination XX, and generates a detailed schedule.
[0522] 4. The server sends the generated plan to the user's terminal.
[0523] 3. Plan presentation and customization
[0524] 1. The terminal displays the detailed travel plan received from the server to the user.
[0525] 2. The user reviews the proposed plan and instructs the user to make changes as needed (e.g., add a route through a local products fair).
[0526] 3. The terminal sends a change instruction to the server.
[0527] 4. The server updates the plan, taking the changes into account, and sends it back to the user's device.
[0528] 4. Additional proposals and historical analysis
[0529] 1. The server analyzes the user's preferences based on the database and incorporates them into suggestions for the next travel plan.
[0530] 2. When users receive new suggestions, it becomes easier for them to explore further preferred locations and routes.
[0531] Specific example
[0532] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[0533] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0534] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[0535] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[0536] 5. The server sends the generated plan to the user, who then confirms it.
[0537] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[0538] 7. The server is updated, the plan is regenerated, and sent to the user's terminal.
[0539] The above describes the specific embodiments and program processing details for implementing the present invention.
[0540] The following describes the processing flow.
[0541] Step 1:
[0542] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[0543] Step 2:
[0544] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[0545] Step 3:
[0546] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[0547] Step 4:
[0548] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[0549] Step 5:
[0550] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[0551] Step 6:
[0552] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[0553] Step 7:
[0554] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[0555] Step 8:
[0556] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[0557] Step 9:
[0558] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[0559] Step 10:
[0560] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[0561] Step 11:
[0562] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[0563] Step 12:
[0564] The server stores the user's past usage history and current plan data in a database and analyzes the user's preferences. This allows for the creation of customized travel suggestions that are tailored to the user's tastes.
[0565] (Example 1)
[0566] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0567] Traditional travel planning systems had the problem of requiring users to manually input detailed information about the places they wanted to visit, which was time-consuming. Furthermore, the generation of travel routes and schedules relied heavily on user input, lacking automation. As a result, users were unable to create travel plans efficiently, resulting in a time-consuming and cumbersome process. Additionally, the lack of customized suggestions based on user preferences made it difficult to improve user satisfaction.
[0568] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0569] In this invention, the server includes means for inputting an image, means for analyzing the input image to identify a geographical location, means for generating a plan that generates a travel route based on the analyzed geographical location information and user instructions, means for presenting the generated travel route and detailed schedule to the user, and means for making additional suggestions based on the user's preferences. This makes it possible for users to easily upload an image and have a travel route and detailed schedule automatically generated, enabling the provision of efficient and customized travel plans.
[0570] "Means of inputting images" refers to the function that allows users to upload image data to the system via their device.
[0571] "Means for analyzing input images to determine geographical location" refers to technologies that analyze uploaded image data to identify geographical information (e.g., longitude and latitude) of the location shown in the image.
[0572] The "means for generating plans" refer to a function that automatically creates the optimal travel route and detailed schedule for visiting multiple locations, based on analyzed geographical location information and user instructions.
[0573] "Means of presentation" refers to functions that visually display the generated travel route and detailed schedule to the user.
[0574] The "suggestion method" refers to a function that provides additional suggestions tailored to the user's preferences, based on the user's past travel history and current travel plan.
[0575] "Image recognition technology" is a technique that uses machine learning and algorithms to identify specific objects or locations within an image.
[0576] "A means for users to set the duration of their stay at each location" refers to a function that allows users to set the number of days or hours they will stay at each designated location.
[0577] This invention relates to a system that identifies a desired destination based on images provided by a user and generates a detailed travel route and schedule. The system includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[0578] First, the user uploads an image of the place they want to go to their device (such as a smartphone or PC). The device then sends this image data to the server. The server uses image recognition technology (such as "Gemini") to analyze the geographical location of the uploaded image and stores the analysis results in a database.
[0579] Next, the user uses text input to specify their desired destination, indicating that they want to travel via designated locations, and to specify the duration of their stay at each point. Upon receiving these instructions, the terminal sends the user's information to the server. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and then to the final destination XX, generating a detailed schedule that takes into account transportation methods and travel time. The generated plan is immediately presented to the user's terminal via the internet.
[0580] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can enter that information. The terminal sends this information to the server, which then updates the plan.
[0581] Furthermore, the server can analyze the user's preferences based on their past usage history and current travel plans, and reflect this in future travel suggestions. This enables customized suggestions tailored to the user's preferences.
[0582] Specific example
[0583] 1. The user uploads a landscape photo of the foot of Mt. Fuji and a night view photo of Tokyo Tower to the system.
[0584] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0585] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[0586] 4. The terminal sends instruction information to the server, which generates the optimal route and detailed schedule.
[0587] 5. The server sends the generated plan to the user's terminal, and the user confirms it.
[0588] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[0589] 7. The server regenerates the updated plan and sends it to the user's terminal.
[0590] Examples of prompts to input into a generative AI model
[0591] A user has uploaded landscape photos of the foot of Mt. Fuji and nighttime photos of Tokyo Tower to the system. Please use these photos to create a travel plan to Kyoto.
[0592] I stayed for one day at the foot of Mt. Fuji.
[0593] I stayed at Tokyo Tower for two days.
[0594] The final destination of the trip is Kyoto
[0595] I'd like to stop by the Osaka product fair on my way home.
[0596] Generate the optimal route and detailed schedule, and present them to the user.
[0597] Thus, this invention enables users to easily upload images, and the server automatically generates travel routes and detailed schedules. Furthermore, by providing customized suggestions based on user preferences, it realizes a system that offers more satisfying travel plans.
[0598] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0599] Step 1:
[0600] The user accesses the system and uploads an image of the place they want to go to their device.
[0601] Input: Image of the place you want to go (e.g., a landscape photo of Mt. Fuji)
[0602] Specific operation: The user accesses the system's web page or application and clicks the image upload button. The upload screen appears, and the user selects an image file from their device and sends it to the server.
[0603] Step 2:
[0604] The device sends the uploaded image data to the server.
[0605] Input: Image data uploaded by the user
[0606] Specific operation: The terminal divides the image data into packets and sends them to the server over the network.
[0607] Step 3:
[0608] The server uses image recognition technology to analyze the geographical location of the input image.
[0609] Input: Image data sent to the server
[0610] Specific operation: The server uses image recognition technology (e.g., "Gemini") to analyze images and identify specific landmarks or geographical features within them. The identified geographical locations (e.g., longitude and latitude) are recorded in a database.
[0611] Step 4:
[0612] The server saves the analysis results to the database.
[0613] Input: Geographic location information obtained using image recognition technology
[0614] Specific operation: The server connects to the database and saves the identified geographic coordinates and related information to the "Location" table.
[0615] Step 5:
[0616] The user provides text input indicating that they want to travel to their destination via specific locations and specifies the length of stay at each point.
[0617] Input: User instructions in text format (Example: "I want to spend one day at Mt. Fuji, two days at Tokyo Tower, and then go to my final destination, Kyoto.")
[0618] Specific operation: The user enters the desired location and length of stay into the system's text input field and clicks the submit button.
[0619] Step 6:
[0620] The terminal sends instructions from the user to the server.
[0621] Input: User instructions
[0622] Specific operation: The terminal sends user instructions in text format to the server.
[0623] Step 7:
[0624] The server searches for the optimal route from the starting point to each identified point and the final destination, and generates a detailed schedule.
[0625] Input: User-specified geographical location and length of stay
[0626] Specific operation: The server uses map databases and traffic information to calculate the optimal route from the starting point (current location) to each specified point and the final destination. It then creates a planned schedule, taking into account the length of stay.
[0627] Step 8:
[0628] The server sends the generated plan to the user's device.
[0629] Input: Detailed travel plan (route and schedule)
[0630] Specific operation: Format the generated plan and send it to the user's terminal via the internet.
[0631] Step 9:
[0632] The terminal displays the detailed travel plan received from the server to the user.
[0633] Input: Travel plan sent from the server
[0634] Specific operation: Visually display the travel route and schedule on the device screen (web page or application).
[0635] Step 10:
[0636] The user reviews the proposed plan and is instructed to make changes as needed.
[0637] Input: Instructions for the change (Example: I want to stop by the Osaka product exhibition on my way home)
[0638] Specific actions: The user reviews the presented plan, enters any necessary changes in text, and clicks the submit button.
[0639] Step 11:
[0640] The terminal sends a change instruction to the server.
[0641] Input: Instructions for change
[0642] Specific action: The terminal sends the user's change instruction to the server.
[0643] Step 12:
[0644] The server updates the plan, taking the changes into account, and sends it back to the user's device.
[0645] Input: Changed instructions
[0646] Specific operation: The server recalculates the travel plan based on the new instructions and sends the updated plan to the user's terminal.
[0647] Step 13:
[0648] The server analyzes user preferences based on a database and incorporates them into suggestions for the next travel plan.
[0649] Input: User's past usage history and current travel plan
[0650] Specific operation: The server analyzes historical data in the database, extracts user preferences and trends, and performs calculations to reflect these in future travel suggestions.
[0651] Step 14:
[0652] When a user receives new suggestions, they will explore further preferred locations and routes.
[0653] Input: New travel suggestions from the server
[0654] Specific operation: The user receives suggestions and considers new travel destinations and routes that suit their preferences.
[0655] (Application Example 1)
[0656] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0657] Traditional food delivery services struggle to identify specific dishes users want, lacking ways to improve the user experience. They also lack efficient support for users searching for specific cuisines or restaurants while traveling or out and about. This results in users having to spend a lot of time choosing meals and deciding on delivery plans, leading to reduced convenience.
[0658] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0659] In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated route and detailed schedule to the user, a suggestion means for making additional suggestions based on the user's preferences, a means for identifying a specific dish and serving location from a food image, and a means for generating an optimal route and delivery plan from the serving location. This makes it possible for the user to identify the restaurant they want to go to or the dish they want to eat using an image, and for the server to efficiently suggest the location and the optimal delivery plan.
[0660] "Image input means" refers to hardware or software functions that allow a user to upload images to a system.
[0661] "Analysis means" refers to technologies for identifying geographical location information and specific dishes from input images.
[0662] A "plan generation means" is a system function for generating routes and delivery plans based on analyzed information and user instructions.
[0663] "Presentation means" refers to methods or devices for displaying the generated plan or detailed schedule to the user.
[0664] The "suggestion method" is a function that makes additional suggestions based on the user's past usage history and preferences.
[0665] "Specific dish" refers to the type or name of food that can be identified from images uploaded by users.
[0666] "Place of service" refers to the location of the restaurant or delivery service where the specific dish is served.
[0667] A "delivery plan" is a plan that includes the optimal delivery route and time from the user's current location to the delivery destination.
[0668] This invention relates to a system that identifies specific locations or dishes based on images input by a user, and efficiently generates and presents plans and delivery plans based on those locations. This system consists of an image input means, an analysis means, a plan generation means, a presentation means, and a suggestion means.
[0669] System Overview
[0670] 1. Image input and analysis
[0671] Users access the system using a smartphone or PC and upload images of restaurants they want to visit or dishes they want to eat. The device sends this image data to the server. The server uses image recognition technology, such as Google's "Gemini," to analyze the specific locations and dishes in the input images and stores that information in a database.
[0672] 2. Plan generation and user specification
[0673] The user instructs the server via text input, "I want the food in the uploaded photo delivered." They also specify the length of stay at each location and the desired delivery time. The device then sends this information to the server.
[0674] 3. Generating the optimal route and detailed schedule
[0675] The server generates the optimal route from the user's current location or starting point to a restaurant serving a specific dish. Considering travel time and transportation methods, the server calculates a detailed delivery plan. The generated plan is immediately presented to the user's device.
[0676] 4. Customize your plan
[0677] Users can review the presented delivery plan and make changes as needed. For example, they can select a different restaurant or add information to their route home if they want to stop at a specific location. The device sends this information to the server, which then updates the plan.
[0678] 5. Suggestions based on user preferences
[0679] The server analyzes the user's past usage history and current plan to understand their preferences and incorporates these into future recommendations. This allows for customized recommendations tailored to the user's tastes.
[0680] Specific processing instructions
[0681] 1. Image input and analysis:
[0682] Users access the system from their smartphones or PCs and upload images of food or restaurants. The device sends the uploaded image data to the server, which analyzes the images using tools such as Google's "Gemini" and stores the information in a database.
[0683] 2. Plan generation and user specification:
[0684] The user instructs the server via text input, "I want the food in the picture delivered," and this information is sent from the terminal to the server. The server then generates the optimal plan based on the restaurants and delivery services that offer the specific dish.
[0685] 3. Generating the optimal route and detailed schedule:
[0686] The server calculates the optimal route from the user's current location to the restaurant and generates a detailed schedule. The generated plan is then displayed on the terminal.
[0687] 4. Customize your plan:
[0688] The user reviews the presented plan and makes changes as needed. The device sends the change information to the server, which then generates and presents the updated plan again.
[0689] 5. Additional proposals and historical analysis:
[0690] The server analyzes the user's past usage history and preferences based on a database to optimize future recommendations.
[0691] Specific example
[0692] For example, if a user uploads a photo of fried rice, the server analyzes the photo and identifies it as fried rice. The user then enters "I want fried rice delivered," and the server generates the optimal delivery plan from a specific restaurant. The user can review this plan and request changes if necessary. The server also generates a plan that reflects any stops along the return journey.
[0693] Example of a prompt
[0694] "Upload a photo of fried rice and suggest restaurants that offer delivery of that dish."
[0695] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0696] Step 1:
[0697] Uploading and sending images
[0698] Users log in to the system using their smartphones or PCs and upload images of the food they want to eat or the restaurants they want to visit. The device receives the image data sent by the user and sends it to the server. The input here is the image taken by the user, and the output is the image data sent to the server.
[0699] Step 2:
[0700] Image analysis
[0701] The server uses image recognition technologies such as Google's "Gemini" to analyze the received image data. In this process, the image data serves as input, and features and content within the image are recognized during the analysis process, identifying specific dishes or locations. The analysis results are stored in a database as geographical location information and information about the specific dish.
[0702] Step 3:
[0703] Text input and sending instructions
[0704] The user enters a text instruction via the system interface, such as "I want the food in the uploaded image delivered." They can also enter additional information, such as their desired delivery time. The terminal then sends these instructions to the server. The input is the user's text instruction, and the output is the data sent to the server.
[0705] Step 4:
[0706] Generating a delivery plan
[0707] The server generates the optimal delivery route and schedule based on the user's current location and information on the nearest restaurants that serve the specified dish. Analyzed information and user instructions serve as input, and the optimal route and detailed schedule are output. This calculation also utilizes data such as transportation methods and travel time.
[0708] Step 5:
[0709] Presentation of the plan
[0710] The generated delivery plan and detailed schedule are sent from the server to the terminal, which then displays them to the user. The input here is the plan data generated on the server side, and the output is the plan information displayed on the user's terminal.
[0711] Step 6:
[0712] Customize your plan
[0713] The user reviews the presented delivery plan and makes changes as needed, such as selecting a different restaurant or specifying additional transit points. The terminal sends the user's change instructions to the server, which then regenerates the plan based on this information. The input is the user's change instructions, and the output is the updated plan.
[0714] Step 7:
[0715] Presentation and approval of the final plan
[0716] The server generates the updated plan and sends it back to the user's device. The user approves the final plan, and the plan is finalized. The input here is the new plan data, and the output is the final plan display and approval on the user's device.
[0717] Step 8:
[0718] Additional proposals and historical analysis
[0719] The server makes delivery suggestions for future deliveries based on the user's past usage history and preferences. In this process, past usage data is used as input, and suggestion data based on the user's preferences is output. The suggested content will be presented to the user the next time they use the service.
[0720] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0721] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. Furthermore, this invention incorporates an emotion engine that recognizes the user's emotions and includes a function to adjust the plan according to the user's emotional state. The system of this invention includes image input means, analysis means, plan generation means, presentation means, suggestion means, and an emotion engine.
[0722] System Overview
[0723] 1. Image Input and Analysis
[0724] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[0725] 2. Plan generation and user specification
[0726] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[0727] 3. Generating the optimal route and detailed schedule
[0728] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[0729] 4. User state recognition using an emotion engine
[0730] This invention incorporates an emotion engine that analyzes the user's facial expressions and voice, and incorporates this into the travel plan. The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits this data to a server. The server analyzes the collected data and recognizes the user's emotional state.
[0731] 5. Adjusting plans based on emotions
[0732] Based on the perceived emotions, the server adjusts travel routes and itineraries in real time. For example, if a user appears tired, the plan can be changed to include more rest periods.
[0733] 6. Plan presentation and customization
[0734] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[0735] 7. Use of long-term emotional data
[0736] The server records the user's emotional state over the long term and uses this history to optimize future travel suggestions. This historical data is used to further refine the user's preferred types of travel plans.
[0737] Explain the program's processing in natural language.
[0738] 1. Image input and analysis
[0739] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[0740] 2. The device sends the uploaded image data to the server.
[0741] 3. The server uses image recognition technology to analyze specific locations within the image.
[0742] 4. The server identifies the geographical location based on the analysis results and stores this information in the database.
[0743] 2. Creating a travel plan
[0744] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[0745] 2. The terminal sends instructions from the user to the server.
[0746] 3. The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route.
[0747] 4. The server calculates the mode of transport and travel time, and determines the detailed schedule.
[0748] 5. The server sends the generated travel plan to the user's device.
[0749] 3. User state recognition using an emotion engine
[0750] 1. The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[0751] 2. The device sends the collected data to the server.
[0752] 3. The server uses an emotion engine to analyze the data and recognize the user's emotional state.
[0753] 4. Adjusting plans based on emotions
[0754] 1. The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine.
[0755] 2. For example, if a user is feeling stressed, the system can regenerate a plan to visit places or facilities where they can relax.
[0756] 3. The server sends the adjusted plan to the user's terminal.
[0757] 5. Plan presentation and customization
[0758] 1. The user reviews the travel plan presented via their device and is instructed to make changes as needed.
[0759] 2. The terminal sends the user's change instruction to the server.
[0760] 3. The server updates the plan to reflect the changed information and sends it to the user's device again.
[0761] 6. Use of long-term emotional data
[0762] 1. The server records the user's emotional state over the long term and stores it in a database.
[0763] 2. The server optimizes future travel suggestions based on accumulated historical data and analyzes user preferences in a more precise manner.
[0764] Specific example
[0765] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[0766] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0767] 3. The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[0768] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[0769] 5. The server sends the generated plan to the user, who then confirms it.
[0770] 6. The device uses its built-in camera and microphone to collect facial expressions and audio recordings of the user as they review the plan, and transmits them to the server.
[0771] 7. The server analyzes the user's emotions using an emotion engine and adjusts the plan accordingly. For example, if it determines that the user is tired, it changes the plan to include more rest time.
[0772] 8. The user enters that they want to stop by the Osaka product exhibition on their way home, and the terminal sends the information to the server.
[0773] 9. The server regenerates the modified plan and sends it to the user's terminal.
[0774] 10. The server records the user's emotional state over the long term and uses it to improve future travel suggestions.
[0775] The above describes the specific embodiments and program processing details for implementing the present invention.
[0776] The following describes the processing flow.
[0777] Step 1:
[0778] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[0779] Step 2:
[0780] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[0781] Step 3:
[0782] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[0783] Step 4:
[0784] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[0785] Step 5:
[0786] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[0787] Step 6:
[0788] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[0789] Step 7:
[0790] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[0791] Step 8:
[0792] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[0793] Step 9:
[0794] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[0795] Step 10:
[0796] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[0797] Step 11:
[0798] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[0799] Step 12:
[0800] The device uses its built-in camera and microphone to collect facial expressions and audio data as the user reviews the plan. This data is collected only with the user's permission.
[0801] Step 13:
[0802] The device sends collected facial and audio data to the server. The server uses an emotion engine to analyze the user's emotional state.
[0803] Step 14:
[0804] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. For example, if the user looks tired, it might increase the rest time.
[0805] Step 15:
[0806] The server sends the adjusted plan to the user's device. The user can review the new plan and confirm it again before starting their trip.
[0807] Step 16:
[0808] The server records the user's emotional state over the long term and stores it in a database. When planning the next trip, the server takes the user's past emotional data into consideration to provide more customized suggestions.
[0809] The above outlines the specific processing steps of the invention that combines an emotion engine.
[0810] (Example 2)
[0811] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0812] Traditional travel planning systems have limitations in identifying desired destinations based on user-provided images and generating optimal travel routes that include those locations. Furthermore, they lack the functionality to adjust plans based on the user's emotional state during the trip, making it difficult to mitigate stress and dissatisfaction. Additionally, they haven't adequately provided mechanisms for optimizing future travel plans using long-term user emotional data.
[0813] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated travel route and detailed schedule to the user, an emotion engine means for analyzing the user's emotional state and adjusting the travel plan, and a suggestion means for making additional suggestions based on the user's preferences. This makes it possible to identify desired destinations and generate optimal travel routes based on images provided by the user, and furthermore, to adjust the travel plan in real time according to the user's emotional state and optimize future travel plans using long-term emotional data.
[0814] The "image input method" is a function that allows users to upload images of places they want to visit to the system.
[0815] "Analysis means" refers to a function for determining geographical location from input image data.
[0816] The "plan generation method" is a function that generates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[0817] "Presentation means" refers to a function for displaying the generated travel route and detailed schedule to the user.
[0818] The "emotional engine" is a function that analyzes the user's facial expressions and voice to adjust the travel plan in real time.
[0819] A "suggestion mechanism" is a function that provides additional suggestions based on the user's preferences.
[0820] "Image recognition technology" is a technology that analyzes the content of an input image and extracts specific information from it.
[0821] "Means for setting the length of stay" refers to a function that allows users to specify the number of days or hours they will stay at each visited location.
[0822] This invention is a system that identifies desired destinations based on images provided by the user and generates the optimal travel route and detailed schedule that passes through those locations. Furthermore, it incorporates a function that recognizes the user's emotions and adjusts the travel plan according to their emotional state. This invention includes the following main functions:
[0823] 1. Image input and analysis
[0824] Users upload images of places they want to visit to the system from their smartphones, PCs, or other devices. The device then sends this image data to the server. The server uses "image recognition technology" to analyze and identify the geographical location of the input image. For example, if a user uploads a photo of the foot of Mount Fuji, the server analyzes the photo to identify the foot of Mount Fuji.
[0825] 2. Creating a travel plan
[0826] The user specifies their desired destination, such as "I want to go to XX via photo A and photo B," through text input, and also specifies the length of stay at each location. The terminal receives this information and sends it to the server. The server generates the optimal route based on the information about the starting point, specific locations, and destination. For example, if the user wants to go from the foot of Mt. Fuji to Kyoto via Tokyo Tower, the server will calculate a detailed schedule considering the means of transportation and travel time at each location.
[0827] 3. Presentation of the generated plan
[0828] The server sends the generated travel plan to the user's device, which then displays it. The user can review the plan and customize it as needed.
[0829] 4. User state recognition using an emotion engine
[0830] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and sends this data to a server. The server uses an "emotion engine" to analyze the data and recognize the user's emotional state. For example, if the user is smiling or looks tired while reviewing a plan, the emotion engine will analyze that data.
[0831] 5. Adjusting plans based on emotions
[0832] Based on the perceived emotions, the server may adjust the travel route and itinerary in real time. For example, if the user appears tired, the plan may be changed to include more rest periods or visits to relaxing locations. The adjusted plan is then sent back to the user's device.
[0833] 6. Accumulation and utilization of long-term emotional data
[0834] The server records the user's emotional state over the long term and stores this data in a database. This accumulated emotional data can be used to optimize future travel suggestions and propose plans that are more tailored to the user's preferences.
[0835] Specific example
[0836] For example, if a user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system, the server analyzes these photos to identify Mt. Fuji and Tokyo Tower. The user then inputs, "I want to go to Kyoto via Photo A and Photo B," and specifies the length of stay at each location. The server generates the optimal route and creates a detailed schedule considering transportation methods and travel time. The generated plan is displayed on the terminal for the user to review. At this time, the terminal uses the camera and microphone to collect the user's facial expressions and voice, which the server analyzes with an emotion engine to adjust the plan. For example, if the user inputs that they want to stop by an Osaka product exhibition on the way back, the plan will be regenerated.
[0837] Example of a prompt
[0838] "Please generate the optimal travel plan based on the following instructions. I would like to travel to Kyoto, passing through the foot of Mt. Fuji and Tokyo Tower along the way. I will stay at the foot of Mt. Fuji for one day and at Tokyo Tower for two days."
[0839] "Based on user sentiment data, please suggest a day trip plan that includes plenty of breaks."
[0840] The above describes specific embodiments for carrying out the present invention.
[0841] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0842] Program processing flow and detailed explanation
[0843] Image input and analysis
[0844] Step 1:
[0845] Users upload images of places they want to visit to the system from their smartphones or PCs. Input is image files (e.g., JPEG or PNG format). Output is the image data being saved on the device. Specifically, the user uses the system's upload function, selects an image from a browser or dedicated app, and clicks the upload button.
[0846] Step 2:
[0847] The device sends uploaded image data to the server. The input is the image file data, and the output is the image data received by the server. The device uses an internet connection to send the image data to the server as packets. Specifically, it sends the image data using an HTTP request.
[0848] Step 3:
[0849] The server uses image recognition technology to analyze the input image. The input is the received image data, and the output is the geographical location information of the analysis result. Specifically, the server uses Google's "image recognition technology" to execute an algorithm that extracts image features and determines the geographical location.
[0850] Step 4:
[0851] The server identifies the geographic location based on the analysis results and stores this information in a database. The input is geographic location information, and the output is the geographic location information stored in the database. The server performs an insert operation into the database to persist the identified location data.
[0852] Travel plan generation
[0853] Step 1:
[0854] The user enters text such as "I want to go to XX via photo A and photo B" and specifies the length of stay at each location. The input is text data (intermediate points and length of stay). The output is the instruction information entered into the terminal. Specifically, the user enters the desired intermediate points and length of stay in the text input field and clicks the submit button.
[0855] Step 2:
[0856] The terminal sends user instructions to the server. The input is text data (intermediate locations and length of stay), and the output is the instructions received by the server. The terminal sends the instructions to the server using an HTTP request.
[0857] Step 3:
[0858] The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route. The input is the user's instruction information (waypoints, starting point, destination, length of stay), and the output is the generated optimal route information. Specifically, the server uses a Geographic Information System (GIS) to calculate the optimal route connecting each point using an algorithm.
[0859] Step 4:
[0860] The server calculates the mode of transport and travel time to determine a detailed schedule. The input is the generated optimal route information, and the output is a detailed schedule (mode of transport, departure time, arrival time, etc.). Specifically, the server uses a traffic information API to calculate the travel time for each segment of the journey.
[0861] Step 5:
[0862] The server generates a travel plan and sends it to the user's terminal. The input is a detailed schedule, and the output is the travel plan displayed on the terminal. The server sends the detailed schedule to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[0863] User state recognition using an emotion engine
[0864] Step 1:
[0865] The device uses its built-in camera and microphone to collect user facial expressions and audio data. The input is the user's facial expressions and voice, and the output is the collected data. Specifically, the device's app activates the camera, takes a picture of the user's face, and records audio using the microphone.
[0866] Step 2:
[0867] The terminal sends the collected data to the server. The input consists of facial expressions and voice data, and the output is the data received by the server. The terminal encrypts the data and sends it to the server via an HTTP request.
[0868] Step 3:
[0869] The server uses an emotion engine to analyze data and recognize the user's emotional state. The input is facial expression and voice data, and the output is recognized emotional state information. Specifically, the server applies an emotion recognition algorithm to analyze the user's emotions (e.g., joy, sadness, fatigue, etc.).
[0870] Adjusting plans based on emotions
[0871] Step 1:
[0872] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. The input is the recognized emotional state information, and the output is the adjusted travel plan. Specifically, the server makes adjustments such as changing the schedule to include more rest.
[0873] Step 2:
[0874] The server sends the adjusted plan to the user's terminal. The input is the adjusted travel plan, and the output is the adjusted plan displayed on the terminal. The server sends the adjusted plan to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[0875] Plan presentation and customization
[0876] Step 1:
[0877] The user reviews the presented travel plan. The input is the travel plan displayed on the device, and the output is the user's review process. Specifically, the user views the page displaying the travel plan.
[0878] Step 2:
[0879] If a user wishes to change their plan, they enter that information. The input is the instruction for the desired change, and the output is the change instruction entered into the terminal. For example, they might enter, "I would like to stop by the Osaka product exhibition on my way home."
[0880] Step 3:
[0881] The terminal sends the user's change instructions to the server. The input is the information of the desired change, and the output is the change instructions received by the server. The terminal sends the change instructions to the server using an HTTP request.
[0882] Step 4:
[0883] The server regenerates the plan to reflect the changed information. The input is the change instruction information, and the output is the regenerated plan. The server recalculates the optimal route and detailed schedule and generates a new plan.
[0884] Step 5:
[0885] The server sends the updated plan to the user's terminal. The input is the regenerated plan, and the output is the updated plan displayed on the terminal. The server sends the updated plan to the terminal via an HTTP response, which the terminal receives and displays.
[0886] Use of long-term emotional data
[0887] Step 1:
[0888] The server records the user's emotional state over the long term. The input is recognized emotional state information, and the output is emotional data stored in a database. The server performs insert operations into the database to save the emotional data.
[0889] Step 2:
[0890] The server optimizes the next travel suggestion based on accumulated historical data. The input is long-term sentiment data, and the output is an optimized travel suggestion. The server analyzes the sentiment data and proposes the best travel plan based on the user's priorities and preferences.
[0891] The above outlines the specific processing steps of the system.
[0892] (Application Example 2)
[0893] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0894] Traditional travel planning systems have been somewhat effective in creating travel plans that meet users' preferences, but they have a weakness in their inability to adjust plans to reflect real-time emotional states during the trip. Furthermore, means for users to smartly review their plans during their trip, and visually supportive devices, are still not fully utilized. Therefore, there is a need to reduce stress during travel and improve the user experience.
[0895] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0896] In this invention, the server includes means for inputting an image, analysis means for identifying a geographical location from the input image, plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, presentation means for presenting the generated travel route and detailed schedule to the user, suggestion means for making additional suggestions based on the user's preferences, means including an emotion engine that analyzes the user's facial expressions and voice data and adjusts the travel plan based on their emotional state, and means for presenting the travel plan to the user using a real-time adaptable mobile or visual device. This makes it possible to adjust the plan in real time to reflect the user's emotional state during the trip and to improve the user's travel experience through visual support from a smart device.
[0897] "Image input means" refers to a device or software that receives image data input by a user and transmits it to a server.
[0898] "Analysis means" refers to a device or software that identifies a specific geographical location from input image data and provides that information to a server.
[0899] A "plan generation means" refers to a device or software that creates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[0900] "Presentation means" refers to a device or software that provides the user with the generated travel route and detailed schedule visually or audibly.
[0901] "Suggestion means" refers to a device or software that provides additional information or recommended plans regarding travel routes based on the user's preferences.
[0902] "Means including an emotion engine" refers to a device or software that analyzes the user's facial expressions and voice data to recognize their emotional state and adjust the travel plan accordingly.
[0903] "Means for presenting a travel plan to a user using a mobile or visual device" refers to a device or software that visually presents a travel plan while reflecting the user's location information and emotional state in real time.
[0904] "Real-time" refers to a state where the user's current status and location information are reflected in real time, allowing for immediate adaptation or modification.
[0905] This invention is a system for efficiently creating travel plans and adjusting them in real time according to the user's emotional state. The system includes means for image input, analysis, plan generation, presentation, suggestion, an emotion engine, and means for presenting the travel plan to the user using a mobile device or visual device.
[0906] Description of the program's overall operation
[0907] 1. Image input and analysis
[0908] Users upload images of places they want to visit from their smartphones or PCs to the system in order to create travel plans. In this case, the image input mechanism is activated. The device sends these images to the server, which uses image recognition technology (e.g., a common image recognition API) to determine the geographical location of the images. The results of this analysis are stored in a database.
[0909] 2. Plan generation and user specification
[0910] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location. Based on this, the terminal sends the instructions to the server. The server generates the optimal travel route and detailed schedule based on the analyzed geographic location information and the user's instructions. This generation process uses a route calculation engine and transportation information.
[0911] 3. Collection and analysis of facial expressions and voice.
[0912] The device uses its built-in camera and microphone to collect user facial expressions and voice data. This data is sent to a server in real time, where the server uses an emotion engine to analyze the user's emotional state. Specifically, facial expressions are analyzed using facial recognition software (e.g., OpenCV or dlib), and the data is sent to an emotion analysis API.
[0913] 4. Adjusting plans based on emotions
[0914] Based on the analysis results from the emotion engine, the server adjusts the travel plan to match the user's emotional state. For example, if the user is tired, the plan will be changed to include relaxing facilities. Conversely, if the user is excited, the plan can be adjusted to include more active spots.
[0915] 5. Real-time plan presentation
[0916] The travel plan, generated in real time, is presented to the user through visual devices such as smart glasses. This allows the user to enjoy their trip while checking a plan that suits their emotional state in real time.
[0917] Specific example
[0918] The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a photo of an urban landmark (Photo B) to the system.
[0919] The server identifies photo A as the foot of Mount Fuji and photo B as an urban area, and saves them in the database.
[0920] The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[0921] The terminal sends instruction information to the server, which then generates the optimal route and a detailed schedule.
[0922] The device uses its built-in camera and microphone to collect the user's facial expressions and voice in real time as they review the plan, and transmits this information to the server.
[0923] The server uses an emotion engine to analyze the user's emotions and adjust the plan accordingly. For example, if it determines that the user is tired, it will change the plan to include more rest time.
[0924] Example of a prompt
[0925] "Since you seem tired this time, please suggest a travel plan that includes places where you can relax."
[0926] This system allows for real-time adjustments to travel plans that reflect the user's emotional state during their trip, and enhances the user's travel experience through visual support via smart devices.
[0927] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0928] Step 1:
[0929] The user uses their device to upload images of the place they want to go to the system.
[0930] The specific input is the user's image data, which the device sends to the server. The server uses image recognition technology to analyze the uploaded image. For example, it uses an image recognition API to determine the geographical location and stores this information in a database. The output is the analyzed geographical location information.
[0931] Step 2:
[0932] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location.
[0933] The input is user text instructions, which the terminal sends to the server. The server generates the optimal travel route and detailed schedule based on the starting point, analyzed geographic location information, and user instructions. It uses a route calculation engine to determine the best route, taking into account transportation methods and travel time. The output is the generated travel plan and detailed schedule.
[0934] Step 3:
[0935] The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[0936] The input consists of the user's real-time facial expressions and voice data, which the terminal collects and sends to the server. The server uses an emotion engine to analyze the data and recognize the user's emotional state. Specifically, it analyzes facial expressions using facial recognition software (e.g., OpenCV or dlib) and sends the data to an emotion analysis API to identify the emotional state. The output is the user's emotional state as determined by the emotion engine.
[0937] Step 4:
[0938] The server adjusts the travel plan in real time according to the user's emotional state, based on the analysis results from the emotion engine.
[0939] The input consists of the user's emotional state, as determined by the emotion engine, and the existing travel plan. Based on this data, the server regenerates the plan, including new suggestions and modifications. For example, if the user is tired, the server might add rest areas and modify the plan to include relaxing facilities. The output is the adjusted travel plan.
[0940] Step 5:
[0941] The device presents the user with travel plans that have been generated or adjusted in real time.
[0942] The input is a pre-configured travel plan, and the device visually presents this information to the user via smart glasses or a smartphone. The user can continue their trip while checking the plan in real time to match their emotional state. The output is the pre-configured travel plan displayed on the user's device.
[0943] Step 6:
[0944] The server records the user's emotional state over the long term to optimize future travel suggestions.
[0945] The input consists of past travel plans and user emotional state data, which the server stores in a database. This data is used to generate optimal travel plans based on the user's preferences for future trips. The output is a more personalized suggestion for the next trip.
[0946] Through these steps, users can enjoy travel plans tailored to their emotional state in real time, enhancing their travel experience.
[0947] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0948] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0949] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0950] [Third Embodiment]
[0951] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0952] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0953] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0954] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0955] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0956] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0957] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0958] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0959] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0960] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0961] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0962] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0963] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. The system of this invention includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[0964] System Overview
[0965] 1. Image Input and Analysis
[0966] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[0967] 2. Plan generation and user specification
[0968] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[0969] 3. Generating the optimal route and detailed schedule
[0970] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[0971] 4. Customize your plan
[0972] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[0973] 5. Suggestions based on user preferences
[0974] The server analyzes the user's preferences based on their past usage history and current travel plan, and incorporates this into future recommendations. This enables customized suggestions tailored to the user's preferences.
[0975] Explain the program's processing in natural language.
[0976] 1. Image input and analysis
[0977] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[0978] 2. The device sends the uploaded image data to the server.
[0979] 3. The server uses image recognition technology to analyze specific locations within the image.
[0980] 4. The server saves the analysis results to the database.
[0981] 2. Creating a travel plan
[0982] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[0983] 2. The terminal sends instructions from the user to the server.
[0984] 3. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and destination XX, and generates a detailed schedule.
[0985] 4. The server sends the generated plan to the user's terminal.
[0986] 3. Plan presentation and customization
[0987] 1. The terminal displays the detailed travel plan received from the server to the user.
[0988] 2. The user reviews the proposed plan and instructs the user to make changes as needed (e.g., add a route through a local products fair).
[0989] 3. The terminal sends a change instruction to the server.
[0990] 4. The server updates the plan, taking the changes into account, and sends it back to the user's device.
[0991] 4. Additional proposals and historical analysis
[0992] 1. The server analyzes the user's preferences based on the database and incorporates them into suggestions for the next travel plan.
[0993] 2. When users receive new suggestions, it becomes easier for them to explore further preferred locations and routes.
[0994] Specific example
[0995] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[0996] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[0997] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[0998] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[0999] 5. The server sends the generated plan to the user, who then confirms it.
[1000] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[1001] 7. The server is updated, the plan is regenerated, and sent to the user's terminal.
[1002] The above describes the specific embodiments and program processing details for implementing the present invention.
[1003] The following describes the processing flow.
[1004] Step 1:
[1005] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[1006] Step 2:
[1007] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[1008] Step 3:
[1009] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[1010] Step 4:
[1011] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[1012] Step 5:
[1013] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[1014] Step 6:
[1015] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[1016] Step 7:
[1017] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[1018] Step 8:
[1019] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[1020] Step 9:
[1021] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[1022] Step 10:
[1023] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[1024] Step 11:
[1025] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[1026] Step 12:
[1027] The server stores the user's past usage history and current plan data in a database and analyzes the user's preferences. This allows for the creation of customized travel suggestions that are tailored to the user's tastes.
[1028] (Example 1)
[1029] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1030] Traditional travel planning systems had the problem of requiring users to manually input detailed information about the places they wanted to visit, which was time-consuming. Furthermore, the generation of travel routes and schedules relied heavily on user input, lacking automation. As a result, users were unable to create travel plans efficiently, resulting in a time-consuming and cumbersome process. Additionally, the lack of customized suggestions based on user preferences made it difficult to improve user satisfaction.
[1031] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1032] In this invention, the server includes means for inputting an image, means for analyzing the input image to identify a geographical location, means for generating a plan that generates a travel route based on the analyzed geographical location information and user instructions, means for presenting the generated travel route and detailed schedule to the user, and means for making additional suggestions based on the user's preferences. This makes it possible for users to easily upload an image and have a travel route and detailed schedule automatically generated, enabling the provision of efficient and customized travel plans.
[1033] "Means of inputting images" refers to the function that allows users to upload image data to the system via their device.
[1034] "Means for analyzing input images to determine geographical location" refers to technologies that analyze uploaded image data to identify geographical information (e.g., longitude and latitude) of the location shown in the image.
[1035] The "means for generating plans" refer to a function that automatically creates the optimal travel route and detailed schedule for visiting multiple locations, based on analyzed geographical location information and user instructions.
[1036] "Means of presentation" refers to functions that visually display the generated travel route and detailed schedule to the user.
[1037] The "suggestion method" refers to a function that provides additional suggestions tailored to the user's preferences, based on the user's past travel history and current travel plan.
[1038] "Image recognition technology" is a technique that uses machine learning and algorithms to identify specific objects or locations within an image.
[1039] "A means for users to set the duration of their stay at each location" refers to a function that allows users to set the number of days or hours they will stay at each designated location.
[1040] This invention relates to a system that identifies a desired destination based on images provided by a user and generates a detailed travel route and schedule. The system includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[1041] First, the user uploads an image of the place they want to go to their device (such as a smartphone or PC). The device then sends this image data to the server. The server uses image recognition technology (such as "Gemini") to analyze the geographical location of the uploaded image and stores the analysis results in a database.
[1042] Next, the user uses text input to specify their desired destination, indicating that they want to travel via designated locations, and to specify the duration of their stay at each point. Upon receiving these instructions, the terminal sends the user's information to the server. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and then to the final destination XX, generating a detailed schedule that takes into account transportation methods and travel time. The generated plan is immediately presented to the user's terminal via the internet.
[1043] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can enter that information. The terminal sends this information to the server, which then updates the plan.
[1044] Furthermore, the server can analyze the user's preferences based on their past usage history and current travel plans, and reflect this in future travel suggestions. This enables customized suggestions tailored to the user's preferences.
[1045] Specific example
[1046] 1. The user uploads a landscape photo of the foot of Mt. Fuji and a night view photo of Tokyo Tower to the system.
[1047] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[1048] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[1049] 4. The terminal sends instruction information to the server, which generates the optimal route and detailed schedule.
[1050] 5. The server sends the generated plan to the user's terminal, and the user confirms it.
[1051] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[1052] 7. The server regenerates the updated plan and sends it to the user's terminal.
[1053] Examples of prompts to input into a generative AI model
[1054] A user has uploaded landscape photos of the foot of Mt. Fuji and nighttime photos of Tokyo Tower to the system. Please use these photos to create a travel plan to Kyoto.
[1055] I stayed for one day at the foot of Mt. Fuji.
[1056] I stayed at Tokyo Tower for two days.
[1057] The final destination of the trip is Kyoto
[1058] I'd like to stop by the Osaka product fair on my way home.
[1059] Generate the optimal route and detailed schedule, and present them to the user.
[1060] Thus, this invention enables users to easily upload images, and the server automatically generates travel routes and detailed schedules. Furthermore, by providing customized suggestions based on user preferences, it realizes a system that offers more satisfying travel plans.
[1061] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1062] Step 1:
[1063] The user accesses the system and uploads an image of the place they want to go to their device.
[1064] Input: Image of the place you want to go (e.g., a landscape photo of Mt. Fuji)
[1065] Specific operation: The user accesses the system's web page or application and clicks the image upload button. The upload screen appears, and the user selects an image file from their device and sends it to the server.
[1066] Step 2:
[1067] The device sends the uploaded image data to the server.
[1068] Input: Image data uploaded by the user
[1069] Specific operation: The terminal divides the image data into packets and sends them to the server over the network.
[1070] Step 3:
[1071] The server uses image recognition technology to analyze the geographical location of the input image.
[1072] Input: Image data sent to the server
[1073] Specific operation: The server uses image recognition technology (e.g., "Gemini") to analyze images and identify specific landmarks or geographical features within them. The identified geographical locations (e.g., longitude and latitude) are recorded in a database.
[1074] Step 4:
[1075] The server saves the analysis results to the database.
[1076] Input: Geographic location information obtained using image recognition technology
[1077] Specific operation: The server connects to the database and saves the identified geographic coordinates and related information to the "Location" table.
[1078] Step 5:
[1079] The user provides text input indicating that they want to travel to their destination via specific locations and specifies the length of stay at each point.
[1080] Input: User instructions in text format (Example: "I want to spend one day at Mt. Fuji, two days at Tokyo Tower, and then go to my final destination, Kyoto.")
[1081] Specific operation: The user enters the desired location and length of stay into the system's text input field and clicks the submit button.
[1082] Step 6:
[1083] The terminal sends instructions from the user to the server.
[1084] Input: User instructions
[1085] Specific operation: The terminal sends user instructions in text format to the server.
[1086] Step 7:
[1087] The server searches for the optimal route from the starting point to each identified point and the final destination, and generates a detailed schedule.
[1088] Input: User-specified geographical location and length of stay
[1089] Specific operation: The server uses map databases and traffic information to calculate the optimal route from the starting point (current location) to each specified point and the final destination. It then creates a planned schedule, taking into account the length of stay.
[1090] Step 8:
[1091] The server sends the generated plan to the user's device.
[1092] Input: Detailed travel plan (route and schedule)
[1093] Specific operation: Format the generated plan and send it to the user's terminal via the internet.
[1094] Step 9:
[1095] The terminal displays the detailed travel plan received from the server to the user.
[1096] Input: Travel plan sent from the server
[1097] Specific operation: Visually display the travel route and schedule on the device screen (web page or application).
[1098] Step 10:
[1099] The user reviews the proposed plan and is instructed to make changes as needed.
[1100] Input: Instructions for the change (Example: I want to stop by the Osaka product exhibition on my way home)
[1101] Specific actions: The user reviews the presented plan, enters any necessary changes in text, and clicks the submit button.
[1102] Step 11:
[1103] The terminal sends a change instruction to the server.
[1104] Input: Instructions for change
[1105] Specific action: The terminal sends the user's change instruction to the server.
[1106] Step 12:
[1107] The server updates the plan, taking the changes into account, and sends it back to the user's device.
[1108] Input: Changed instructions
[1109] Specific operation: The server recalculates the travel plan based on the new instructions and sends the updated plan to the user's terminal.
[1110] Step 13:
[1111] The server analyzes user preferences based on a database and incorporates them into suggestions for the next travel plan.
[1112] Input: User's past usage history and current travel plan
[1113] Specific operation: The server analyzes historical data in the database, extracts user preferences and trends, and performs calculations to reflect these in future travel suggestions.
[1114] Step 14:
[1115] When a user receives new suggestions, they will explore further preferred locations and routes.
[1116] Input: New travel suggestions from the server
[1117] Specific operation: The user receives suggestions and considers new travel destinations and routes that suit their preferences.
[1118] (Application Example 1)
[1119] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1120] Traditional food delivery services struggle to identify specific dishes users want, lacking ways to improve the user experience. They also lack efficient support for users searching for specific cuisines or restaurants while traveling or out and about. This results in users having to spend a lot of time choosing meals and deciding on delivery plans, leading to reduced convenience.
[1121] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1122] In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated route and detailed schedule to the user, a suggestion means for making additional suggestions based on the user's preferences, a means for identifying a specific dish and serving location from a food image, and a means for generating an optimal route and delivery plan from the serving location. This makes it possible for the user to identify the restaurant they want to go to or the dish they want to eat using an image, and for the server to efficiently suggest the location and the optimal delivery plan.
[1123] "Image input means" refers to hardware or software functions that allow a user to upload images to a system.
[1124] "Analysis means" refers to technologies for identifying geographical location information and specific dishes from input images.
[1125] A "plan generation means" is a system function for generating routes and delivery plans based on analyzed information and user instructions.
[1126] "Presentation means" refers to methods or devices for displaying the generated plan or detailed schedule to the user.
[1127] The "suggestion method" is a function that makes additional suggestions based on the user's past usage history and preferences.
[1128] "Specific dish" refers to the type or name of food that can be identified from images uploaded by users.
[1129] "Place of service" refers to the location of the restaurant or delivery service where the specific dish is served.
[1130] A "delivery plan" is a plan that includes the optimal delivery route and time from the user's current location to the delivery destination.
[1131] This invention relates to a system that identifies specific locations or dishes based on images input by a user, and efficiently generates and presents plans and delivery plans based on those locations. This system consists of an image input means, an analysis means, a plan generation means, a presentation means, and a suggestion means.
[1132] System Overview
[1133] 1. Image input and analysis
[1134] Users access the system using a smartphone or PC and upload images of restaurants they want to visit or dishes they want to eat. The device sends this image data to the server. The server uses image recognition technology, such as Google's "Gemini," to analyze the specific locations and dishes in the input images and stores that information in a database.
[1135] 2. Plan generation and user specification
[1136] The user instructs the server via text input, "I want the food in the uploaded photo delivered." They also specify the length of stay at each location and the desired delivery time. The device then sends this information to the server.
[1137] 3. Generating the optimal route and detailed schedule
[1138] The server generates the optimal route from the user's current location or starting point to a restaurant serving a specific dish. Considering travel time and transportation methods, the server calculates a detailed delivery plan. The generated plan is immediately presented to the user's device.
[1139] 4. Customize your plan
[1140] Users can review the presented delivery plan and make changes as needed. For example, they can select a different restaurant or add information to their route home if they want to stop at a specific location. The device sends this information to the server, which then updates the plan.
[1141] 5. Suggestions based on user preferences
[1142] The server analyzes the user's past usage history and current plan to understand their preferences and incorporates these into future recommendations. This allows for customized recommendations tailored to the user's tastes.
[1143] Specific processing instructions
[1144] 1. Image input and analysis:
[1145] Users access the system from their smartphones or PCs and upload images of food or restaurants. The device sends the uploaded image data to the server, which analyzes the images using tools such as Google's "Gemini" and stores the information in a database.
[1146] 2. Plan generation and user specification:
[1147] The user instructs the server via text input, "I want the food in the picture delivered," and this information is sent from the terminal to the server. The server then generates the optimal plan based on the restaurants and delivery services that offer the specific dish.
[1148] 3. Generating the optimal route and detailed schedule:
[1149] The server calculates the optimal route from the user's current location to the restaurant and generates a detailed schedule. The generated plan is then displayed on the terminal.
[1150] 4. Customize your plan:
[1151] The user reviews the presented plan and makes changes as needed. The device sends the change information to the server, which then generates and presents the updated plan again.
[1152] 5. Additional proposals and historical analysis:
[1153] The server analyzes the user's past usage history and preferences based on a database to optimize future recommendations.
[1154] Specific example
[1155] For example, if a user uploads a photo of fried rice, the server analyzes the photo and identifies it as fried rice. The user then enters "I want fried rice delivered," and the server generates the optimal delivery plan from a specific restaurant. The user can review this plan and request changes if necessary. The server also generates a plan that reflects any stops along the return journey.
[1156] Example of a prompt
[1157] "Upload a photo of fried rice and suggest restaurants that offer delivery of that dish."
[1158] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1159] Step 1:
[1160] Uploading and sending images
[1161] Users log in to the system using their smartphones or PCs and upload images of the food they want to eat or the restaurants they want to visit. The device receives the image data sent by the user and sends it to the server. The input here is the image taken by the user, and the output is the image data sent to the server.
[1162] Step 2:
[1163] Image analysis
[1164] The server uses image recognition technologies such as Google's "Gemini" to analyze the received image data. In this process, the image data serves as input, and features and content within the image are recognized during the analysis process, identifying specific dishes or locations. The analysis results are stored in a database as geographical location information and information about the specific dish.
[1165] Step 3:
[1166] Text input and sending instructions
[1167] The user enters a text instruction via the system interface, such as "I want the food in the uploaded image delivered." They can also enter additional information, such as their desired delivery time. The terminal then sends these instructions to the server. The input is the user's text instruction, and the output is the data sent to the server.
[1168] Step 4:
[1169] Generating a delivery plan
[1170] The server generates the optimal delivery route and schedule based on the user's current location and information on the nearest restaurants that serve the specified dish. Analyzed information and user instructions serve as input, and the optimal route and detailed schedule are output. This calculation also utilizes data such as transportation methods and travel time.
[1171] Step 5:
[1172] Presentation of the plan
[1173] The generated delivery plan and detailed schedule are sent from the server to the terminal, which then displays them to the user. The input here is the plan data generated on the server side, and the output is the plan information displayed on the user's terminal.
[1174] Step 6:
[1175] Customize your plan
[1176] The user reviews the presented delivery plan and makes changes as needed, such as selecting a different restaurant or specifying additional transit points. The terminal sends the user's change instructions to the server, which then regenerates the plan based on this information. The input is the user's change instructions, and the output is the updated plan.
[1177] Step 7:
[1178] Presentation and approval of the final plan
[1179] The server generates the updated plan and sends it back to the user's device. The user approves the final plan, and the plan is finalized. The input here is the new plan data, and the output is the final plan display and approval on the user's device.
[1180] Step 8:
[1181] Additional proposals and historical analysis
[1182] The server makes delivery suggestions for future deliveries based on the user's past usage history and preferences. In this process, past usage data is used as input, and suggestion data based on the user's preferences is output. The suggested content will be presented to the user the next time they use the service.
[1183] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1184] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. Furthermore, this invention incorporates an emotion engine that recognizes the user's emotions and includes a function to adjust the plan according to the user's emotional state. The system of this invention includes image input means, analysis means, plan generation means, presentation means, suggestion means, and an emotion engine.
[1185] System Overview
[1186] 1. Image Input and Analysis
[1187] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[1188] 2. Plan generation and user specification
[1189] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[1190] 3. Generating the optimal route and detailed schedule
[1191] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[1192] 4. User state recognition using an emotion engine
[1193] This invention incorporates an emotion engine that analyzes the user's facial expressions and voice, and incorporates this into the travel plan. The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits this data to a server. The server analyzes the collected data and recognizes the user's emotional state.
[1194] 5. Adjusting plans based on emotions
[1195] Based on the perceived emotions, the server adjusts travel routes and itineraries in real time. For example, if a user appears tired, the plan can be changed to include more rest periods.
[1196] 6. Plan presentation and customization
[1197] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[1198] 7. Use of long-term emotional data
[1199] The server records the user's emotional state over the long term and uses this history to optimize future travel suggestions. This historical data is used to further refine the user's preferred types of travel plans.
[1200] Explain the program's processing in natural language.
[1201] 1. Image input and analysis
[1202] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[1203] 2. The device sends the uploaded image data to the server.
[1204] 3. The server uses image recognition technology to analyze specific locations within the image.
[1205] 4. The server identifies the geographical location based on the analysis results and stores this information in the database.
[1206] 2. Creating a travel plan
[1207] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[1208] 2. The terminal sends instructions from the user to the server.
[1209] 3. The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route.
[1210] 4. The server calculates the mode of transport and travel time, and determines the detailed schedule.
[1211] 5. The server sends the generated travel plan to the user's device.
[1212] 3. User state recognition using an emotion engine
[1213] 1. The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[1214] 2. The device sends the collected data to the server.
[1215] 3. The server uses an emotion engine to analyze the data and recognize the user's emotional state.
[1216] 4. Adjusting plans based on emotions
[1217] 1. The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine.
[1218] 2. For example, if a user is feeling stressed, the system can regenerate a plan to visit places or facilities where they can relax.
[1219] 3. The server sends the adjusted plan to the user's terminal.
[1220] 5. Plan presentation and customization
[1221] 1. The user reviews the travel plan presented via their device and is instructed to make changes as needed.
[1222] 2. The terminal sends the user's change instruction to the server.
[1223] 3. The server updates the plan to reflect the changed information and sends it to the user's device again.
[1224] 6. Use of long-term emotional data
[1225] 1. The server records the user's emotional state over the long term and stores it in a database.
[1226] 2. The server optimizes future travel suggestions based on accumulated historical data and analyzes user preferences in a more precise manner.
[1227] Specific example
[1228] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[1229] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[1230] 3. The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[1231] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[1232] 5. The server sends the generated plan to the user, who then confirms it.
[1233] 6. The device uses its built-in camera and microphone to collect facial expressions and audio recordings of the user as they review the plan, and transmits them to the server.
[1234] 7. The server analyzes the user's emotions using an emotion engine and adjusts the plan accordingly. For example, if it determines that the user is tired, it changes the plan to include more rest time.
[1235] 8. The user enters that they want to stop by the Osaka product exhibition on their way home, and the terminal sends the information to the server.
[1236] 9. The server regenerates the modified plan and sends it to the user's terminal.
[1237] 10. The server records the user's emotional state over the long term and uses it to improve future travel suggestions.
[1238] The above describes the specific embodiments and program processing details for implementing the present invention.
[1239] The following describes the processing flow.
[1240] Step 1:
[1241] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[1242] Step 2:
[1243] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[1244] Step 3:
[1245] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[1246] Step 4:
[1247] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[1248] Step 5:
[1249] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[1250] Step 6:
[1251] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[1252] Step 7:
[1253] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[1254] Step 8:
[1255] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[1256] Step 9:
[1257] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[1258] Step 10:
[1259] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[1260] Step 11:
[1261] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[1262] Step 12:
[1263] The device uses its built-in camera and microphone to collect facial expressions and audio data as the user reviews the plan. This data is collected only with the user's permission.
[1264] Step 13:
[1265] The device sends collected facial and audio data to the server. The server uses an emotion engine to analyze the user's emotional state.
[1266] Step 14:
[1267] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. For example, if the user looks tired, it might increase the rest time.
[1268] Step 15:
[1269] The server sends the adjusted plan to the user's device. The user can review the new plan and confirm it again before starting their trip.
[1270] Step 16:
[1271] The server records the user's emotional state over the long term and stores it in a database. When planning the next trip, the server takes the user's past emotional data into consideration to provide more customized suggestions.
[1272] The above outlines the specific processing steps of the invention that combines an emotion engine.
[1273] (Example 2)
[1274] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1275] Traditional travel planning systems have limitations in identifying desired destinations based on user-provided images and generating optimal travel routes that include those locations. Furthermore, they lack the functionality to adjust plans based on the user's emotional state during the trip, making it difficult to mitigate stress and dissatisfaction. Additionally, they haven't adequately provided mechanisms for optimizing future travel plans using long-term user emotional data.
[1276] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated travel route and detailed schedule to the user, an emotion engine means for analyzing the user's emotional state and adjusting the travel plan, and a suggestion means for making additional suggestions based on the user's preferences. This makes it possible to identify desired destinations and generate optimal travel routes based on images provided by the user, and furthermore, to adjust the travel plan in real time according to the user's emotional state and optimize future travel plans using long-term emotional data.
[1277] The "image input method" is a function that allows users to upload images of places they want to visit to the system.
[1278] "Analysis means" refers to a function for determining geographical location from input image data.
[1279] The "plan generation method" is a function that generates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[1280] "Presentation means" refers to a function for displaying the generated travel route and detailed schedule to the user.
[1281] The "emotional engine" is a function that analyzes the user's facial expressions and voice to adjust the travel plan in real time.
[1282] A "suggestion mechanism" is a function that provides additional suggestions based on the user's preferences.
[1283] "Image recognition technology" is a technology that analyzes the content of an input image and extracts specific information from it.
[1284] "Means for setting the length of stay" refers to a function that allows users to specify the number of days or hours they will stay at each visited location.
[1285] This invention is a system that identifies desired destinations based on images provided by the user and generates the optimal travel route and detailed schedule that passes through those locations. Furthermore, it incorporates a function that recognizes the user's emotions and adjusts the travel plan according to their emotional state. This invention includes the following main functions:
[1286] 1. Image input and analysis
[1287] Users upload images of places they want to visit to the system from their smartphones, PCs, or other devices. The device then sends this image data to the server. The server uses "image recognition technology" to analyze and identify the geographical location of the input image. For example, if a user uploads a photo of the foot of Mount Fuji, the server analyzes the photo to identify the foot of Mount Fuji.
[1288] 2. Creating a travel plan
[1289] The user specifies their desired destination, such as "I want to go to XX via photo A and photo B," through text input, and also specifies the length of stay at each location. The terminal receives this information and sends it to the server. The server generates the optimal route based on the information about the starting point, specific locations, and destination. For example, if the user wants to go from the foot of Mt. Fuji to Kyoto via Tokyo Tower, the server will calculate a detailed schedule considering the means of transportation and travel time at each location.
[1290] 3. Presentation of the generated plan
[1291] The server sends the generated travel plan to the user's device, which then displays it. The user can review the plan and customize it as needed.
[1292] 4. User state recognition using an emotion engine
[1293] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and sends this data to a server. The server uses an "emotion engine" to analyze the data and recognize the user's emotional state. For example, if the user is smiling or looks tired while reviewing a plan, the emotion engine will analyze that data.
[1294] 5. Adjusting plans based on emotions
[1295] Based on the perceived emotions, the server may adjust the travel route and itinerary in real time. For example, if the user appears tired, the plan may be changed to include more rest periods or visits to relaxing locations. The adjusted plan is then sent back to the user's device.
[1296] 6. Accumulation and utilization of long-term emotional data
[1297] The server records the user's emotional state over the long term and stores this data in a database. This accumulated emotional data can be used to optimize future travel suggestions and propose plans that are more tailored to the user's preferences.
[1298] Specific example
[1299] For example, if a user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system, the server analyzes these photos to identify Mt. Fuji and Tokyo Tower. The user then inputs, "I want to go to Kyoto via Photo A and Photo B," and specifies the length of stay at each location. The server generates the optimal route and creates a detailed schedule considering transportation methods and travel time. The generated plan is displayed on the terminal for the user to review. At this time, the terminal uses the camera and microphone to collect the user's facial expressions and voice, which the server analyzes with an emotion engine to adjust the plan. For example, if the user inputs that they want to stop by an Osaka product exhibition on the way back, the plan will be regenerated.
[1300] Example of a prompt
[1301] "Please generate the optimal travel plan based on the following instructions. I would like to travel to Kyoto, passing through the foot of Mt. Fuji and Tokyo Tower along the way. I will stay at the foot of Mt. Fuji for one day and at Tokyo Tower for two days."
[1302] "Based on user sentiment data, please suggest a day trip plan that includes plenty of breaks."
[1303] The above describes specific embodiments for carrying out the present invention.
[1304] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1305] Program processing flow and detailed explanation
[1306] Image input and analysis
[1307] Step 1:
[1308] Users upload images of places they want to visit to the system from their smartphones or PCs. Input is image files (e.g., JPEG or PNG format). Output is the image data being saved on the device. Specifically, the user uses the system's upload function, selects an image from a browser or dedicated app, and clicks the upload button.
[1309] Step 2:
[1310] The device sends uploaded image data to the server. The input is the image file data, and the output is the image data received by the server. The device uses an internet connection to send the image data to the server as packets. Specifically, it sends the image data using an HTTP request.
[1311] Step 3:
[1312] The server uses image recognition technology to analyze the input image. The input is the received image data, and the output is the geographical location information of the analysis result. Specifically, the server uses Google's "image recognition technology" to execute an algorithm that extracts image features and determines the geographical location.
[1313] Step 4:
[1314] The server identifies the geographic location based on the analysis results and stores this information in a database. The input is geographic location information, and the output is the geographic location information stored in the database. The server performs an insert operation into the database to persist the identified location data.
[1315] Travel plan generation
[1316] Step 1:
[1317] The user enters text such as "I want to go to XX via photo A and photo B" and specifies the length of stay at each location. The input is text data (intermediate points and length of stay). The output is the instruction information entered into the terminal. Specifically, the user enters the desired intermediate points and length of stay in the text input field and clicks the submit button.
[1318] Step 2:
[1319] The terminal sends user instructions to the server. The input is text data (intermediate locations and length of stay), and the output is the instructions received by the server. The terminal sends the instructions to the server using an HTTP request.
[1320] Step 3:
[1321] The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route. The input is the user's instruction information (waypoints, starting point, destination, length of stay), and the output is the generated optimal route information. Specifically, the server uses a Geographic Information System (GIS) to calculate the optimal route connecting each point using an algorithm.
[1322] Step 4:
[1323] The server calculates the mode of transport and travel time to determine a detailed schedule. The input is the generated optimal route information, and the output is a detailed schedule (mode of transport, departure time, arrival time, etc.). Specifically, the server uses a traffic information API to calculate the travel time for each segment of the journey.
[1324] Step 5:
[1325] The server generates a travel plan and sends it to the user's terminal. The input is a detailed schedule, and the output is the travel plan displayed on the terminal. The server sends the detailed schedule to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[1326] User state recognition using an emotion engine
[1327] Step 1:
[1328] The device uses its built-in camera and microphone to collect user facial expressions and audio data. The input is the user's facial expressions and voice, and the output is the collected data. Specifically, the device's app activates the camera, takes a picture of the user's face, and records audio using the microphone.
[1329] Step 2:
[1330] The terminal sends the collected data to the server. The input consists of facial expressions and voice data, and the output is the data received by the server. The terminal encrypts the data and sends it to the server via an HTTP request.
[1331] Step 3:
[1332] The server uses an emotion engine to analyze data and recognize the user's emotional state. The input is facial expression and voice data, and the output is recognized emotional state information. Specifically, the server applies an emotion recognition algorithm to analyze the user's emotions (e.g., joy, sadness, fatigue, etc.).
[1333] Adjusting plans based on emotions
[1334] Step 1:
[1335] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. The input is the recognized emotional state information, and the output is the adjusted travel plan. Specifically, the server makes adjustments such as changing the schedule to include more rest.
[1336] Step 2:
[1337] The server sends the adjusted plan to the user's terminal. The input is the adjusted travel plan, and the output is the adjusted plan displayed on the terminal. The server sends the adjusted plan to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[1338] Plan presentation and customization
[1339] Step 1:
[1340] The user reviews the presented travel plan. The input is the travel plan displayed on the device, and the output is the user's review process. Specifically, the user views the page displaying the travel plan.
[1341] Step 2:
[1342] If a user wishes to change their plan, they enter that information. The input is the instruction for the desired change, and the output is the change instruction entered into the terminal. For example, they might enter, "I would like to stop by the Osaka product exhibition on my way home."
[1343] Step 3:
[1344] The terminal sends the user's change instructions to the server. The input is the information of the desired change, and the output is the change instructions received by the server. The terminal sends the change instructions to the server using an HTTP request.
[1345] Step 4:
[1346] The server regenerates the plan to reflect the changed information. The input is the change instruction information, and the output is the regenerated plan. The server recalculates the optimal route and detailed schedule and generates a new plan.
[1347] Step 5:
[1348] The server sends the updated plan to the user's terminal. The input is the regenerated plan, and the output is the updated plan displayed on the terminal. The server sends the updated plan to the terminal via an HTTP response, which the terminal receives and displays.
[1349] Use of long-term emotional data
[1350] Step 1:
[1351] The server records the user's emotional state over the long term. The input is recognized emotional state information, and the output is emotional data stored in a database. The server performs insert operations into the database to save the emotional data.
[1352] Step 2:
[1353] The server optimizes the next travel suggestion based on accumulated historical data. The input is long-term sentiment data, and the output is an optimized travel suggestion. The server analyzes the sentiment data and proposes the best travel plan based on the user's priorities and preferences.
[1354] The above outlines the specific processing steps of the system.
[1355] (Application Example 2)
[1356] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1357] Traditional travel planning systems have been somewhat effective in creating travel plans that meet users' preferences, but they have a weakness in their inability to adjust plans to reflect real-time emotional states during the trip. Furthermore, means for users to smartly review their plans during their trip, and visually supportive devices, are still not fully utilized. Therefore, there is a need to reduce stress during travel and improve the user experience.
[1358] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1359] In this invention, the server includes means for inputting an image, analysis means for identifying a geographical location from the input image, plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, presentation means for presenting the generated travel route and detailed schedule to the user, suggestion means for making additional suggestions based on the user's preferences, means including an emotion engine that analyzes the user's facial expressions and voice data and adjusts the travel plan based on their emotional state, and means for presenting the travel plan to the user using a real-time adaptable mobile or visual device. This makes it possible to adjust the plan in real time to reflect the user's emotional state during the trip and to improve the user's travel experience through visual support from a smart device.
[1360] "Image input means" refers to a device or software that receives image data input by a user and transmits it to a server.
[1361] "Analysis means" refers to a device or software that identifies a specific geographical location from input image data and provides that information to a server.
[1362] A "plan generation means" refers to a device or software that creates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[1363] "Presentation means" refers to a device or software that provides the user with the generated travel route and detailed schedule visually or audibly.
[1364] "Suggestion means" refers to a device or software that provides additional information or recommended plans regarding travel routes based on the user's preferences.
[1365] "Means including an emotion engine" refers to a device or software that analyzes the user's facial expressions and voice data to recognize their emotional state and adjust the travel plan accordingly.
[1366] "Means for presenting a travel plan to a user using a mobile or visual device" refers to a device or software that visually presents a travel plan while reflecting the user's location information and emotional state in real time.
[1367] "Real-time" refers to a state where the user's current status and location information are reflected in real time, allowing for immediate adaptation or modification.
[1368] This invention is a system for efficiently creating travel plans and adjusting them in real time according to the user's emotional state. The system includes means for image input, analysis, plan generation, presentation, suggestion, an emotion engine, and means for presenting the travel plan to the user using a mobile device or visual device.
[1369] Description of the program's overall operation
[1370] 1. Image input and analysis
[1371] Users upload images of places they want to visit from their smartphones or PCs to the system in order to create travel plans. In this case, the image input mechanism is activated. The device sends these images to the server, which uses image recognition technology (e.g., a common image recognition API) to determine the geographical location of the images. The results of this analysis are stored in a database.
[1372] 2. Plan generation and user specification
[1373] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location. Based on this, the terminal sends the instructions to the server. The server generates the optimal travel route and detailed schedule based on the analyzed geographic location information and the user's instructions. This generation process uses a route calculation engine and transportation information.
[1374] 3. Collection and analysis of facial expressions and voice.
[1375] The device uses its built-in camera and microphone to collect user facial expressions and voice data. This data is sent to a server in real time, where the server uses an emotion engine to analyze the user's emotional state. Specifically, facial expressions are analyzed using facial recognition software (e.g., OpenCV or dlib), and the data is sent to an emotion analysis API.
[1376] 4. Adjusting plans based on emotions
[1377] Based on the analysis results from the emotion engine, the server adjusts the travel plan to match the user's emotional state. For example, if the user is tired, the plan will be changed to include relaxing facilities. Conversely, if the user is excited, the plan can be adjusted to include more active spots.
[1378] 5. Real-time plan presentation
[1379] The travel plan, generated in real time, is presented to the user through visual devices such as smart glasses. This allows the user to enjoy their trip while checking a plan that suits their emotional state in real time.
[1380] Specific example
[1381] The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a photo of an urban landmark (Photo B) to the system.
[1382] The server identifies photo A as the foot of Mount Fuji and photo B as an urban area, and saves them in the database.
[1383] The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[1384] The terminal sends instruction information to the server, which then generates the optimal route and a detailed schedule.
[1385] The device uses its built-in camera and microphone to collect the user's facial expressions and voice in real time as they review the plan, and transmits this information to the server.
[1386] The server uses an emotion engine to analyze the user's emotions and adjust the plan accordingly. For example, if it determines that the user is tired, it will change the plan to include more rest time.
[1387] Example of a prompt
[1388] "Since you seem tired this time, please suggest a travel plan that includes places where you can relax."
[1389] This system allows for real-time adjustments to travel plans that reflect the user's emotional state during their trip, and enhances the user's travel experience through visual support via smart devices.
[1390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1391] Step 1:
[1392] The user uses their device to upload images of the place they want to go to the system.
[1393] The specific input is the user's image data, which the device sends to the server. The server uses image recognition technology to analyze the uploaded image. For example, it uses an image recognition API to determine the geographical location and stores this information in a database. The output is the analyzed geographical location information.
[1394] Step 2:
[1395] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location.
[1396] The input is user text instructions, which the terminal sends to the server. The server generates the optimal travel route and detailed schedule based on the starting point, analyzed geographic location information, and user instructions. It uses a route calculation engine to determine the best route, taking into account transportation methods and travel time. The output is the generated travel plan and detailed schedule.
[1397] Step 3:
[1398] The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[1399] The input consists of the user's real-time facial expressions and voice data, which the terminal collects and sends to the server. The server uses an emotion engine to analyze the data and recognize the user's emotional state. Specifically, it analyzes facial expressions using facial recognition software (e.g., OpenCV or dlib) and sends the data to an emotion analysis API to identify the emotional state. The output is the user's emotional state as determined by the emotion engine.
[1400] Step 4:
[1401] The server adjusts the travel plan in real time according to the user's emotional state, based on the analysis results from the emotion engine.
[1402] The input consists of the user's emotional state, as determined by the emotion engine, and the existing travel plan. Based on this data, the server regenerates the plan, including new suggestions and modifications. For example, if the user is tired, the server might add rest areas and modify the plan to include relaxing facilities. The output is the adjusted travel plan.
[1403] Step 5:
[1404] The device presents the user with travel plans that have been generated or adjusted in real time.
[1405] The input is a pre-configured travel plan, and the device visually presents this information to the user via smart glasses or a smartphone. The user can continue their trip while checking the plan in real time to match their emotional state. The output is the pre-configured travel plan displayed on the user's device.
[1406] Step 6:
[1407] The server records the user's emotional state over the long term to optimize future travel suggestions.
[1408] The input consists of past travel plans and user emotional state data, which the server stores in a database. This data is used to generate optimal travel plans based on the user's preferences for future trips. The output is a more personalized suggestion for the next trip.
[1409] Through these steps, users can enjoy travel plans tailored to their emotional state in real time, enhancing their travel experience.
[1410] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1411] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1412] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1413] [Fourth Embodiment]
[1414] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1415] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1416] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1417] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1418] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1420] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1421] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1422] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1423] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1424] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1425] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1426] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1427] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. The system of this invention includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[1428] System Overview
[1429] 1. Image Input and Analysis
[1430] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[1431] 2. Plan generation and user specification
[1432] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[1433] 3. Generating the optimal route and detailed schedule
[1434] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[1435] 4. Customize your plan
[1436] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[1437] 5. Suggestions based on user preferences
[1438] The server analyzes the user's preferences based on their past usage history and current travel plan, and incorporates this into future recommendations. This enables customized suggestions tailored to the user's preferences.
[1439] Explain the program's processing in natural language.
[1440] 1. Image input and analysis
[1441] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[1442] 2. The device sends the uploaded image data to the server.
[1443] 3. The server uses image recognition technology to analyze specific locations within the image.
[1444] 4. The server saves the analysis results to the database.
[1445] 2. Creating a travel plan
[1446] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[1447] 2. The terminal sends instructions from the user to the server.
[1448] 3. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and destination XX, and generates a detailed schedule.
[1449] 4. The server sends the generated plan to the user's terminal.
[1450] 3. Plan presentation and customization
[1451] 1. The terminal displays the detailed travel plan received from the server to the user.
[1452] 2. The user reviews the proposed plan and instructs the user to make changes as needed (e.g., add a route through a local products fair).
[1453] 3. The terminal sends a change instruction to the server.
[1454] 4. The server updates the plan, taking the changes into account, and sends it back to the user's device.
[1455] 4. Additional proposals and historical analysis
[1456] 1. The server analyzes the user's preferences based on the database and incorporates them into suggestions for the next travel plan.
[1457] 2. When users receive new suggestions, it becomes easier for them to explore further preferred locations and routes.
[1458] Specific example
[1459] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[1460] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[1461] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[1462] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[1463] 5. The server sends the generated plan to the user, who then confirms it.
[1464] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[1465] 7. The server is updated, the plan is regenerated, and sent to the user's terminal.
[1466] The above describes the specific embodiments and program processing details for implementing the present invention.
[1467] The following describes the processing flow.
[1468] Step 1:
[1469] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[1470] Step 2:
[1471] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[1472] Step 3:
[1473] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[1474] Step 4:
[1475] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[1476] Step 5:
[1477] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[1478] Step 6:
[1479] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[1480] Step 7:
[1481] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[1482] Step 8:
[1483] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[1484] Step 9:
[1485] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[1486] Step 10:
[1487] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[1488] Step 11:
[1489] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[1490] Step 12:
[1491] The server stores the user's past usage history and current plan data in a database and analyzes the user's preferences. This allows for the creation of customized travel suggestions that are tailored to the user's tastes.
[1492] (Example 1)
[1493] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1494] Traditional travel planning systems had the problem of requiring users to manually input detailed information about the places they wanted to visit, which was time-consuming. Furthermore, the generation of travel routes and schedules relied heavily on user input, lacking automation. As a result, users were unable to create travel plans efficiently, resulting in a time-consuming and cumbersome process. Additionally, the lack of customized suggestions based on user preferences made it difficult to improve user satisfaction.
[1495] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1496] In this invention, the server includes means for inputting an image, means for analyzing the input image to identify a geographical location, means for generating a plan that generates a travel route based on the analyzed geographical location information and user instructions, means for presenting the generated travel route and detailed schedule to the user, and means for making additional suggestions based on the user's preferences. This makes it possible for users to easily upload an image and have a travel route and detailed schedule automatically generated, enabling the provision of efficient and customized travel plans.
[1497] "Means of inputting images" refers to the function that allows users to upload image data to the system via their device.
[1498] "Means for analyzing input images to determine geographical location" refers to technologies that analyze uploaded image data to identify geographical information (e.g., longitude and latitude) of the location shown in the image.
[1499] The "means for generating plans" refer to a function that automatically creates the optimal travel route and detailed schedule for visiting multiple locations, based on analyzed geographical location information and user instructions.
[1500] "Means of presentation" refers to functions that visually display the generated travel route and detailed schedule to the user.
[1501] The "suggestion method" refers to a function that provides additional suggestions tailored to the user's preferences, based on the user's past travel history and current travel plan.
[1502] "Image recognition technology" is a technique that uses machine learning and algorithms to identify specific objects or locations within an image.
[1503] "A means for users to set the duration of their stay at each location" refers to a function that allows users to set the number of days or hours they will stay at each designated location.
[1504] This invention relates to a system that identifies a desired destination based on images provided by a user and generates a detailed travel route and schedule. The system includes image input means, analysis means, plan generation means, presentation means, and suggestion means.
[1505] First, the user uploads an image of the place they want to go to their device (such as a smartphone or PC). The device then sends this image data to the server. The server uses image recognition technology (such as "Gemini") to analyze the geographical location of the uploaded image and stores the analysis results in a database.
[1506] Next, the user uses text input to specify their desired destination, indicating that they want to travel via designated locations, and to specify the duration of their stay at each point. Upon receiving these instructions, the terminal sends the user's information to the server. The server searches for the optimal route from the starting point (current location) to specific locations A and B, and then to the final destination XX, generating a detailed schedule that takes into account transportation methods and travel time. The generated plan is immediately presented to the user's terminal via the internet.
[1507] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can enter that information. The terminal sends this information to the server, which then updates the plan.
[1508] Furthermore, the server can analyze the user's preferences based on their past usage history and current travel plans, and reflect this in future travel suggestions. This enables customized suggestions tailored to the user's preferences.
[1509] Specific example
[1510] 1. The user uploads a landscape photo of the foot of Mt. Fuji and a night view photo of Tokyo Tower to the system.
[1511] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[1512] 3. The user enters "I want to go to Kyoto via Photo A and Photo B" and specifies "I want to stay for 1 day at Photo A and 2 days at Photo B".
[1513] 4. The terminal sends instruction information to the server, which generates the optimal route and detailed schedule.
[1514] 5. The server sends the generated plan to the user's terminal, and the user confirms it.
[1515] 6. The user requests to stop by an Osaka product exhibition on their way home, and the terminal sends this request to the server.
[1516] 7. The server regenerates the updated plan and sends it to the user's terminal.
[1517] Examples of prompts to input into a generative AI model
[1518] A user has uploaded landscape photos of the foot of Mt. Fuji and nighttime photos of Tokyo Tower to the system. Please use these photos to create a travel plan to Kyoto.
[1519] I stayed for one day at the foot of Mt. Fuji.
[1520] I stayed at Tokyo Tower for two days.
[1521] The final destination of the trip is Kyoto
[1522] I'd like to stop by the Osaka product fair on my way home.
[1523] Generate the optimal route and detailed schedule, and present them to the user.
[1524] Thus, this invention enables users to easily upload images, and the server automatically generates travel routes and detailed schedules. Furthermore, by providing customized suggestions based on user preferences, it realizes a system that offers more satisfying travel plans.
[1525] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1526] Step 1:
[1527] The user accesses the system and uploads an image of the place they want to go to their device.
[1528] Input: Image of the place you want to go (e.g., a landscape photo of Mt. Fuji)
[1529] Specific operation: The user accesses the system's web page or application and clicks the image upload button. The upload screen appears, and the user selects an image file from their device and sends it to the server.
[1530] Step 2:
[1531] The device sends the uploaded image data to the server.
[1532] Input: Image data uploaded by the user
[1533] Specific operation: The terminal divides the image data into packets and sends them to the server over the network.
[1534] Step 3:
[1535] The server uses image recognition technology to analyze the geographical location of the input image.
[1536] Input: Image data sent to the server
[1537] Specific operation: The server uses image recognition technology (e.g., "Gemini") to analyze images and identify specific landmarks or geographical features within them. The identified geographical locations (e.g., longitude and latitude) are recorded in a database.
[1538] Step 4:
[1539] The server saves the analysis results to the database.
[1540] Input: Geographic location information obtained using image recognition technology
[1541] Specific operation: The server connects to the database and saves the identified geographic coordinates and related information to the "Location" table.
[1542] Step 5:
[1543] The user provides text input indicating that they want to travel to their destination via specific locations and specifies the length of stay at each point.
[1544] Input: User instructions in text format (Example: "I want to spend one day at Mt. Fuji, two days at Tokyo Tower, and then go to my final destination, Kyoto.")
[1545] Specific operation: The user enters the desired location and length of stay into the system's text input field and clicks the submit button.
[1546] Step 6:
[1547] The terminal sends instructions from the user to the server.
[1548] Input: User instructions
[1549] Specific operation: The terminal sends user instructions in text format to the server.
[1550] Step 7:
[1551] The server searches for the optimal route from the starting point to each identified point and the final destination, and generates a detailed schedule.
[1552] Input: User-specified geographical location and length of stay
[1553] Specific operation: The server uses map databases and traffic information to calculate the optimal route from the starting point (current location) to each specified point and the final destination. It then creates a planned schedule, taking into account the length of stay.
[1554] Step 8:
[1555] The server sends the generated plan to the user's device.
[1556] Input: Detailed travel plan (route and schedule)
[1557] Specific operation: Format the generated plan and send it to the user's terminal via the internet.
[1558] Step 9:
[1559] The terminal displays the detailed travel plan received from the server to the user.
[1560] Input: Travel plan sent from the server
[1561] Specific operation: Visually display the travel route and schedule on the device screen (web page or application).
[1562] Step 10:
[1563] The user reviews the proposed plan and is instructed to make changes as needed.
[1564] Input: Instructions for the change (Example: I want to stop by the Osaka product exhibition on my way home)
[1565] Specific actions: The user reviews the presented plan, enters any necessary changes in text, and clicks the submit button.
[1566] Step 11:
[1567] The terminal sends a change instruction to the server.
[1568] Input: Instructions for change
[1569] Specific action: The terminal sends the user's change instruction to the server.
[1570] Step 12:
[1571] The server updates the plan, taking the changes into account, and sends it back to the user's device.
[1572] Input: Changed instructions
[1573] Specific operation: The server recalculates the travel plan based on the new instructions and sends the updated plan to the user's terminal.
[1574] Step 13:
[1575] The server analyzes user preferences based on a database and incorporates them into suggestions for the next travel plan.
[1576] Input: User's past usage history and current travel plan
[1577] Specific operation: The server analyzes historical data in the database, extracts user preferences and trends, and performs calculations to reflect these in future travel suggestions.
[1578] Step 14:
[1579] When a user receives new suggestions, they will explore further preferred locations and routes.
[1580] Input: New travel suggestions from the server
[1581] Specific operation: The user receives suggestions and considers new travel destinations and routes that suit their preferences.
[1582] (Application Example 1)
[1583] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1584] Traditional food delivery services struggle to identify specific dishes users want, lacking ways to improve the user experience. They also lack efficient support for users searching for specific cuisines or restaurants while traveling or out and about. This results in users having to spend a lot of time choosing meals and deciding on delivery plans, leading to reduced convenience.
[1585] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1586] In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated route and detailed schedule to the user, a suggestion means for making additional suggestions based on the user's preferences, a means for identifying a specific dish and serving location from a food image, and a means for generating an optimal route and delivery plan from the serving location. This makes it possible for the user to identify the restaurant they want to go to or the dish they want to eat using an image, and for the server to efficiently suggest the location and the optimal delivery plan.
[1587] "Image input means" refers to hardware or software functions that allow a user to upload images to a system.
[1588] "Analysis means" refers to technologies for identifying geographical location information and specific dishes from input images.
[1589] A "plan generation means" is a system function for generating routes and delivery plans based on analyzed information and user instructions.
[1590] "Presentation means" refers to methods or devices for displaying the generated plan or detailed schedule to the user.
[1591] The "suggestion method" is a function that makes additional suggestions based on the user's past usage history and preferences.
[1592] "Specific dish" refers to the type or name of food that can be identified from images uploaded by users.
[1593] "Place of service" refers to the location of the restaurant or delivery service where the specific dish is served.
[1594] A "delivery plan" is a plan that includes the optimal delivery route and time from the user's current location to the delivery destination.
[1595] This invention relates to a system that identifies specific locations or dishes based on images input by a user, and efficiently generates and presents plans and delivery plans based on those locations. This system consists of an image input means, an analysis means, a plan generation means, a presentation means, and a suggestion means.
[1596] System Overview
[1597] 1. Image input and analysis
[1598] Users access the system using a smartphone or PC and upload images of restaurants they want to visit or dishes they want to eat. The device sends this image data to the server. The server uses image recognition technology, such as Google's "Gemini," to analyze the specific locations and dishes in the input images and stores that information in a database.
[1599] 2. Plan generation and user specification
[1600] The user instructs the server via text input, "I want the food in the uploaded photo delivered." They also specify the length of stay at each location and the desired delivery time. The device then sends this information to the server.
[1601] 3. Generating the optimal route and detailed schedule
[1602] The server generates the optimal route from the user's current location or starting point to a restaurant serving a specific dish. Considering travel time and transportation methods, the server calculates a detailed delivery plan. The generated plan is immediately presented to the user's device.
[1603] 4. Customize your plan
[1604] Users can review the presented delivery plan and make changes as needed. For example, they can select a different restaurant or add information to their route home if they want to stop at a specific location. The device sends this information to the server, which then updates the plan.
[1605] 5. Suggestions based on user preferences
[1606] The server analyzes the user's past usage history and current plan to understand their preferences and incorporates these into future recommendations. This allows for customized recommendations tailored to the user's tastes.
[1607] Specific processing instructions
[1608] 1. Image input and analysis:
[1609] Users access the system from their smartphones or PCs and upload images of food or restaurants. The device sends the uploaded image data to the server, which analyzes the images using tools such as Google's "Gemini" and stores the information in a database.
[1610] 2. Plan generation and user specification:
[1611] The user instructs the server via text input, "I want the food in the picture delivered," and this information is sent from the terminal to the server. The server then generates the optimal plan based on the restaurants and delivery services that offer the specific dish.
[1612] 3. Generating the optimal route and detailed schedule:
[1613] The server calculates the optimal route from the user's current location to the restaurant and generates a detailed schedule. The generated plan is then displayed on the terminal.
[1614] 4. Customize your plan:
[1615] The user reviews the presented plan and makes changes as needed. The device sends the change information to the server, which then generates and presents the updated plan again.
[1616] 5. Additional proposals and historical analysis:
[1617] The server analyzes the user's past usage history and preferences based on a database to optimize future recommendations.
[1618] Specific example
[1619] For example, if a user uploads a photo of fried rice, the server analyzes the photo and identifies it as fried rice. The user then enters "I want fried rice delivered," and the server generates the optimal delivery plan from a specific restaurant. The user can review this plan and request changes if necessary. The server also generates a plan that reflects any stops along the return journey.
[1620] Example of a prompt
[1621] "Upload a photo of fried rice and suggest restaurants that offer delivery of that dish."
[1622] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1623] Step 1:
[1624] Uploading and sending images
[1625] Users log in to the system using their smartphones or PCs and upload images of the food they want to eat or the restaurants they want to visit. The device receives the image data sent by the user and sends it to the server. The input here is the image taken by the user, and the output is the image data sent to the server.
[1626] Step 2:
[1627] Image analysis
[1628] The server uses image recognition technologies such as Google's "Gemini" to analyze the received image data. In this process, the image data serves as input, and features and content within the image are recognized during the analysis process, identifying specific dishes or locations. The analysis results are stored in a database as geographical location information and information about the specific dish.
[1629] Step 3:
[1630] Text input and sending instructions
[1631] The user enters a text instruction via the system interface, such as "I want the food in the uploaded image delivered." They can also enter additional information, such as their desired delivery time. The terminal then sends these instructions to the server. The input is the user's text instruction, and the output is the data sent to the server.
[1632] Step 4:
[1633] Generating a delivery plan
[1634] The server generates the optimal delivery route and schedule based on the user's current location and information on the nearest restaurants that serve the specified dish. Analyzed information and user instructions serve as input, and the optimal route and detailed schedule are output. This calculation also utilizes data such as transportation methods and travel time.
[1635] Step 5:
[1636] Presentation of the plan
[1637] The generated delivery plan and detailed schedule are sent from the server to the terminal, which then displays them to the user. The input here is the plan data generated on the server side, and the output is the plan information displayed on the user's terminal.
[1638] Step 6:
[1639] Customize your plan
[1640] The user reviews the presented delivery plan and makes changes as needed, such as selecting a different restaurant or specifying additional transit points. The terminal sends the user's change instructions to the server, which then regenerates the plan based on this information. The input is the user's change instructions, and the output is the updated plan.
[1641] Step 7:
[1642] Presentation and approval of the final plan
[1643] The server generates the updated plan and sends it back to the user's device. The user approves the final plan, and the plan is finalized. The input here is the new plan data, and the output is the final plan display and approval on the user's device.
[1644] Step 8:
[1645] Additional proposals and historical analysis
[1646] The server makes delivery suggestions for future deliveries based on the user's past usage history and preferences. In this process, past usage data is used as input, and suggestion data based on the user's preferences is output. The suggested content will be presented to the user the next time they use the service.
[1647] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1648] This invention is a system for efficiently creating travel plans, which identifies desired destinations based on images provided by the user and generates a detailed travel route and schedule. Furthermore, this invention incorporates an emotion engine that recognizes the user's emotions and includes a function to adjust the plan according to the user's emotional state. The system of this invention includes image input means, analysis means, plan generation means, presentation means, suggestion means, and an emotion engine.
[1649] System Overview
[1650] 1. Image Input and Analysis
[1651] First, the user inputs an image of the place they want to go into the system. The device then sends this image to the server. The server uses Google's "Gemini" or other image recognition technologies to analyze and identify the geographical location of the input image.
[1652] 2. Plan generation and user specification
[1653] The user uses text input to specify that they want to reach their destination via designated locations. They can also specify the length of stay at each location. Upon receiving these instructions, the terminal sends the information to the server.
[1654] 3. Generating the optimal route and detailed schedule
[1655] The server generates the optimal route from the user's current location or starting point to each specified point and the final destination. Considering the mode of transport and travel time, the server calculates a detailed schedule. The generated plan is immediately presented to the user's device.
[1656] 4. User state recognition using an emotion engine
[1657] This invention incorporates an emotion engine that analyzes the user's facial expressions and voice, and incorporates this into the travel plan. The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and transmits this data to a server. The server analyzes the collected data and recognizes the user's emotional state.
[1658] 5. Adjusting plans based on emotions
[1659] Based on the perceived emotions, the server adjusts travel routes and itineraries in real time. For example, if a user appears tired, the plan can be changed to include more rest periods.
[1660] 6. Plan presentation and customization
[1661] Users can review the presented travel plan and make changes as needed. For example, if they want to stop by a specific local products fair on their way back, they can add that information. The terminal sends this information to the server, which then updates the plan.
[1662] 7. Use of long-term emotional data
[1663] The server records the user's emotional state over the long term and uses this history to optimize future travel suggestions. This historical data is used to further refine the user's preferred types of travel plans.
[1664] Explain the program's processing in natural language.
[1665] 1. Image input and analysis
[1666] 1. Users access the system from a device such as a smartphone or PC and upload an image of the place they want to go.
[1667] 2. The device sends the uploaded image data to the server.
[1668] 3. The server uses image recognition technology to analyze specific locations within the image.
[1669] 4. The server identifies the geographical location based on the analysis results and stores this information in the database.
[1670] 2. Creating a travel plan
[1671] 1. The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the number of days to stay at each location.
[1672] 2. The terminal sends instructions from the user to the server.
[1673] 3. The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route.
[1674] 4. The server calculates the mode of transport and travel time, and determines the detailed schedule.
[1675] 5. The server sends the generated travel plan to the user's device.
[1676] 3. User state recognition using an emotion engine
[1677] 1. The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[1678] 2. The device sends the collected data to the server.
[1679] 3. The server uses an emotion engine to analyze the data and recognize the user's emotional state.
[1680] 4. Adjusting plans based on emotions
[1681] 1. The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine.
[1682] 2. For example, if a user is feeling stressed, the system can regenerate a plan to visit places or facilities where they can relax.
[1683] 3. The server sends the adjusted plan to the user's terminal.
[1684] 5. Plan presentation and customization
[1685] 1. The user reviews the travel plan presented via their device and is instructed to make changes as needed.
[1686] 2. The terminal sends the user's change instruction to the server.
[1687] 3. The server updates the plan to reflect the changed information and sends it to the user's device again.
[1688] 6. Use of long-term emotional data
[1689] 1. The server records the user's emotional state over the long term and stores it in a database.
[1690] 2. The server optimizes future travel suggestions based on accumulated historical data and analyzes user preferences in a more precise manner.
[1691] Specific example
[1692] 1. The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system.
[1693] 2. The server identifies photo A as the foot of Mount Fuji and photo B as Tokyo Tower, and saves them in the database.
[1694] 3. The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[1695] 4. The terminal sends instruction information to the server, and the server generates the optimal route and detailed schedule.
[1696] 5. The server sends the generated plan to the user, who then confirms it.
[1697] 6. The device uses its built-in camera and microphone to collect facial expressions and audio recordings of the user as they review the plan, and transmits them to the server.
[1698] 7. The server analyzes the user's emotions using an emotion engine and adjusts the plan accordingly. For example, if it determines that the user is tired, it changes the plan to include more rest time.
[1699] 8. The user enters that they want to stop by the Osaka product exhibition on their way home, and the terminal sends the information to the server.
[1700] 9. The server regenerates the modified plan and sends it to the user's terminal.
[1701] 10. The server records the user's emotional state over the long term and uses it to improve future travel suggestions.
[1702] The above describes the specific embodiments and program processing details for implementing the present invention.
[1703] The following describes the processing flow.
[1704] Step 1:
[1705] The user uploads images (Photo A and Photo B) of the place they want to go to the system. The user selects the images through the system interface from their smartphone or PC and presses the upload button.
[1706] Step 2:
[1707] The device sends the selected image data to the server. The image data is sent to the server via the internet and temporarily stored on the server side.
[1708] Step 3:
[1709] The server uses Google's Gemini and other image recognition technologies to begin analyzing the uploaded images. The server extracts features from each image and matches them against known geographical location data in the database.
[1710] Step 4:
[1711] The server identifies specific geographical locations based on the image analysis results. For example, it might identify that photo A is at the foot of Mount Fuji and photo B is near Tokyo Tower. The identified information is stored in a database.
[1712] Step 5:
[1713] The user specifies their destination via text input, such as "I want to go to XX via photo A and photo B," and also enters their request including the length of stay at each location. The terminal provides an input form for the user to enter the details.
[1714] Step 6:
[1715] The device sends the user's text instructions and length of stay to the server. The server receives this data and organizes information about the departure point, each specific location, and the destination.
[1716] Step 7:
[1717] The server analyzes the optimal route from the starting point to photo A, photo B, and destination XX, and calculates the mode of transport and travel time. It generates the optimal travel route and time schedule, taking into account different modes of transport such as trains, buses, and cars, and the travel time.
[1718] Step 8:
[1719] The server sends a detailed travel plan to the user's device. The plan includes information such as departure time, arrival time, mode of transportation, and duration of stay at each location.
[1720] Step 9:
[1721] The user reviews the travel plan presented through their device. They can then specify additional places they wish to visit or stop at on their return journey, if necessary. For example, they might enter that they would like to visit an Osaka product exhibition on their way back.
[1722] Step 10:
[1723] The terminal sends the user's change instructions to the server. The server re-analyzes the plan based on the additional information and generates an updated optimal route and schedule.
[1724] Step 11:
[1725] The server resends the updated plan to the user's device. The user reviews the new plan and makes their final travel decision.
[1726] Step 12:
[1727] The device uses its built-in camera and microphone to collect facial expressions and audio data as the user reviews the plan. This data is collected only with the user's permission.
[1728] Step 13:
[1729] The device sends collected facial and audio data to the server. The server uses an emotion engine to analyze the user's emotional state.
[1730] Step 14:
[1731] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. For example, if the user looks tired, it might increase the rest time.
[1732] Step 15:
[1733] The server sends the adjusted plan to the user's device. The user can review the new plan and confirm it again before starting their trip.
[1734] Step 16:
[1735] The server records the user's emotional state over the long term and stores it in a database. When planning the next trip, the server takes the user's past emotional data into consideration to provide more customized suggestions.
[1736] The above outlines the specific processing steps of the invention that combines an emotion engine.
[1737] (Example 2)
[1738] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1739] Traditional travel planning systems have limitations in identifying desired destinations based on user-provided images and generating optimal travel routes that include those locations. Furthermore, they lack the functionality to adjust plans based on the user's emotional state during the trip, making it difficult to mitigate stress and dissatisfaction. Additionally, they haven't adequately provided mechanisms for optimizing future travel plans using long-term user emotional data.
[1740] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an image input means, an analysis means for identifying a geographical location from the input image, a plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, a presentation means for presenting the generated travel route and detailed schedule to the user, an emotion engine means for analyzing the user's emotional state and adjusting the travel plan, and a suggestion means for making additional suggestions based on the user's preferences. This makes it possible to identify desired destinations and generate optimal travel routes based on images provided by the user, and furthermore, to adjust the travel plan in real time according to the user's emotional state and optimize future travel plans using long-term emotional data.
[1741] The "image input method" is a function that allows users to upload images of places they want to visit to the system.
[1742] "Analysis means" refers to a function for determining geographical location from input image data.
[1743] The "plan generation method" is a function that generates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[1744] "Presentation means" refers to a function for displaying the generated travel route and detailed schedule to the user.
[1745] The "emotional engine" is a function that analyzes the user's facial expressions and voice to adjust the travel plan in real time.
[1746] A "suggestion mechanism" is a function that provides additional suggestions based on the user's preferences.
[1747] "Image recognition technology" is a technology that analyzes the content of an input image and extracts specific information from it.
[1748] "Means for setting the length of stay" refers to a function that allows users to specify the number of days or hours they will stay at each visited location.
[1749] This invention is a system that identifies desired destinations based on images provided by the user and generates the optimal travel route and detailed schedule that passes through those locations. Furthermore, it incorporates a function that recognizes the user's emotions and adjusts the travel plan according to their emotional state. This invention includes the following main functions:
[1750] 1. Image input and analysis
[1751] Users upload images of places they want to visit to the system from their smartphones, PCs, or other devices. The device then sends this image data to the server. The server uses "image recognition technology" to analyze and identify the geographical location of the input image. For example, if a user uploads a photo of the foot of Mount Fuji, the server analyzes the photo to identify the foot of Mount Fuji.
[1752] 2. Creating a travel plan
[1753] The user specifies their desired destination, such as "I want to go to XX via photo A and photo B," through text input, and also specifies the length of stay at each location. The terminal receives this information and sends it to the server. The server generates the optimal route based on the information about the starting point, specific locations, and destination. For example, if the user wants to go from the foot of Mt. Fuji to Kyoto via Tokyo Tower, the server will calculate a detailed schedule considering the means of transportation and travel time at each location.
[1754] 3. Presentation of the generated plan
[1755] The server sends the generated travel plan to the user's device, which then displays it. The user can review the plan and customize it as needed.
[1756] 4. User state recognition using an emotion engine
[1757] The device uses its built-in camera and microphone to collect the user's facial expressions and voice, and sends this data to a server. The server uses an "emotion engine" to analyze the data and recognize the user's emotional state. For example, if the user is smiling or looks tired while reviewing a plan, the emotion engine will analyze that data.
[1758] 5. Adjusting plans based on emotions
[1759] Based on the perceived emotions, the server may adjust the travel route and itinerary in real time. For example, if the user appears tired, the plan may be changed to include more rest periods or visits to relaxing locations. The adjusted plan is then sent back to the user's device.
[1760] 6. Accumulation and utilization of long-term emotional data
[1761] The server records the user's emotional state over the long term and stores this data in a database. This accumulated emotional data can be used to optimize future travel suggestions and propose plans that are more tailored to the user's preferences.
[1762] Specific example
[1763] For example, if a user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a night view photo of Tokyo Tower (Photo B) to the system, the server analyzes these photos to identify Mt. Fuji and Tokyo Tower. The user then inputs, "I want to go to Kyoto via Photo A and Photo B," and specifies the length of stay at each location. The server generates the optimal route and creates a detailed schedule considering transportation methods and travel time. The generated plan is displayed on the terminal for the user to review. At this time, the terminal uses the camera and microphone to collect the user's facial expressions and voice, which the server analyzes with an emotion engine to adjust the plan. For example, if the user inputs that they want to stop by an Osaka product exhibition on the way back, the plan will be regenerated.
[1764] Example of a prompt
[1765] "Please generate the optimal travel plan based on the following instructions. I would like to travel to Kyoto, passing through the foot of Mt. Fuji and Tokyo Tower along the way. I will stay at the foot of Mt. Fuji for one day and at Tokyo Tower for two days."
[1766] "Based on user sentiment data, please suggest a day trip plan that includes plenty of breaks."
[1767] The above describes specific embodiments for carrying out the present invention.
[1768] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1769] Program processing flow and detailed explanation
[1770] Image input and analysis
[1771] Step 1:
[1772] Users upload images of places they want to visit to the system from their smartphones or PCs. Input is image files (e.g., JPEG or PNG format). Output is the image data being saved on the device. Specifically, the user uses the system's upload function, selects an image from a browser or dedicated app, and clicks the upload button.
[1773] Step 2:
[1774] The device sends uploaded image data to the server. The input is the image file data, and the output is the image data received by the server. The device uses an internet connection to send the image data to the server as packets. Specifically, it sends the image data using an HTTP request.
[1775] Step 3:
[1776] The server uses image recognition technology to analyze the input image. The input is the received image data, and the output is the geographical location information of the analysis result. Specifically, the server uses Google's "image recognition technology" to execute an algorithm that extracts image features and determines the geographical location.
[1777] Step 4:
[1778] The server identifies the geographic location based on the analysis results and stores this information in a database. The input is geographic location information, and the output is the geographic location information stored in the database. The server performs an insert operation into the database to persist the identified location data.
[1779] Travel plan generation
[1780] Step 1:
[1781] The user enters text such as "I want to go to XX via photo A and photo B" and specifies the length of stay at each location. The input is text data (intermediate points and length of stay). The output is the instruction information entered into the terminal. Specifically, the user enters the desired intermediate points and length of stay in the text input field and clicks the submit button.
[1782] Step 2:
[1783] The terminal sends user instructions to the server. The input is text data (intermediate locations and length of stay), and the output is the instructions received by the server. The terminal sends the instructions to the server using an HTTP request.
[1784] Step 3:
[1785] The server organizes information about the starting point, each specific location, and the destination, and generates the optimal route. The input is the user's instruction information (waypoints, starting point, destination, length of stay), and the output is the generated optimal route information. Specifically, the server uses a Geographic Information System (GIS) to calculate the optimal route connecting each point using an algorithm.
[1786] Step 4:
[1787] The server calculates the mode of transport and travel time to determine a detailed schedule. The input is the generated optimal route information, and the output is a detailed schedule (mode of transport, departure time, arrival time, etc.). Specifically, the server uses a traffic information API to calculate the travel time for each segment of the journey.
[1788] Step 5:
[1789] The server generates a travel plan and sends it to the user's terminal. The input is a detailed schedule, and the output is the travel plan displayed on the terminal. The server sends the detailed schedule to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[1790] User state recognition using an emotion engine
[1791] Step 1:
[1792] The device uses its built-in camera and microphone to collect user facial expressions and audio data. The input is the user's facial expressions and voice, and the output is the collected data. Specifically, the device's app activates the camera, takes a picture of the user's face, and records audio using the microphone.
[1793] Step 2:
[1794] The terminal sends the collected data to the server. The input consists of facial expressions and voice data, and the output is the data received by the server. The terminal encrypts the data and sends it to the server via an HTTP request.
[1795] Step 3:
[1796] The server uses an emotion engine to analyze data and recognize the user's emotional state. The input is facial expression and voice data, and the output is recognized emotional state information. Specifically, the server applies an emotion recognition algorithm to analyze the user's emotions (e.g., joy, sadness, fatigue, etc.).
[1797] Adjusting plans based on emotions
[1798] Step 1:
[1799] The server adjusts the travel plan to reflect the user's current state based on the analysis results from the emotion engine. The input is the recognized emotional state information, and the output is the adjusted travel plan. Specifically, the server makes adjustments such as changing the schedule to include more rest.
[1800] Step 2:
[1801] The server sends the adjusted plan to the user's terminal. The input is the adjusted travel plan, and the output is the adjusted plan displayed on the terminal. The server sends the adjusted plan to the terminal via an HTTP response, which the terminal receives and displays on its screen.
[1802] Plan presentation and customization
[1803] Step 1:
[1804] The user reviews the presented travel plan. The input is the travel plan displayed on the device, and the output is the user's review process. Specifically, the user views the page displaying the travel plan.
[1805] Step 2:
[1806] If a user wishes to change their plan, they enter that information. The input is the instruction for the desired change, and the output is the change instruction entered into the terminal. For example, they might enter, "I would like to stop by the Osaka product exhibition on my way home."
[1807] Step 3:
[1808] The terminal sends the user's change instructions to the server. The input is the information of the desired change, and the output is the change instructions received by the server. The terminal sends the change instructions to the server using an HTTP request.
[1809] Step 4:
[1810] The server regenerates the plan to reflect the changed information. The input is the change instruction information, and the output is the regenerated plan. The server recalculates the optimal route and detailed schedule and generates a new plan.
[1811] Step 5:
[1812] The server sends the updated plan to the user's terminal. The input is the regenerated plan, and the output is the updated plan displayed on the terminal. The server sends the updated plan to the terminal via an HTTP response, which the terminal receives and displays.
[1813] Use of long-term emotional data
[1814] Step 1:
[1815] The server records the user's emotional state over the long term. The input is recognized emotional state information, and the output is emotional data stored in a database. The server performs insert operations into the database to save the emotional data.
[1816] Step 2:
[1817] The server optimizes the next travel suggestion based on accumulated historical data. The input is long-term sentiment data, and the output is an optimized travel suggestion. The server analyzes the sentiment data and proposes the best travel plan based on the user's priorities and preferences.
[1818] The above outlines the specific processing steps of the system.
[1819] (Application Example 2)
[1820] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1821] Traditional travel planning systems have been somewhat effective in creating travel plans that meet users' preferences, but they have a weakness in their inability to adjust plans to reflect real-time emotional states during the trip. Furthermore, means for users to smartly review their plans during their trip, and visually supportive devices, are still not fully utilized. Therefore, there is a need to reduce stress during travel and improve the user experience.
[1822] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1823] In this invention, the server includes means for inputting an image, analysis means for identifying a geographical location from the input image, plan generation means for generating a travel route based on the analyzed geographical location information and user instructions, presentation means for presenting the generated travel route and detailed schedule to the user, suggestion means for making additional suggestions based on the user's preferences, means including an emotion engine that analyzes the user's facial expressions and voice data and adjusts the travel plan based on their emotional state, and means for presenting the travel plan to the user using a real-time adaptable mobile or visual device. This makes it possible to adjust the plan in real time to reflect the user's emotional state during the trip and to improve the user's travel experience through visual support from a smart device.
[1824] "Image input means" refers to a device or software that receives image data input by a user and transmits it to a server.
[1825] "Analysis means" refers to a device or software that identifies a specific geographical location from input image data and provides that information to a server.
[1826] A "plan generation means" refers to a device or software that creates the optimal travel route and detailed schedule based on analyzed geographic location information and user instructions.
[1827] "Presentation means" refers to a device or software that provides the user with the generated travel route and detailed schedule visually or audibly.
[1828] "Suggestion means" refers to a device or software that provides additional information or recommended plans regarding travel routes based on the user's preferences.
[1829] "Means including an emotion engine" refers to a device or software that analyzes the user's facial expressions and voice data to recognize their emotional state and adjust the travel plan accordingly.
[1830] "Means for presenting a travel plan to a user using a mobile or visual device" refers to a device or software that visually presents a travel plan while reflecting the user's location information and emotional state in real time.
[1831] "Real-time" refers to a state where the user's current status and location information are reflected in real time, allowing for immediate adaptation or modification.
[1832] This invention is a system for efficiently creating travel plans and adjusting them in real time according to the user's emotional state. The system includes means for image input, analysis, plan generation, presentation, suggestion, an emotion engine, and means for presenting the travel plan to the user using a mobile device or visual device.
[1833] Description of the program's overall operation
[1834] 1. Image input and analysis
[1835] Users upload images of places they want to visit from their smartphones or PCs to the system in order to create travel plans. In this case, the image input mechanism is activated. The device sends these images to the server, which uses image recognition technology (e.g., a common image recognition API) to determine the geographical location of the images. The results of this analysis are stored in a database.
[1836] 2. Plan generation and user specification
[1837] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location. Based on this, the terminal sends the instructions to the server. The server generates the optimal travel route and detailed schedule based on the analyzed geographic location information and the user's instructions. This generation process uses a route calculation engine and transportation information.
[1838] 3. Collection and analysis of facial expressions and voice.
[1839] The device uses its built-in camera and microphone to collect user facial expressions and voice data. This data is sent to a server in real time, where the server uses an emotion engine to analyze the user's emotional state. Specifically, facial expressions are analyzed using facial recognition software (e.g., OpenCV or dlib), and the data is sent to an emotion analysis API.
[1840] 4. Adjusting plans based on emotions
[1841] Based on the analysis results from the emotion engine, the server adjusts the travel plan to match the user's emotional state. For example, if the user is tired, the plan will be changed to include relaxing facilities. Conversely, if the user is excited, the plan can be adjusted to include more active spots.
[1842] 5. Real-time plan presentation
[1843] The travel plan, generated in real time, is presented to the user through visual devices such as smart glasses. This allows the user to enjoy their trip while checking a plan that suits their emotional state in real time.
[1844] Specific example
[1845] The user uploads a landscape photo of the foot of Mt. Fuji (Photo A) and a photo of an urban landmark (Photo B) to the system.
[1846] The server identifies photo A as the foot of Mount Fuji and photo B as an urban area, and saves them in the database.
[1847] The user enters "I want to go to Kyoto via Photo A and Photo B," and "I want to stay for one day at Photo A and two days at Photo B."
[1848] The terminal sends instruction information to the server, which then generates the optimal route and a detailed schedule.
[1849] The device uses its built-in camera and microphone to collect the user's facial expressions and voice in real time as they review the plan, and transmits this information to the server.
[1850] The server uses an emotion engine to analyze the user's emotions and adjust the plan accordingly. For example, if it determines that the user is tired, it will change the plan to include more rest time.
[1851] Example of a prompt
[1852] "Since you seem tired this time, please suggest a travel plan that includes places where you can relax."
[1853] This system allows for real-time adjustments to travel plans that reflect the user's emotional state during their trip, and enhances the user's travel experience through visual support via smart devices.
[1854] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1855] Step 1:
[1856] The user uses their device to upload images of the place they want to go to the system.
[1857] The specific input is the user's image data, which the device sends to the server. The server uses image recognition technology to analyze the uploaded image. For example, it uses an image recognition API to determine the geographical location and stores this information in a database. The output is the analyzed geographical location information.
[1858] Step 2:
[1859] The user provides instructions via text input, such as "I want to go to XX via photo A and photo B," and specifies the length of stay at each location.
[1860] The input is user text instructions, which the terminal sends to the server. The server generates the optimal travel route and detailed schedule based on the starting point, analyzed geographic location information, and user instructions. It uses a route calculation engine to determine the best route, taking into account transportation methods and travel time. The output is the generated travel plan and detailed schedule.
[1861] Step 3:
[1862] The device uses its built-in camera and microphone to collect user facial expressions and voice data.
[1863] The input consists of the user's real-time facial expressions and voice data, which the terminal collects and sends to the server. The server uses an emotion engine to analyze the data and recognize the user's emotional state. Specifically, it analyzes facial expressions using facial recognition software (e.g., OpenCV or dlib) and sends the data to an emotion analysis API to identify the emotional state. The output is the user's emotional state as determined by the emotion engine.
[1864] Step 4:
[1865] The server adjusts the travel plan in real time according to the user's emotional state, based on the analysis results from the emotion engine.
[1866] The input consists of the user's emotional state, as determined by the emotion engine, and the existing travel plan. Based on this data, the server regenerates the plan, including new suggestions and modifications. For example, if the user is tired, the server might add rest areas and modify the plan to include relaxing facilities. The output is the adjusted travel plan.
[1867] Step 5:
[1868] The device presents the user with travel plans that have been generated or adjusted in real time.
[1869] The input is a pre-configured travel plan, and the device visually presents this information to the user via smart glasses or a smartphone. The user can continue their trip while checking the plan in real time to match their emotional state. The output is the pre-configured travel plan displayed on the user's device.
[1870] Step 6:
[1871] The server records the user's emotional state over the long term to optimize future travel suggestions.
[1872] The input consists of past travel plans and user emotional state data, which the server stores in a database. This data is used to generate optimal travel plans based on the user's preferences for future trips. The output is a more personalized suggestion for the next trip.
[1873] Through these steps, users can enjoy travel plans tailored to their emotional state in real time, enhancing their travel experience.
[1874] The specific processing unit 290 transmits the result of the specific processing to the robot 414...
Claims
1. Image as input method, An analysis means for determining geographical location from an input image, A plan generation means that generates a travel route based on analyzed geographic location information and user instructions, A presentation means for presenting the generated travel route and detailed schedule to the user, A suggestion method that provides additional suggestions based on user preferences, A system that includes this.
2. The system according to claim 1, which uses image recognition technology to identify a geographic location from an image.
3. The system according to claim 1, further comprising means for setting the duration of stay for the user at each location.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A