system
The system uses multimodal AI to analyze images and text to quickly and accurately identify locations, overcoming conventional challenges by providing precise location information.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional technologies face difficulties in quickly and accurately identifying specific locations from still images, moving images, or text describing conditions.
A system comprising a reception unit, analysis unit, and provision unit that utilizes multimodal AI to analyze still images, videos, or text containing conditions, and provides identified locations using map data.
Enables rapid and precise identification of locations from various input types, providing detailed information such as addresses, coordinates, and additional location-specific details.
Smart Images

Figure 2026073563000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that it is difficult to quickly and accurately identify a specific location from a still image, a moving image, or a text describing conditions.
[0005] The system according to the embodiment aims to quickly and accurately identify a specific location from a still image, a moving image, or a text describing conditions.
Means for Solving the Problems
[0006] The system according to the embodiment includes a reception unit, an analysis unit, and a provision unit. The reception unit receives an input of a still image, a moving image, or a text describing conditions from a user. The analysis unit analyzes the information received by the reception unit and identifies a specific location. The provision unit provides the location identified by the analysis unit. [Effects of the Invention]
[0007] The system according to this embodiment can quickly and accurately identify a specific location from still images, videos, or text describing conditions. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The location identification system according to an embodiment of the present invention is a system that uses map data to identify a location from the background in still images or videos, or from text describing conditions. The location identification system works by having the user input still images, videos, or text describing conditions. Next, a multimodal AI analyzes the input information to identify a specific location. Based on the analysis results, the identified location is provided to the user. This mechanism allows users to easily identify locations from still images, videos, or text describing conditions. For example, users can identify the location of a photograph taken during a trip or search for a location that meets specific conditions. First, the user inputs still images, videos, or text describing conditions. For example, the user might upload a still image with the caption "Where is this photo taken?" or input text describing conditions such as "A cafe with an ocean view." This information is input to the multimodal AI. Next, the multimodal AI analyzes the input information. The multimodal AI integrates and analyzes various information such as images, audio, and text to identify a specific location. For example, it identifies a location based on map data, using the background in a still image or text describing conditions. Based on the analysis results, the system provides users with identified locations. For example, if a location is identified in a still image uploaded by a user, detailed information about that location is provided. It also provides a list of identified locations based on text describing specific conditions. This mechanism allows users to easily identify locations from still images, videos, or text describing conditions. For example, users can identify the location of a photograph taken during a trip or search for a location that meets specific criteria. Thus, the location identification system allows users to easily identify locations from still images, videos, or text describing conditions.
[0029] The location identification system according to this embodiment comprises a reception unit, an analysis unit, and a provision unit. The reception unit receives input from the user in the form of still images, videos, or text describing conditions. The reception unit can accept, for example, still images in JPEG format, videos in MP4 format, and text in text file format. For example, the reception unit can receive still images uploaded by the user. The reception unit can also receive videos recorded by the user. The reception unit can also receive text describing conditions entered by the user. For example, the reception unit can receive a still image with the caption "Where is this photo taken?" or text describing conditions such as "A cafe with an ocean view." The analysis unit analyzes the information received by the reception unit to identify a specific location. The analysis unit analyzes the background in the still images and videos, and the text describing the conditions, for example, using multimodal AI. The analysis unit analyzes the background in the still images and videos using, for example, an image recognition algorithm. The analysis unit can also analyze the text describing the conditions using natural language processing technology. For example, the analysis unit analyzes the background in a still image and identifies the location based on map data. The analysis unit can also analyze text containing conditions and identify a specific location. The provision unit provides the location identified by the analysis unit. The provision unit provides the user with detailed information about the identified location. For example, the provision unit provides the address or geographical coordinates of the identified location. The provision unit can also provide the business hours or contact information of the identified location. For example, the provision unit provides the user with detailed information about the identified location. The provision unit can also provide a list of locations identified based on text containing conditions. For example, the provision unit provides a list of locations identified based on text containing conditions. As a result, the location identification system according to this embodiment allows users to easily identify locations from still images, videos, or text containing conditions.
[0030] The reception unit accepts still images, videos, or text containing conditions from users. For example, it can accept still images in JPEG format, videos in MP4 format, and text files. Specifically, it provides an interface for receiving still images and videos uploaded by users from smartphones or computers. For example, users can easily select and upload files through a web browser or dedicated application. Upon receiving these files, the reception unit immediately transfers them to the server for access by the analysis unit. The reception unit can also accept text containing conditions entered by users. For example, users can enter questions such as "Where is this photo taken?" or text describing conditions such as "A cafe with an ocean view." This allows the reception unit to efficiently receive diverse forms of information provided by users, improving the overall flexibility and usability of the system. Furthermore, the reception unit can appropriately classify the received information and perform preprocessing to enable efficient processing by the analysis unit. For example, adjusting the resolution of still images and videos, or standardizing the format of text data, can reduce the load on the analysis unit and improve processing speed. This allows the reception desk to smoothly process user input and play a role in improving the overall efficiency of the system.
[0031] The analysis unit analyzes information received by the reception unit to identify specific locations. For example, the analysis unit uses multimodal AI to analyze backgrounds in still images and videos, as well as text describing conditions. Specifically, it uses image recognition algorithms to analyze backgrounds in still images and videos, and natural language processing technology to analyze text describing conditions. For example, it identifies buildings and landmarks in still images and matches them with map data to determine the location. In the case of videos, it can analyze multiple frames and identify the location based on dynamic information. Furthermore, the analysis unit analyzes text describing conditions and extracts the characteristics of the location the user is looking for. For example, it analyzes the condition "cafe with an ocean view" and generates a list of cafes with ocean views. This involves a process of analyzing text using natural language processing technology to understand keywords and context. By combining these technologies, the analysis unit can efficiently analyze diverse information provided by users and accurately identify locations. Furthermore, the analysis unit can continuously improve analysis accuracy by utilizing past data and user feedback. For example, by referring to a database of previously identified locations and detecting similar patterns, the accuracy of the analysis can be improved. Furthermore, the analysis algorithm can be adjusted based on user feedback to provide more accurate results. This allows the analysis unit to quickly and accurately analyze the information provided by the user and identify specific locations.
[0032] The service provider provides locations identified by the analysis unit. For example, the service provider provides users with detailed information about the identified locations. Specifically, it provides the address and geographical coordinates of the identified locations. The service provider can also provide information such as business hours and contact details for the identified locations. For example, if the identified location is a restaurant, it can provide its business hours, contact information, and menu information. Furthermore, the service provider can provide a list of identified locations based on a written description of the conditions. For example, based on the condition "cafes with ocean views," it can generate a list of multiple cafes and provide detailed information for each. To provide this information to users in a visually easy-to-understand format, the service provider can provide interfaces such as map displays and list displays. For example, identified locations can be displayed as pins on a map, allowing users to click and view detailed information. In a list display, the names, addresses, and ratings of identified locations can be displayed in a list format to facilitate comparison by the user. Furthermore, the service provider can collect user feedback and continuously improve the accuracy and quality of the information provided. For example, users can leave ratings and comments on the provided information, and the information can be updated and improved based on that feedback. This allows the service provider to offer users accurate and useful information, improving the overall reliability and usability of the system.
[0033] The analysis unit includes an image analysis unit that can analyze the background in still images and videos. The analysis unit can, for example, use an image recognition algorithm to analyze the background in still images and videos. For example, the analysis unit can analyze backgrounds such as landscapes, buildings, and natural environments. The analysis unit can also use an image recognition algorithm to analyze the background in still images and videos. For example, the analysis unit can analyze backgrounds such as landscapes, buildings, and natural environments. This makes it possible to identify locations by analyzing the background in still images and videos. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can take the background in a still image or video as input and analyze the background using an AI model that analyzes backgrounds.
[0034] The analysis unit includes a text analysis unit and can analyze text containing conditions. The analysis unit can analyze text containing conditions using, for example, natural language processing technology. For example, the analysis unit can extract keywords from text containing conditions using keyword extraction technology. The analysis unit can also analyze the context of text containing conditions using context analysis technology. For example, the analysis unit extracts keywords from text containing conditions and analyzes the context. This makes it possible to identify locations by analyzing text containing conditions. Some or all of the above processing in the analysis unit may be performed using, for example, AI, or without AI. For example, the analysis unit can take text containing conditions as input and analyze the text using an AI model that analyzes text.
[0035] The service provider can provide users with detailed information about a specified location. For example, the service provider can provide the address or geographical coordinates of the specified location. The service provider can also provide the geographical coordinates of the specified location. The service provider can also provide the business hours and contact information of the specified location. The service provider can also provide the business hours and contact information of the specified location. The service provider can also provide the contact information of the specified location. By providing detailed information about the specified location, users can understand the details of the location. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take detailed information about a specified location as input and provide detailed information using an AI model that provides detailed information.
[0036] The service provider can provide a list of locations identified based on a text describing the conditions. For example, the service provider can provide a list of locations identified based on a text describing the conditions. For example, the service provider can provide a list of locations identified based on a text describing the conditions. The service provider can also provide the list of identified locations in the form of a text list, a link list, an image list, etc. For example, the service provider can provide the list of identified locations in the form of a text list. The service provider can also provide the list of identified locations in the form of a link list. For example, the service provider can provide the list of identified locations in the form of an image list. For example, the service provider can provide the list of identified locations in the form of an image list. This allows the user to choose a location from multiple options by providing a list of locations identified based on a text describing the conditions. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can take a list of locations identified based on a text describing the conditions as input and provide a list using an AI model that provides lists.
[0037] The reception desk can analyze the user's past input history and select the optimal input method. For example, based on the user's past input history, the reception desk can prioritize suggesting frequently used input methods. For example, the reception desk can prioritize suggesting voice input or text input, which the user has frequently used in the past. The reception desk can also predict and suggest input methods to be used at specific times based on the user's past input history. For example, if the reception desk tends to use voice input at a particular time, it will suggest voice input at that time. The reception desk can also suggest similar input methods by referring to content the user has entered in the past. For example, the reception desk will suggest similar input methods based on content the user has entered in the past. In this way, by analyzing the user's past input history, the optimal input method can be suggested. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can select an input method using an AI model that takes the user's past input history as input and selects the optimal input method.
[0038] The input system can filter input based on the user's current areas of interest. For example, the input system can identify the user's current areas of interest based on their past search history or survey results. The input system can also prioritize input that is highly relevant to the user's current areas of interest. For example, if the user is interested in travel, the input system will prioritize displaying travel-related input options. Similarly, if the user is interested in food, the input system can prioritize displaying input options related to restaurants and cafes. For example, if the user is interested in shopping, the input system can prioritize displaying shopping-related input options. This allows the system to prioritize input that is highly relevant by filtering based on the user's current areas of interest. Some or all of the above-described processing in the reception area may be performed using AI, for example, or without AI. For example, the reception area can take the user's current areas of interest as input and use an AI model to identify areas of interest.
[0039] The reception system can prioritize receiving inputs that are highly relevant based on the user's geographical location information. For example, the reception system can obtain the user's geographical location information using GPS data or an IP address. For instance, it can obtain geographical location information using the user's smartphone's GPS data. Alternatively, the reception system can obtain geographical location information using the user's IP address. Furthermore, the reception system can prioritize receiving inputs that are highly relevant based on the user's geographical location information. For example, it might prioritize inputs related to locations close to the user's current location. Also, if the user is interested in a particular region, the reception system can prioritize inputs related to that region. Furthermore, if the user is traveling, the reception system can prioritize inputs related to their travel destination. For example, if the user is traveling, the reception system will prioritize inputs related to their travel destination. This allows for the provision of more appropriate information by prioritizing the acceptance of highly relevant inputs based on the user's geographical location information. Some or all of the processing described above in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can accept input using an AI model that takes the user's geographical location information as input and prioritizes the acceptance of highly relevant inputs.
[0040] The reception unit can analyze the user's social media activity and accept relevant inputs when receiving input. For example, the reception unit can analyze the content of the user's social media posts and the number of followers. For example, the reception unit can analyze photos and videos that the user has shared on social media. The reception unit can also analyze places and events that the user follows on social media. For example, the reception unit can analyze places and events that the user follows on social media. The reception unit can also analyze topics that the user has shown interest in on social media. For example, the reception unit can analyze topics that the user has shown interest in on social media. By analyzing the user's social media activity, the reception unit can prioritize the acceptance of relevant inputs. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can accept input using an AI model that takes the user's social media activity as input and accepts relevant inputs.
[0041] The analysis unit can optimize the analysis algorithm by referring to past analysis data during image analysis. For example, the analysis unit can refer to past image data and analysis results. For example, the analysis unit can apply the optimal algorithm to similar images based on previously analyzed image data. The analysis unit can also refer to past analysis results and adjust parameters to improve analysis accuracy. For example, the analysis unit can adjust parameters to improve analysis accuracy based on past analysis results. The analysis unit can also detect specific patterns based on past analysis data and optimize the analysis algorithm. For example, the analysis unit can detect specific patterns based on past analysis data and optimize the analysis algorithm. In this way, by referring to past analysis data, the analysis algorithm can be optimized and analysis accuracy can be improved. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can take past analysis data as input and optimize the algorithm using an AI model that optimizes the analysis algorithm.
[0042] The analysis unit can apply different analysis methods depending on the image category during image analysis. For example, the analysis unit can classify images into categories such as landscapes, people, and buildings. The analysis unit can also apply different analysis methods depending on the image category. For example, the analysis unit can apply a landscape analysis algorithm to images of natural landscapes. For example, the analysis unit can apply a building analysis algorithm to images of urban landscapes. For example, the analysis unit can apply an interior analysis algorithm to images of interiors. By applying different analysis methods depending on the image category, the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the image category as input and applies an analysis method appropriate to the category.
[0043] The analysis unit can adjust the level of detail of its analysis based on the importance of the text. For example, the analysis unit can evaluate the importance of a text based on the frequency of keyword occurrences or the importance of context. The analysis unit can also evaluate the importance of a text based on the importance of context. The analysis unit can also adjust the level of detail of its analysis based on the importance of the text. For example, the analysis unit performs a detailed analysis on texts with high importance. The analysis unit can also perform a simplified analysis on texts with low importance. By adjusting the level of detail of the analysis based on the importance of the text, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the importance of text as input and adjusts the level of detail of the analysis based on the importance.
[0044] The analysis unit can apply different analysis algorithms depending on the category of the text during text analysis. For example, the analysis unit classifies texts into categories such as technical documents, news articles, and blog posts. The analysis unit can also apply different analysis algorithms depending on the category of the text. For example, the analysis unit applies a technical document analysis algorithm to technical documents. The analysis unit can also apply a news article analysis algorithm to news articles. For example, the analysis unit applies a news article analysis algorithm to news articles. The analysis unit can also apply a blog post analysis algorithm to blog posts. For example, the analysis unit applies a blog post analysis algorithm to blog posts. By applying different analysis algorithms depending on the category of the text, the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the category of the text as input and applies an analysis algorithm appropriate to the category.
[0045] The service provider can adjust the level of detail provided based on the importance of the specified location at the time of delivery. For example, the service provider can evaluate the importance of the specified location based on the user's level of interest or the urgency of the information. The service provider can also evaluate the importance of the specified location based on the urgency of the information. The service provider can also evaluate the importance of the specified location based on the urgency of the information. The service provider can also adjust the level of detail provided based on the importance of the specified location. For example, if it is an important tourist destination, the service provider can provide detailed information. The service provider can also provide concise information if it is a general location. The service provider can also adjust the level of detail provided based on the importance specified by the user. By adjusting the level of detail provided based on the importance of the specified location, more appropriate information can be provided. Some or all of the processing described above in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can provide information using an AI model that takes the importance of a specified location as input and adjusts the level of detail of the delivery based on the importance.
[0046] The service provider can apply different service delivery methods depending on the category of the specified location. For example, the service provider can classify the categories of the specified location into tourist destinations, restaurants, shopping malls, etc. The service provider can also apply different service delivery methods depending on the category of the specified location. For example, in the case of a tourist destination, the service provider can provide information in the form of a tourist guide. For example, in the case of a restaurant, the service provider can provide information including menus and reviews. For example, in the case of a shopping mall, the service provider can provide store information and sales information. By applying different service delivery methods depending on the category of the specified location, more appropriate information can be provided. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide information using an AI model that takes the category of the specified location as input and applies a service delivery method appropriate to the category.
[0047] The service provider can adjust the order of information provided based on the relevance of the identified locations. For example, the service provider can evaluate the relevance of identified locations based on the user's search history or current location information. The service provider can also evaluate the relevance of identified locations based on the user's current location information. The service provider can also adjust the order of information provided based on the relevance of the identified locations. For example, the service provider can prioritize providing locations that are most relevant to the conditions specified by the user. The service provider can also prioritize providing locations that are more relevant based on the user's past input history. The service provider can also prioritize providing locations that are more relevant based on the user's current areas of interest. By adjusting the order of information provided based on the relevance of the identified locations, more appropriate information can be provided. Some or all of the processing described above in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can provide information using an AI model that takes the relationships of identified locations as input and adjusts the order of delivery based on those relationships.
[0048] The service provider can improve the accuracy of its service by referring to relevant literature for the specified location at the time of service provision. For example, the service provider can refer to tourist guides and reviews related to the specified location. The service provider can also refer to reviews related to the specified location. The service provider can also refer to historical background and cultural information related to the specified location. For example, the service provider can refer to historical background related to the specified location. The service provider can also refer to cultural information related to the specified location. The service provider can also refer to the latest news and event information related to the specified location. For example, the service provider can refer to the latest news related to the specified location. The service provider can also refer to event information related to the specified location. This allows for improved accuracy of the service by referring to relevant literature for the specified location. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes relevant literature for a specified location as input and improves the accuracy of the information provision by referring to the relevant literature.
[0049] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0050] The reception desk can automatically suggest relevant past search history based on the user's input. For example, if a user enters "cafe with an ocean view," the reception desk will display search history such as "seaside restaurant" or "beach cafe" that the user has searched for in the past. Also, if a user uploads a still image and asks "Where is this photo taken?", the reception desk can suggest similar photos that the user has uploaded in the past. This allows users to identify locations more efficiently by referring to their past search history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can take the user's past search history as input and suggest history using an AI model that suggests relevant history.
[0051] The reception desk can automatically suggest relevant news and event information based on user input. For example, if a user enters "cafe with an ocean view," the reception desk will display events and the latest news in that area. Similarly, if a user uploads a still image with the question "Where is this photo taken?", the reception desk can suggest the latest news and event information related to that location. This allows users to identify locations more efficiently while referring to relevant and up-to-date information. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk could use an AI model that takes user input as input and suggests relevant news and event information to propose information.
[0052] The analysis unit can optimize its analysis algorithm by referring to the user's past analysis results during image analysis. For example, it can apply the optimal algorithm to similar images based on the analysis results of images previously uploaded by the user. It can also apply the optimal algorithm to similar conditions based on the analysis results of text containing conditions previously entered by the user. In this way, by referring to the user's past analysis results, the analysis algorithm can be optimized and the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can optimize its algorithm using an AI model that optimizes the analysis algorithm based on the user's past analysis results as input.
[0053] The information provider can prioritize providing highly relevant information based on the user's current geographical location when providing detailed information about a specified location. For example, it can prioritize displaying information about locations close to the user's current location, and if the user is interested in a particular region, it can prioritize displaying information related to that region. Furthermore, if the user is traveling, it can prioritize providing information related to their travel destination. This allows for the provision of more appropriate information by providing highly relevant information based on the user's current geographical location. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes the user's geographical location as input and provides highly relevant information.
[0054] The reception desk can automatically suggest relevant social media posts based on the user's input. For example, if a user enters "cafe with an ocean view," the reception desk will display popular social media posts in that area. Similarly, if a user uploads a still image with the question "Where is this photo taken?", the reception desk can suggest social media posts related to that location. This allows users to identify locations more efficiently by referring to relevant social media posts. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk could suggest posts using an AI model that takes the user's input as input and suggests relevant social media posts.
[0055] The information provider can prioritize providing highly relevant information based on the user's past search history when providing detailed information about a specified location. For example, it can prioritize displaying information about relevant locations based on the user's past search history, such as "seaside restaurants" or "beach cafes." It can also prioritize providing information about relevant locations based on the user's past search history, including the text containing the conditions they have entered. This allows for the provision of more appropriate information by providing highly relevant information based on the user's past search history. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes the user's past search history as input and provides highly relevant information.
[0056] The following briefly describes the processing flow for example form 1.
[0057] Step 1: The reception desk accepts still images, videos, or text files containing conditions from users. For example, it can accept still images in JPEG format, videos in MP4 format, and text files. It accepts still images, recorded videos, or text files containing conditions uploaded by users. Step 2: The analysis unit analyzes the information received by the reception unit to identify a specific location. For example, it uses multimodal AI to analyze the background in still images and videos, as well as the text describing the conditions. It uses image recognition algorithms and natural language processing technology to analyze the background in still images and videos, as well as the text describing the conditions, and identifies the location based on map data. Step 3: The provisioning unit provides the locations identified by the analysis unit. For example, it provides the user with detailed information about the identified locations, such as addresses, geographical coordinates, business hours, and contact information. It can also provide a list of identified locations based on a written description of the conditions.
[0058] (Example of form 2) The location identification system according to an embodiment of the present invention is a system that uses map data to identify a location from the background in still images or videos, or from text describing conditions. The location identification system works by having the user input still images, videos, or text describing conditions. Next, a multimodal AI analyzes the input information to identify a specific location. Based on the analysis results, the identified location is provided to the user. This mechanism allows users to easily identify locations from still images, videos, or text describing conditions. For example, users can identify the location of a photograph taken during a trip or search for a location that meets specific conditions. First, the user inputs still images, videos, or text describing conditions. For example, the user might upload a still image with the caption "Where is this photo taken?" or input text describing conditions such as "A cafe with an ocean view." This information is input to the multimodal AI. Next, the multimodal AI analyzes the input information. The multimodal AI integrates and analyzes various information such as images, audio, and text to identify a specific location. For example, it identifies a location based on map data, using the background in a still image or text describing conditions. Based on the analysis results, the system provides users with identified locations. For example, if a location is identified in a still image uploaded by a user, detailed information about that location is provided. It also provides a list of identified locations based on text describing specific conditions. This mechanism allows users to easily identify locations from still images, videos, or text describing conditions. For example, users can identify the location of a photograph taken during a trip or search for a location that meets specific criteria. Thus, the location identification system allows users to easily identify locations from still images, videos, or text describing conditions.
[0059] The location identification system according to this embodiment comprises a reception unit, an analysis unit, and a provision unit. The reception unit receives input from the user in the form of still images, videos, or text describing conditions. The reception unit can accept, for example, still images in JPEG format, videos in MP4 format, and text in text file format. For example, the reception unit can receive still images uploaded by the user. The reception unit can also receive videos recorded by the user. The reception unit can also receive text describing conditions entered by the user. For example, the reception unit can receive a still image with the caption "Where is this photo taken?" or text describing conditions such as "A cafe with an ocean view." The analysis unit analyzes the information received by the reception unit to identify a specific location. The analysis unit analyzes the background in the still images and videos, and the text describing the conditions, for example, using multimodal AI. The analysis unit analyzes the background in the still images and videos using, for example, an image recognition algorithm. The analysis unit can also analyze the text describing the conditions using natural language processing technology. For example, the analysis unit analyzes the background in a still image and identifies the location based on map data. The analysis unit can also analyze text containing conditions and identify a specific location. The provision unit provides the location identified by the analysis unit. The provision unit provides the user with detailed information about the identified location. For example, the provision unit provides the address or geographical coordinates of the identified location. The provision unit can also provide the business hours or contact information of the identified location. For example, the provision unit provides the user with detailed information about the identified location. The provision unit can also provide a list of locations identified based on text containing conditions. For example, the provision unit provides a list of locations identified based on text containing conditions. As a result, the location identification system according to this embodiment allows users to easily identify locations from still images, videos, or text containing conditions.
[0060] The reception unit accepts still images, videos, or text containing conditions from users. For example, it can accept still images in JPEG format, videos in MP4 format, and text files. Specifically, it provides an interface for receiving still images and videos uploaded by users from smartphones or computers. For example, users can easily select and upload files through a web browser or dedicated application. Upon receiving these files, the reception unit immediately transfers them to the server for access by the analysis unit. The reception unit can also accept text containing conditions entered by users. For example, users can enter questions such as "Where is this photo taken?" or text describing conditions such as "A cafe with an ocean view." This allows the reception unit to efficiently receive diverse forms of information provided by users, improving the overall flexibility and usability of the system. Furthermore, the reception unit can appropriately classify the received information and perform preprocessing to enable efficient processing by the analysis unit. For example, adjusting the resolution of still images and videos, or standardizing the format of text data, can reduce the load on the analysis unit and improve processing speed. This allows the reception desk to smoothly process user input and play a role in improving the overall efficiency of the system.
[0061] The analysis unit analyzes information received by the reception unit to identify specific locations. For example, the analysis unit uses multimodal AI to analyze backgrounds in still images and videos, as well as text describing conditions. Specifically, it uses image recognition algorithms to analyze backgrounds in still images and videos, and natural language processing technology to analyze text describing conditions. For example, it identifies buildings and landmarks in still images and matches them with map data to determine the location. In the case of videos, it can analyze multiple frames and identify the location based on dynamic information. Furthermore, the analysis unit analyzes text describing conditions and extracts the characteristics of the location the user is looking for. For example, it analyzes the condition "cafe with an ocean view" and generates a list of cafes with ocean views. This involves a process of analyzing text using natural language processing technology to understand keywords and context. By combining these technologies, the analysis unit can efficiently analyze diverse information provided by users and accurately identify locations. Furthermore, the analysis unit can continuously improve analysis accuracy by utilizing past data and user feedback. For example, by referring to a database of previously identified locations and detecting similar patterns, the accuracy of the analysis can be improved. Furthermore, the analysis algorithm can be adjusted based on user feedback to provide more accurate results. This allows the analysis unit to quickly and accurately analyze the information provided by the user and identify specific locations.
[0062] The service provider provides locations identified by the analysis unit. For example, the service provider provides users with detailed information about the identified locations. Specifically, it provides the address and geographical coordinates of the identified locations. The service provider can also provide information such as business hours and contact details for the identified locations. For example, if the identified location is a restaurant, it can provide its business hours, contact information, and menu information. Furthermore, the service provider can provide a list of identified locations based on a written description of the conditions. For example, based on the condition "cafes with ocean views," it can generate a list of multiple cafes and provide detailed information for each. To provide this information to users in a visually easy-to-understand format, the service provider can provide interfaces such as map displays and list displays. For example, identified locations can be displayed as pins on a map, allowing users to click and view detailed information. In a list display, the names, addresses, and ratings of identified locations can be displayed in a list format to facilitate comparison by the user. Furthermore, the service provider can collect user feedback and continuously improve the accuracy and quality of the information provided. For example, users can leave ratings and comments on the provided information, and the information can be updated and improved based on that feedback. This allows the service provider to offer users accurate and useful information, improving the overall reliability and usability of the system.
[0063] The analysis unit includes an image analysis unit that can analyze the background in still images and videos. The analysis unit can, for example, use an image recognition algorithm to analyze the background in still images and videos. For example, the analysis unit can analyze backgrounds such as landscapes, buildings, and natural environments. The analysis unit can also use an image recognition algorithm to analyze the background in still images and videos. For example, the analysis unit can analyze backgrounds such as landscapes, buildings, and natural environments. This makes it possible to identify locations by analyzing the background in still images and videos. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can take the background in a still image or video as input and analyze the background using an AI model that analyzes backgrounds.
[0064] The analysis unit includes a text analysis unit and can analyze text containing conditions. The analysis unit can analyze text containing conditions using, for example, natural language processing technology. For example, the analysis unit can extract keywords from text containing conditions using keyword extraction technology. The analysis unit can also analyze the context of text containing conditions using context analysis technology. For example, the analysis unit extracts keywords from text containing conditions and analyzes the context. This makes it possible to identify locations by analyzing text containing conditions. Some or all of the above processing in the analysis unit may be performed using, for example, AI, or without AI. For example, the analysis unit can take text containing conditions as input and analyze the text using an AI model that analyzes text.
[0065] The service provider can provide users with detailed information about a specified location. For example, the service provider can provide the address or geographical coordinates of the specified location. The service provider can also provide the geographical coordinates of the specified location. The service provider can also provide the business hours and contact information of the specified location. The service provider can also provide the business hours and contact information of the specified location. The service provider can also provide the contact information of the specified location. By providing detailed information about the specified location, users can understand the details of the location. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take detailed information about a specified location as input and provide detailed information using an AI model that provides detailed information.
[0066] The service provider can provide a list of locations identified based on a text describing the conditions. For example, the service provider can provide a list of locations identified based on a text describing the conditions. For example, the service provider can provide a list of locations identified based on a text describing the conditions. The service provider can also provide the list of identified locations in the form of a text list, a link list, an image list, etc. For example, the service provider can provide the list of identified locations in the form of a text list. The service provider can also provide the list of identified locations in the form of a link list. For example, the service provider can provide the list of identified locations in the form of an image list. For example, the service provider can provide the list of identified locations in the form of an image list. This allows the user to choose a location from multiple options by providing a list of locations identified based on a text describing the conditions. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can take a list of locations identified based on a text describing the conditions as input and provide a list using an AI model that provides lists.
[0067] The reception unit can estimate the user's emotions and adjust the timing of input acceptance based on the estimated emotions. For example, the reception unit can estimate the user's emotions using facial recognition technology. For example, the reception unit can capture the user's facial expression with a camera and estimate the emotions using a facial recognition algorithm. The reception unit can also estimate the user's emotions using voice analysis technology. For example, the reception unit can record the user's voice and estimate the emotions using a voice analysis algorithm. The reception unit can also estimate the user's emotions using text analysis technology. For example, the reception unit can analyze the text entered by the user and estimate the emotions using a text analysis algorithm. This allows for input acceptance at a more appropriate time by adjusting the timing of input acceptance according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception desk can take user emotion data as input and use an AI model to estimate emotions.
[0068] The reception desk can analyze the user's past input history and select the optimal input method. For example, based on the user's past input history, the reception desk can prioritize suggesting frequently used input methods. For example, the reception desk can prioritize suggesting voice input or text input, which the user has frequently used in the past. The reception desk can also predict and suggest input methods to be used at specific times based on the user's past input history. For example, if the reception desk tends to use voice input at a particular time, it will suggest voice input at that time. The reception desk can also suggest similar input methods by referring to content the user has entered in the past. For example, the reception desk will suggest similar input methods based on content the user has entered in the past. In this way, by analyzing the user's past input history, the optimal input method can be suggested. Some or all of the above processing in the reception desk may be performed using AI, for example, or without AI. For example, the reception desk can select an input method using an AI model that takes the user's past input history as input and selects the optimal input method.
[0069] The input system can filter input based on the user's current areas of interest. For example, the input system can identify the user's current areas of interest based on their past search history or survey results. The input system can also prioritize input that is highly relevant to the user's current areas of interest. For example, if the user is interested in travel, the input system will prioritize displaying travel-related input options. Similarly, if the user is interested in food, the input system can prioritize displaying input options related to restaurants and cafes. For example, if the user is interested in shopping, the input system can prioritize displaying shopping-related input options. This allows the system to prioritize input that is highly relevant by filtering based on the user's current areas of interest. Some or all of the above-described processing in the reception area may be performed using AI, for example, or without AI. For example, the reception area can take the user's current areas of interest as input and use an AI model to identify areas of interest.
[0070] The reception unit can estimate the user's emotions and determine the priority of input reception based on the estimated user emotions. The reception unit can estimate the user's emotions using, for example, facial recognition technology. For example, the reception unit can capture the user's facial expression with a camera and estimate the emotions using a facial recognition algorithm. The reception unit can also estimate the user's emotions using voice analysis technology. For example, the reception unit can record the user's voice and estimate the emotions using a voice analysis algorithm. The reception unit can also estimate the user's emotions using text analysis technology. For example, the reception unit can analyze the text entered by the user and estimate the emotions using a text analysis algorithm. This allows for input reception in a more appropriate order by determining the priority of input reception according to the user's emotions. Emotion estimation is implemented using, for example, an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using, for example, AI, or not using AI. For example, the reception desk can take user emotion data as input and use an AI model to estimate emotions.
[0071] The reception system can prioritize receiving inputs that are highly relevant based on the user's geographical location information. For example, the reception system can obtain the user's geographical location information using GPS data or an IP address. For instance, it can obtain geographical location information using the user's smartphone's GPS data. Alternatively, the reception system can obtain geographical location information using the user's IP address. Furthermore, the reception system can prioritize receiving inputs that are highly relevant based on the user's geographical location information. For example, it might prioritize inputs related to locations close to the user's current location. Also, if the user is interested in a particular region, the reception system can prioritize inputs related to that region. Furthermore, if the user is traveling, the reception system can prioritize inputs related to their travel destination. For example, if the user is traveling, the reception system will prioritize inputs related to their travel destination. This allows for the provision of more appropriate information by prioritizing the acceptance of highly relevant inputs based on the user's geographical location information. Some or all of the processing described above in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can accept input using an AI model that takes the user's geographical location information as input and prioritizes the acceptance of highly relevant inputs.
[0072] The reception unit can analyze the user's social media activity and accept relevant inputs when receiving input. For example, the reception unit can analyze the content of the user's social media posts and the number of followers. For example, the reception unit can analyze photos and videos that the user has shared on social media. The reception unit can also analyze places and events that the user follows on social media. For example, the reception unit can analyze places and events that the user follows on social media. The reception unit can also analyze topics that the user has shown interest in on social media. For example, the reception unit can analyze topics that the user has shown interest in on social media. By analyzing the user's social media activity, the reception unit can prioritize the acceptance of relevant inputs. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can accept input using an AI model that takes the user's social media activity as input and accepts relevant inputs.
[0073] The analysis unit can estimate the user's emotions and adjust the accuracy of the image analysis based on the estimated emotions. For example, the analysis unit can estimate the user's emotions using facial recognition technology. For example, the analysis unit can capture the user's facial expression with a camera and estimate the emotions using a facial recognition algorithm. The analysis unit can also estimate the user's emotions using speech analysis technology. For example, the analysis unit can record the user's voice and estimate the emotions using a speech analysis algorithm. The analysis unit can also estimate the user's emotions using text analysis technology. For example, the analysis unit can analyze text entered by the user and estimate the emotions using a text analysis algorithm. By adjusting the accuracy of the image analysis according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without using AI. For example, the analysis unit can take user emotion data as input and estimate emotions using an AI model that estimates emotions.
[0074] The analysis unit can optimize the analysis algorithm by referring to past analysis data during image analysis. For example, the analysis unit can refer to past image data and analysis results. For example, the analysis unit can apply the optimal algorithm to similar images based on previously analyzed image data. The analysis unit can also refer to past analysis results and adjust parameters to improve analysis accuracy. For example, the analysis unit can adjust parameters to improve analysis accuracy based on past analysis results. The analysis unit can also detect specific patterns based on past analysis data and optimize the analysis algorithm. For example, the analysis unit can detect specific patterns based on past analysis data and optimize the analysis algorithm. In this way, by referring to past analysis data, the analysis algorithm can be optimized and analysis accuracy can be improved. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can take past analysis data as input and optimize the algorithm using an AI model that optimizes the analysis algorithm.
[0075] The analysis unit can apply different analysis methods depending on the image category during image analysis. For example, the analysis unit can classify images into categories such as landscapes, people, and buildings. The analysis unit can also apply different analysis methods depending on the image category. For example, the analysis unit can apply a landscape analysis algorithm to images of natural landscapes. For example, the analysis unit can apply a building analysis algorithm to images of urban landscapes. For example, the analysis unit can apply an interior analysis algorithm to images of interiors. By applying different analysis methods depending on the image category, the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the image category as input and applies an analysis method appropriate to the category.
[0076] The analysis unit can estimate the user's emotions and adjust the text analysis representation based on the estimated user emotions. For example, the analysis unit can estimate the user's emotions using facial recognition technology. For example, the analysis unit can capture the user's facial expressions with a camera and estimate the emotions using a facial recognition algorithm. The analysis unit can also estimate the user's emotions using speech analysis technology. For example, the analysis unit can record the user's voice and estimate the emotions using a speech analysis algorithm. The analysis unit can also estimate the user's emotions using text analysis technology. For example, the analysis unit can analyze the text entered by the user and estimate the emotions using a text analysis algorithm. By adjusting the text analysis representation according to the user's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without using AI. For example, the analysis unit can take user emotion data as input and estimate emotions using an AI model that estimates emotions.
[0077] The analysis unit can adjust the level of detail of its analysis based on the importance of the text. For example, the analysis unit can evaluate the importance of a text based on the frequency of keyword occurrences or the importance of context. The analysis unit can also evaluate the importance of a text based on the importance of context. The analysis unit can also adjust the level of detail of its analysis based on the importance of the text. For example, the analysis unit performs a detailed analysis on texts with high importance. The analysis unit can also perform a simplified analysis on texts with low importance. By adjusting the level of detail of the analysis based on the importance of the text, more appropriate analysis results can be provided. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the importance of text as input and adjusts the level of detail of the analysis based on the importance.
[0078] The analysis unit can apply different analysis algorithms depending on the category of the text during text analysis. For example, the analysis unit classifies texts into categories such as technical documents, news articles, and blog posts. The analysis unit can also apply different analysis algorithms depending on the category of the text. For example, the analysis unit applies a technical document analysis algorithm to technical documents. The analysis unit can also apply a news article analysis algorithm to news articles. For example, the analysis unit applies a news article analysis algorithm to news articles. The analysis unit can also apply a blog post analysis algorithm to blog posts. For example, the analysis unit applies a blog post analysis algorithm to blog posts. By applying different analysis algorithms depending on the category of the text, the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the category of the text as input and applies an analysis algorithm appropriate to the category.
[0079] The service provider can estimate the user's emotions and adjust the way the information is presented based on the estimated emotions. For example, the service provider can estimate the user's emotions using facial recognition technology. For example, the service provider can capture the user's facial expression with a camera and estimate the emotions using a facial recognition algorithm. The service provider can also estimate the user's emotions using speech analysis technology. For example, the service provider can record the user's voice and estimate the emotions using a speech analysis algorithm. The service provider can also estimate the user's emotions using text analysis technology. For example, the service provider can analyze the text entered by the user and estimate the emotions using a text analysis algorithm. This allows the service provider to provide more appropriate information by adjusting the way the information is presented according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take user emotion data as input and use an AI model to estimate emotions.
[0080] The service provider can adjust the level of detail provided based on the importance of the specified location at the time of delivery. For example, the service provider can evaluate the importance of the specified location based on the user's level of interest or the urgency of the information. The service provider can also evaluate the importance of the specified location based on the urgency of the information. The service provider can also evaluate the importance of the specified location based on the urgency of the information. The service provider can also adjust the level of detail provided based on the importance of the specified location. For example, if it is an important tourist destination, the service provider can provide detailed information. The service provider can also provide concise information if it is a general location. The service provider can also adjust the level of detail provided based on the importance specified by the user. By adjusting the level of detail provided based on the importance of the specified location, more appropriate information can be provided. Some or all of the processing described above in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can provide information using an AI model that takes the importance of a specified location as input and adjusts the level of detail of the delivery based on the importance.
[0081] The service provider can apply different service delivery methods depending on the category of the specified location. For example, the service provider can classify the categories of the specified location into tourist destinations, restaurants, shopping malls, etc. The service provider can also apply different service delivery methods depending on the category of the specified location. For example, in the case of a tourist destination, the service provider can provide information in the form of a tourist guide. For example, in the case of a restaurant, the service provider can provide information including menus and reviews. For example, in the case of a shopping mall, the service provider can provide store information and sales information. By applying different service delivery methods depending on the category of the specified location, more appropriate information can be provided. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can provide information using an AI model that takes the category of the specified location as input and applies a service delivery method appropriate to the category.
[0082] The service provider can estimate the user's emotions and determine the priority of the information to be provided based on the estimated emotions. For example, the service provider can estimate the user's emotions using facial recognition technology. For example, the service provider can capture the user's facial expression with a camera and estimate the emotions using a facial recognition algorithm. The service provider can also estimate the user's emotions using voice analysis technology. For example, the service provider can record the user's voice and estimate the emotions using a voice analysis algorithm. The service provider can also estimate the user's emotions using text analysis technology. For example, the service provider can analyze the text entered by the user and estimate the emotions using a text analysis algorithm. This allows the service provider to provide more appropriate information by determining the priority of the information to be provided according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can take user emotion data as input and use an AI model to estimate emotions.
[0083] The service provider can adjust the order of information provided based on the relevance of the identified locations. For example, the service provider can evaluate the relevance of identified locations based on the user's search history or current location information. The service provider can also evaluate the relevance of identified locations based on the user's current location information. The service provider can also adjust the order of information provided based on the relevance of the identified locations. For example, the service provider can prioritize providing locations that are most relevant to the conditions specified by the user. The service provider can also prioritize providing locations that are more relevant based on the user's past input history. The service provider can also prioritize providing locations that are more relevant based on the user's current areas of interest. By adjusting the order of information provided based on the relevance of the identified locations, more appropriate information can be provided. Some or all of the processing described above in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can provide information using an AI model that takes the relationships of identified locations as input and adjusts the order of delivery based on those relationships.
[0084] The service provider can improve the accuracy of its service by referring to relevant literature for the specified location at the time of service provision. For example, the service provider can refer to tourist guides and reviews related to the specified location. The service provider can also refer to reviews related to the specified location. The service provider can also refer to historical background and cultural information related to the specified location. For example, the service provider can refer to historical background related to the specified location. The service provider can also refer to cultural information related to the specified location. The service provider can also refer to the latest news and event information related to the specified location. For example, the service provider can refer to the latest news related to the specified location. The service provider can also refer to event information related to the specified location. This allows for improved accuracy of the service by referring to relevant literature for the specified location. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes relevant literature for a specified location as input and improves the accuracy of the information provision by referring to the relevant literature.
[0085] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0086] The reception desk can automatically suggest relevant past search history based on the user's input. For example, if a user enters "cafe with an ocean view," the reception desk will display search history such as "seaside restaurant" or "beach cafe" that the user has searched for in the past. Also, if a user uploads a still image and asks "Where is this photo taken?", the reception desk can suggest similar photos that the user has uploaded in the past. This allows users to identify locations more efficiently by referring to their past search history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can take the user's past search history as input and suggest history using an AI model that suggests relevant history.
[0087] The analysis unit can estimate the user's emotions during image analysis and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is excited, the analysis results can be displayed in more detail, and if the user is calm, they can be displayed concisely. Furthermore, if the user is feeling anxious, supplementary information can be added to explain the analysis results in an easy-to-understand manner. This helps the user understand the results by providing an appropriate display method according to their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI, or not. For example, the analysis unit can take user emotion data as input and estimate emotions using an AI model that estimates emotions.
[0088] The analysis unit can estimate the user's emotions during text analysis and adjust the analysis priority based on the estimated emotions. For example, if the user is excited, the analysis can be performed quickly, while if the user is calm, the analysis can be performed at a normal speed. Furthermore, if the user is feeling anxious, the analysis results can be provided in more detail. This improves user satisfaction by providing appropriate analysis speed and detail according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI, or not. For example, the analysis unit can take user emotion data as input and estimate emotions using an emotion estimation AI model.
[0089] The service provider can estimate the user's emotions when providing detailed information about a specified location and adjust the way the information is presented based on the estimated emotions. For example, if the user is excited, the information may be displayed in a visually appealing way; if the user is calm, the information may be displayed concisely. Furthermore, if the user is feeling anxious, supplementary information may be added to make the information easier to understand. This helps the user understand the information by providing an appropriate method of presentation tailored to their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, or not. For example, the service provider can take user emotion data as input and estimate emotions using an AI model that estimates emotions.
[0090] The service provider can estimate the user's emotions when providing a list of specified locations and adjust the display order of the list based on the estimated emotions. For example, if the user is excited, visually appealing locations can be displayed at the top of the list; if the user is calm, the list can be displayed alphabetically. Furthermore, if the user is feeling anxious, safer locations can be displayed at the top of the list. This improves user satisfaction by providing an appropriate list display method tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the processing described above in the service provider may be performed using AI, or not. For example, the service provider can take user emotion data as input and estimate emotions using an AI model that estimates emotions.
[0091] The reception desk can automatically suggest relevant news and event information based on user input. For example, if a user enters "cafe with an ocean view," the reception desk will display events and the latest news in that area. Similarly, if a user uploads a still image with the question "Where is this photo taken?", the reception desk can suggest the latest news and event information related to that location. This allows users to identify locations more efficiently while referring to relevant and up-to-date information. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk could use an AI model that takes user input as input and suggests relevant news and event information to propose information.
[0092] The analysis unit can optimize its analysis algorithm by referring to the user's past analysis results during image analysis. For example, it can apply the optimal algorithm to similar images based on the analysis results of images previously uploaded by the user. It can also apply the optimal algorithm to similar conditions based on the analysis results of text containing conditions previously entered by the user. In this way, by referring to the user's past analysis results, the analysis algorithm can be optimized and the analysis accuracy can be improved. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can optimize its algorithm using an AI model that optimizes the analysis algorithm based on the user's past analysis results as input.
[0093] The information provider can prioritize providing highly relevant information based on the user's current geographical location when providing detailed information about a specified location. For example, it can prioritize displaying information about locations close to the user's current location, and if the user is interested in a particular region, it can prioritize displaying information related to that region. Furthermore, if the user is traveling, it can prioritize providing information related to their travel destination. This allows for the provision of more appropriate information by providing highly relevant information based on the user's current geographical location. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes the user's geographical location as input and provides highly relevant information.
[0094] The reception desk can automatically suggest relevant social media posts based on the user's input. For example, if a user enters "cafe with an ocean view," the reception desk will display popular social media posts in that area. Similarly, if a user uploads a still image with the question "Where is this photo taken?", the reception desk can suggest social media posts related to that location. This allows users to identify locations more efficiently by referring to relevant social media posts. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk could suggest posts using an AI model that takes the user's input as input and suggests relevant social media posts.
[0095] The information provider can prioritize providing highly relevant information based on the user's past search history when providing detailed information about a specified location. For example, it can prioritize displaying information about relevant locations based on the user's past search history, such as "seaside restaurants" or "beach cafes." It can also prioritize providing information about relevant locations based on the user's past search history, including the text containing the conditions they have entered. This allows for the provision of more appropriate information by providing highly relevant information based on the user's past search history. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the information provider can use an AI model that takes the user's past search history as input and provides highly relevant information.
[0096] The following briefly describes the processing flow for example form 2.
[0097] Step 1: The reception desk accepts still images, videos, or text files containing conditions from users. For example, it can accept still images in JPEG format, videos in MP4 format, and text files. It accepts still images, recorded videos, or text files containing conditions uploaded by users. Step 2: The analysis unit analyzes the information received by the reception unit to identify a specific location. For example, it uses multimodal AI to analyze the background in still images and videos, as well as the text describing the conditions. It uses image recognition algorithms and natural language processing technology to analyze the background in still images and videos, as well as the text describing the conditions, and identifies the location based on map data. Step 3: The provisioning unit provides the locations identified by the analysis unit. For example, it provides the user with detailed information about the identified locations, such as addresses, geographical coordinates, business hours, and contact information. It can also provide a list of identified locations based on a written description of the conditions.
[0098] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0099] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0100] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0101] Each of the multiple elements described above, including the reception unit, analysis unit, and provision unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart device 14 and receives input from the user in the form of still images, videos, or text describing conditions. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12 and analyzes the input information using multimodal AI to identify a specific location. The provision unit is implemented by the control unit 46A of the smart device 14 or the identification processing unit 290 of the data processing unit 12 and provides the user with detailed information about the identified location. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0102] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0103] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0104] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0105] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0106] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0107] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0108] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0109] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0110] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0111] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0112] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0113] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0114] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0115] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0116] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0117] Each of the multiple elements described above, including the reception unit, analysis unit, and provision unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart glasses 214 and receives input from the user of still images, videos, or text describing conditions. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12 and analyzes the input information using multimodal AI to identify a specific location. The provision unit is implemented by the control unit 46A of the smart glasses 214 or the identification processing unit 290 of the data processing unit 12 and provides the user with detailed information about the identified location. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0118] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0119] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0120] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0121] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0122] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0123] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0124] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0125] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0126] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0127] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0128] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0129] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0130] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0131] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0132] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0133] Each of the multiple elements described above, including the reception unit, analysis unit, and provision unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the headset terminal 314 and receives input from the user in the form of still images, videos, or text describing conditions. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12 and analyzes the input information using multimodal AI to identify a specific location. The provision unit is implemented by the control unit 46A of the headset terminal 314 or the identification processing unit 290 of the data processing unit 12 and provides the user with detailed information about the identified location. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0134] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0135] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0136] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0137] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0138] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0140] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0141] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0142] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0143] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0144] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0145] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0146] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0147] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0148] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0149] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0150] Each of the multiple elements described above, including the reception unit, analysis unit, and provision unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the robot 414 and receives input from the user in the form of still images, videos, or text describing conditions. The analysis unit is implemented by the identification processing unit 290 of the data processing unit 12 and analyzes the input information using multimodal AI to identify a specific location. The provision unit is implemented by either the control unit 46A of the robot 414 or the identification processing unit 290 of the data processing unit 12 and provides the user with detailed information about the identified location. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0151] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0152] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0153] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0154] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0155] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0156] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0157] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0158] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0159] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0160] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0161] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0162] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0163] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0164] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0165] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0166] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0167] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0168] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0169] (Note 1) A reception area that accepts still images, videos, or text containing conditions from users, An analysis unit analyzes the information received by the reception unit and identifies a specific location, The system comprises a providing unit that provides a location identified by the analysis unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, It is equipped with an image analysis unit that analyzes the background in still images and videos. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, It includes a text analysis unit that analyzes text containing specified conditions. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned supply unit is, Provide users with detailed information about the identified location. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned supply unit is, Provides a list of locations identified based on a document describing the conditions. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is When receiving input, filtering is performed based on the user's current areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is The system estimates the user's emotions and determines the priority of input requests based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When receiving input, the system analyzes the user's social media activity and accepts relevant input. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, It estimates the user's emotions and adjusts the accuracy of image analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, During image analysis, the analysis algorithm is optimized by referring to past analysis data. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, When analyzing images, different analysis methods are applied depending on the image category. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, It estimates the user's emotions and adjusts the text analysis representation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, During text analysis, adjust the level of detail based on the importance of the text. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, When analyzing text, different analysis algorithms are applied depending on the category of the text. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned supply unit is, It estimates the user's emotions and adjusts how the information provided is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned supply unit is, When providing the service, adjust the level of detail based on the importance of the identified location. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned supply unit is, When providing the service, different delivery methods will be applied depending on the category of the specified location. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned supply unit is, It estimates the user's emotions and prioritizes the information provided based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned supply unit is, When serving, the order of serving will be adjusted based on the relevance of the identified location. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned supply unit is, When providing the information, we will improve the accuracy of the provision by referring to relevant literature for the identified location. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0170] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception area that accepts still images, videos, or text containing conditions from users, An analysis unit analyzes the information received by the reception unit and identifies a specific location, The system comprises a providing unit that provides a location identified by the analysis unit. A system characterized by the following features.
2. The aforementioned analysis unit, It is equipped with an image analysis unit that analyzes the background in still images and videos. The system according to feature 1.
3. The aforementioned analysis unit, It includes a text analysis unit that analyzes text containing specified conditions. The system according to feature 1.
4. The aforementioned supply unit is, Provide users with detailed information about the identified location. The system according to feature 1.
5. The aforementioned supply unit is, Provides a list of locations identified based on a document describing the conditions. The system according to feature 1.
6. The aforementioned reception unit is The system estimates the user's emotions and adjusts the timing of input acceptance based on the estimated emotions. The system according to feature 1.
7. The aforementioned reception unit is Analyze the user's past input history and select the optimal input method. The system according to feature 1.
8. The aforementioned reception unit is When receiving input, filtering is performed based on the user's current areas of interest. The system according to feature 1.
9. The aforementioned reception unit is The system estimates the user's emotions and determines the priority of input requests based on those estimated emotions. The system according to feature 1.
10. The aforementioned reception unit is When receiving input, the system prioritizes accepting inputs that are highly relevant based on the user's geographical location. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A