system
The system uses AI to analyze camera footage and provide audio responses, addressing the challenge of visually impaired individuals recalling past events, enabling richer memory recall through audio-based visual information retrieval and storage.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Visually impaired individuals have difficulty recalling past events due to reliance on audio information alone, lacking visual memories.
A system comprising a reception unit, analysis unit, provision unit, storage unit, and search unit, utilizing AI technology to analyze camera footage, provide audio responses, store information, and enable visually impaired individuals to recall past events through a database search.
Enables visually impaired individuals to recall past events in a more enriching way by providing audio-based visual information retrieval and storage, enhancing their memory recall.
Smart Images

Figure 2026072760000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that when visually impaired people recall past events, they have no choice but to rely solely on audio information and it is difficult for them to have visual memories.
[0005] The system according to the embodiment aims to enable visually impaired people to recall past events more richly.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a reception unit, an analysis unit, a provision unit, a storage unit, and a search unit. The reception unit receives inquiries from visually impaired persons. The analysis unit analyzes the camera footage based on the inquiries received by the reception unit. The provision unit provides the information analyzed by the analysis unit in audio format. The storage unit stores the information provided by the provision unit in a database. The search unit searches the information stored in the storage unit and provides it to visually impaired persons. [Effects of the Invention]
[0007] The system according to this embodiment can enable visually impaired individuals to reflect on past events in a more enriching way. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) An embodiment of the present invention provides a system for assisting the visually impaired that utilizes AI technology to easily retrieve past visual information, enabling visually impaired individuals to recall more events. When a visually impaired person asks a question using the device, the AI analyzes the information captured by the camera and provides an audible response. For example, if the user asks, "What was that building I saw on my last trip?", the AI analyzes the camera's image and responds, "It was a historical church." This response information is stored in a database to aid the visually impaired person's memory. Furthermore, if the visually impaired person tells the device the date, time, or event they want to recall, the AI organizes the response information and provides the answer. For example, if the user asks, "What was I doing one morning a week ago?", the AI searches the database and responds, "I was looking out the window to check the weather. It was raining." This system allows visually impaired individuals to retain more memories and reflect on the past. For example, if the user asks, "What was I having for dinner yesterday?", the AI responds, "It was fish and miso soup. The fish was well-grilled saury." Furthermore, when asked, "When did I meet so-and-so?", the AI replies, "I met them a month ago at XX. We had a conversation like this..." In this way, visually impaired individuals can recall everyday events in detail and develop rich memories. Assistance systems for the visually impaired are important tools for improving the quality of life for them. This allows visually impaired individuals to easily recall past visual information and develop rich memories.
[0029] The visually impaired assistance system according to this embodiment comprises a reception unit, an analysis unit, a provision unit, a storage unit, and a search unit. The reception unit receives inquiries from visually impaired persons. Inquiries from visually impaired persons include, but are not limited to, voice input or text input. The reception unit receives inquiries from visually impaired persons using, for example, speech recognition technology. The reception unit can also receive inquiries from visually impaired persons using text input. For example, speech recognition technology recognizes the voice of a visually impaired person with high accuracy and converts it into text data. Text input allows visually impaired persons to input inquiries using a keyboard or touchscreen. The analysis unit analyzes the camera image based on the inquiries received by the reception unit. The analysis unit analyzes the camera image using, for example, image recognition technology. The analysis unit can also analyze the camera image using an object detection algorithm. For example, image recognition technology detects specific objects from the camera image and analyzes their information. An object detection algorithm identifies the position and shape of objects from the camera image and analyzes their information. The provision unit provides the information analyzed by the analysis unit in audio format. The provision unit provides the analyzed information in audio format, for example, using speech synthesis technology. The provision unit can also adjust the audio quality before providing the information. For example, speech synthesis technology converts text data into natural-sounding audio and provides it to visually impaired individuals. The audio quality is adjusted to be easily understood by visually impaired individuals. The storage unit stores the information provided by the provision unit in a database. The storage unit stores the information in the database, for example, by specifying the data format. The storage unit can also store the information in the database with a set retention period. For example, the data format includes text format and audio format. The retention period is set so that the information is stored for a certain period of time. The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. The search unit searches the database, for example, using a search algorithm. The search unit can also search the database with set filtering conditions. For example, the search algorithm searches for relevant information based on the visually impaired person's inquiry. The filtering conditions are set so that the visually impaired person can search for information based on specific conditions.As a result, the visual impairment support system according to this embodiment allows visually impaired individuals to easily recall past visual information and maintain a rich memory.
[0030] The reception desk accepts inquiries from visually impaired individuals. These inquiries may include, but are not limited to, voice input or text input. The reception desk uses, for example, speech recognition technology to accept inquiries from visually impaired individuals. Specifically, the speech recognition technology accurately recognizes the voices of visually impaired individuals and converts them into text data. This speech recognition technology has a noise-canceling function, which removes ambient noise and allows for clear recognition of the voices of visually impaired individuals. Furthermore, the speech recognition technology supports multiple languages, so it can accurately recognize inquiries made by visually impaired individuals in different languages. Text input allows visually impaired individuals to input inquiries using a keyboard or touchscreen. For keyboard input, braille keyboards and keyboards with voice feedback functions are provided to make them easier for visually impaired individuals to use. For touchscreen input, vibration functions and voice guidance are incorporated so that visually impaired individuals can input while receiving tactile feedback. In this way, the reception desk provides an environment in which visually impaired individuals can easily make inquiries using voice or text. Furthermore, when receiving inquiries from visually impaired individuals, the reception desk can refer to the user's profile information and suggest the most suitable input method for each individual user. For example, based on past usage history and user settings, it can recommend voice input to users who are proficient in voice input, and text input to users who are proficient in text input. In this way, the reception desk supports visually impaired individuals in making inquiries in the most user-friendly way.
[0031] The analysis unit analyzes the camera footage based on the questions received by the reception unit. The analysis unit analyzes the camera footage using, for example, image recognition technology. Specifically, image recognition technology detects specific objects from the camera footage and analyzes that information. For example, if a visually impaired person asks, "What is that object in front of me?", the analysis unit analyzes the camera footage and identifies the object in front of them. The object detection algorithm identifies the position and shape of objects from the camera footage and analyzes that information. For example, if a visually impaired person asks, "Is it okay to go down this road?", the analysis unit analyzes the camera footage and identifies the road conditions and the presence or absence of obstacles. The analysis unit performs these analyses using AI. The AI analyzes image data using a deep learning model and recognizes objects with high accuracy. For example, the AI can learn from millions of image data and extract object features to quickly and accurately identify the object that the visually impaired person is asking about. Furthermore, the analysis unit can automatically adjust the camera's field of view in response to the visually impaired person's questions. For example, if a visually impaired person asks, "What is on the right?", the analysis unit will adjust the camera's field of view to the right and analyze the image. This allows the analysis unit to provide appropriate information in response to the visually impaired person's question.
[0032] The information provider unit provides the information analyzed by the analysis unit in audio format. For example, the information provider unit provides the analyzed information in audio format using speech synthesis technology. Specifically, speech synthesis technology converts text data into natural-sounding speech and provides it to visually impaired individuals. Speech synthesis technology can adjust the speed, volume, and sound quality of the speech to make it easier for visually impaired individuals to understand. For example, if a visually impaired person is elderly, the speech speed can be slowed and the volume increased to make it easier to understand. Furthermore, the information provider unit can change the gender and tone of the voice according to the visually impaired person's preferences. For example, if a visually impaired person prefers a female voice, the information provider unit will provide the information in a female voice. In addition, the information provider unit can devise methods of providing information to make it easier for visually impaired individuals to understand. For example, if a visually impaired person requests directions, the information provider unit will explain the route from the current location to the destination step by step, reiterating important points. The information provider unit also provides a repeat function in case the visually impaired person wants to hear the information again. This allows the information provider unit to provide visually impaired individuals with the necessary information accurately and clearly.
[0033] The storage unit stores information provided by the service provider in a database. The storage unit stores information in the database, for example, by specifying the data format. Specifically, data formats include text format and audio format. For example, information received by a visually impaired person can be saved in the database in text format so that it can be reviewed later. It can also be saved in the database in audio format so that the information can be played back by the visually impaired person. The storage period is set so that the information is stored for a certain period of time. For example, the storage period can be set so that a visually impaired person can review past information within a certain period of time. Furthermore, the storage unit records the usage history of visually impaired people and stores data to provide information that is optimal for each individual user. For example, it records what questions a visually impaired person asked and what information they received in the past, so that this can be used as a reference the next time they use the service. In this way, the storage unit supports visually impaired people in easily recalling past visual information and developing rich memories.
[0034] The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. The search unit searches the database using, for example, a search algorithm. Specifically, the search algorithm searches for relevant information based on the visually impaired person's inquiry. For example, if a visually impaired person asks, "Can you tell me yesterday's directions again?", the search unit searches the database for information about yesterday's directions and provides it to the visually impaired person through the provision unit. Filtering conditions are set so that visually impaired people can search for information based on specific conditions. For example, if a visually impaired person asks, "Can you tell me the information for the past week?", the search unit filters the information to include the past week and searches for relevant information. Furthermore, the search unit organizes and provides the search results so that visually impaired people can easily understand them. For example, the search results are organized by category so that visually impaired people can quickly find the information they need. In addition, the search unit provides the search results audibly using speech synthesis technology so that visually impaired people can confirm the search results by voice. In this way, the search unit supports visually impaired people in easily recalling past visual information and developing rich memories.
[0035] The reception area can provide an interface for visually impaired individuals to ask questions. For example, the reception area can provide a voice interface. For example, the reception area can be equipped with a microphone for visually impaired individuals to ask questions by voice. The reception area can also provide a touch interface. For example, the reception area can be equipped with an interface for visually impaired individuals to ask questions using a touchscreen. This makes it easier for visually impaired individuals to ask questions. Some or all of the above processing in the reception area may be performed using AI, for example, or without AI. For example, the reception area can input the voice input of a visually impaired person into the AI and have the AI perform voice recognition.
[0036] The analysis unit can analyze camera footage and identify information useful to visually impaired individuals. For example, the analysis unit can analyze camera footage using image recognition technology. For instance, it can detect specific objects from camera footage and analyze their information. Alternatively, the analysis unit can analyze camera footage using object detection algorithms. For example, it can identify the location and shape of objects from camera footage and analyze their information. This allows for the identification of information useful to visually impaired individuals, enabling the provision of appropriate information. Some or all of the above-described processes in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input camera footage data into an AI and have the AI perform object detection.
[0037] The service provider can provide the analyzed information in audio format. For example, the service provider can provide the analyzed information in audio format using speech synthesis technology. For example, the service provider can convert text data into natural-sounding audio and provide it to visually impaired individuals. The service provider can also adjust the audio quality before providing information. For example, the service provider can adjust the audio quality so that it is easy for visually impaired individuals to hear. This allows visually impaired individuals to receive information in audio format. Some or all of the above-described processes in the service provider may be performed using, for example, a generative AI, or without a generative AI. For example, the service provider can input text data into a generative AI and have the generative AI generate audio data.
[0038] The storage unit can store the provided information in a database. The storage unit can store information in the database by specifying the data format, for example. For example, the storage unit can store information by specifying the data format, such as text format or audio format. The storage unit can also store information in the database with a set retention period. For example, the storage unit can set a retention period so that the information is stored for a certain period of time. This allows past information of visually impaired individuals to be stored in the database and referenced later. Some or all of the above processing in the storage unit may be performed using AI, for example, or without using AI. For example, the storage unit can input the provided information into AI and have AI perform the storage of the information in the database.
[0039] The search unit can search a database and provide useful information for visually impaired individuals. The search unit can search the database using, for example, a search algorithm. For example, the search unit can search for relevant information based on a visually impaired person's inquiry. The search unit can also search the database by setting filtering conditions. For example, the search unit can set filtering conditions so that a visually impaired person can search for information based on specific criteria. This allows visually impaired individuals to easily search and refer to past information. Some or all of the above-described processes in the search unit may be performed using, for example, AI, or not. For example, the search unit can input the visually impaired person's inquiry data into the AI and have the AI perform the search for relevant information.
[0040] The reception desk can analyze the user's past inquiry history and select the optimal reception method. For example, the reception desk can automatically display frequently asked questions from the user's past as candidates. The reception desk can also prioritize suggesting reception methods (voice, text, etc.) that the user has used in the past. Furthermore, the reception desk can predict and suggest questions to be used during specific time periods based on the user's past inquiry history. This allows the reception desk to provide the optimal reception method based on past history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past inquiry data into an AI and have the AI select the optimal reception method.
[0041] The reception unit can filter inquiries based on the user's current situation and areas of interest. For example, the reception unit can prioritize inquiries based on the user's current location. It can also filter and prioritize inquiries based on the user's areas of interest. Furthermore, the reception unit can suggest appropriate inquiries based on the user's current situation (e.g., traveling, working). This ensures that the reception unit receives appropriate inquiries tailored to the user's situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input the user's current situation data into the AI and have the AI perform the filtering.
[0042] The reception desk can prioritize receiving inquiries that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific location, the reception desk will prioritize inquiries related to that location. Furthermore, if the user is traveling, the reception desk can prioritize inquiries related to their travel destination. Additionally, if the user is at home, the reception desk can prioritize inquiries related to their home. This allows the reception desk to receive appropriate inquiries based on the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's geographical location data into an AI and have the AI select highly relevant inquiries.
[0043] The reception desk can analyze the user's social media activity when receiving an inquiry and accept relevant inquiries. For example, the reception desk can prioritize accepting relevant inquiries based on information the user has shared on social media. It can also suggest relevant inquiries based on the accounts the user follows on social media. Furthermore, the reception desk can analyze the user's social media activity history and prioritize accepting relevant inquiries. This allows for the acceptance of appropriate inquiries based on the user's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media data into an AI and have the AI select relevant inquiries.
[0044] The analysis unit can adjust the level of detail of the analysis based on the importance of the video during the analysis. For example, the analysis unit performs a detailed analysis for important video. It can also perform a simplified analysis for general video. Furthermore, if the video relates to a specific event, the analysis unit can perform an analysis that includes details of the event. This allows for the provision of appropriate analysis results according to the importance of the video. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform an analysis based on importance.
[0045] The analysis unit can apply different analysis algorithms depending on the category of the video during analysis. For example, in the case of landscape video, the analysis unit can apply an algorithm that analyzes the features of the landscape. In the case of video of people, the analysis unit can also apply a face recognition algorithm. Furthermore, in the case of video of buildings, the analysis unit can apply an algorithm that analyzes the structure of the building. This allows for the provision of appropriate analysis results according to the category of the video. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform analysis according to the category.
[0046] The analysis unit can determine the priority of analysis based on the video's shooting date during the analysis. For example, the analysis unit may prioritize analyzing recently shot video. It can also prioritize analyzing video related to a specific event. Furthermore, the analysis unit can prioritize analyzing video within a period specified by the user. This allows for the provision of appropriate analysis results based on the video's shooting date. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform analysis based on the shooting date.
[0047] The analysis unit can adjust the order of analysis based on the relevance of the videos during the analysis process. For example, the analysis unit may prioritize analyzing videos related to a theme specified by the user. It can also prioritize analyzing videos related to the user's past question history. Furthermore, it can prioritize analyzing videos related to the user's current areas of interest. This allows for the provision of appropriate analysis results based on the relevance of the videos. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform relevance-based analysis.
[0048] The information delivery unit can adjust the level of detail provided based on the importance of the information at the time of delivery. For example, the delivery unit can provide a detailed explanation for important information. It can also provide a concise explanation for general information. Furthermore, if the information relates to a specific event, the delivery unit can provide an explanation that includes details of the event. This enables the provision of appropriate information according to its importance. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not. For example, the delivery unit can input information data into AI and have the AI perform the delivery based on importance.
[0049] The information provider can apply different information provision algorithms depending on the category of information at the time of provision. For example, in the case of landscape information, the provider can apply an algorithm that describes the characteristics of the landscape. In the case of person information, the provider can also apply an algorithm that describes the characteristics of the person. Furthermore, in the case of building information, the provider can also apply an algorithm that describes the structure of the building. This enables the provision of appropriate information according to the category of information. Some or all of the above processing in the information provider may be performed using AI, for example, or without using AI. For example, the information provider can input information data into AI and have the AI perform provision according to the category.
[0050] The information delivery unit can determine the priority of information delivery based on when the information was acquired. For example, the delivery unit may prioritize the delivery of recently acquired information. It may also prioritize the delivery of information related to a specific event. Furthermore, the delivery unit may prioritize the delivery of information within a period specified by the user. This enables the delivery of appropriate information based on when the information was acquired. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not using AI. For example, the delivery unit can input information data into AI and have the AI perform delivery based on the acquisition time.
[0051] The information provider can adjust the order of information delivery based on its relevance. For example, the provider may prioritize information related to a theme specified by the user. It can also prioritize information based on the user's past inquiry history. Furthermore, it can prioritize information based on the user's current areas of interest. This enables the provision of appropriate information based on its relevance. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the provider can input information data into AI and have the AI perform relevance-based delivery.
[0052] The storage unit can optimize its storage algorithm by referring to past stored data during storage. For example, the storage unit can analyze past stored data and select the optimal storage algorithm. It can also apply an algorithm to eliminate duplicate data from past stored data. Furthermore, the storage unit can propose an efficient data storage method based on past stored data. This enables optimal data storage based on past data. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input past stored data into AI and have the AI select the optimal storage algorithm.
[0053] The storage unit can weight the stored data based on when the information was acquired. For example, the storage unit can assign a higher weight to recently acquired information. It can also assign a higher weight to information related to a specific event. Furthermore, the storage unit can assign a higher weight to information within a period specified by the user. This enables appropriate data storage based on when the information was acquired. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input information data into AI and have the AI perform weighting based on the acquisition time.
[0054] The search unit can select the optimal search method by referring to the user's past search history during a search. For example, the search unit can automatically display keywords that the user has frequently searched for in the past as suggestions. The search unit can also prioritize suggesting search methods (voice, text, etc.) that the user has used in the past. Furthermore, the search unit can predict and suggest search methods to be used at specific times based on the user's past search history. This allows for the provision of optimal search results based on past search history. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's past search data into AI and have the AI select the optimal search method.
[0055] The search unit can customize the search method based on the user's current situation during a search. For example, the search unit can prioritize providing relevant search results based on the user's current location. It can also filter relevant search results based on the user's current areas of interest. Furthermore, the search unit can suggest appropriate search methods based on the user's current situation (e.g., traveling, working, etc.). This allows for the provision of appropriate search results tailored to the user's current situation. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's current situation data into the AI and have the AI perform the customization of the search method.
[0056] The search unit can select the optimal search method by considering the user's geographical location information during a search. For example, if the user is in a specific location, the search unit can prioritize providing search results related to that location. Furthermore, if the user is traveling, the search unit can prioritize providing search results related to their travel destination. Additionally, if the user is at home, the search unit can prioritize providing search results related to their home. This allows the system to provide optimal search results based on the user's geographical location information. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's geographical location data into AI and have the AI select the optimal search method.
[0057] The search unit can analyze the user's social media activity during a search and suggest search methods. For example, the search unit can prioritize providing relevant search results based on information the user has shared on social media. It can also suggest relevant search results based on accounts the user follows on social media. Furthermore, the search unit can analyze the user's social media activity history and prioritize providing relevant search results. This allows for the provision of optimal search results based on the user's social media activity. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's social media data into AI and have the AI perform the search method suggestion.
[0058] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0059] The visual impairment assistance system can also be equipped with a health management unit that monitors the user's health status. For example, the health management unit can measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. It can also record the user's exercise level and provide appropriate exercise advice. Furthermore, it can record the user's diet and provide advice on nutritional balance. This makes health management easier for visually impaired individuals and improves their quality of life.
[0060] The visually impaired assistance system can also include a navigation unit that acquires the user's location information and provides information about the surrounding environment based on that location. For example, if the user is in a specific location, the navigation unit can provide information about buildings and facilities around that location. It can also provide directions to the destination if the user is on the move. Furthermore, if the user is using public transportation, it can provide transfer information and service status. This makes travel smoother for visually impaired people, allowing them to go out with peace of mind.
[0061] The visual impairment assistance system can also include a personalization section that provides customized information based on the user's hobbies and interests. For example, if the user is interested in music, the personalization section can provide the latest music information and recommended artists. If the user is interested in sports, it can provide match results and player information. Furthermore, if the user is interested in cooking, it can provide new recipes and cooking tips. This makes it possible to provide information tailored to the interests and hobbies of visually impaired individuals.
[0062] The support system for the visually impaired can also include an educational section to further assist the user's learning. For example, if the user wants to learn a new language, the educational section can provide language learning materials and practice exercises. It can also provide information and materials related to a specific field of interest. Furthermore, if the user wants to prepare for an exam, it can provide mock exams and past exam questions. This effectively supports the learning of visually impaired individuals.
[0063] The system for assisting the visually impaired can also include a social support section to further support the user's social activities. For example, if the user wants to contact friends or family, the social support section can provide an easy-to-use interface. It can also provide relevant information if the user wants to participate in events or gatherings. Furthermore, if the user wants to make new friends, it can match them with people who share common hobbies and interests. This can stimulate social activity for the visually impaired and reduce feelings of isolation.
[0064] The following briefly describes the processing flow for example form 1.
[0065] Step 1: The reception desk receives inquiries from visually impaired individuals. These inquiries may include voice input or text input. For example, speech recognition technology can be used to accurately recognize the voices of visually impaired individuals and convert them into text data. Text input allows visually impaired individuals to enter their inquiries using a keyboard or touchscreen. Step 2: The analysis unit analyzes the camera footage based on the questions received by the reception unit. For example, it analyzes the camera footage using image recognition technology and object detection algorithms to detect specific objects and analyze their information. Step 3: The delivery unit provides the information analyzed by the analysis unit in audio format. For example, it converts the analyzed information into natural-sounding audio using speech synthesis technology and provides it to visually impaired individuals. The audio quality is adjusted to be easily understood by visually impaired individuals. Step 4: The storage unit stores the information provided by the provision unit in a database. For example, it stores the information in a database by specifying the data format and sets a retention period to store the information for a certain period. Step 5: The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. For example, it searches the database using a search algorithm and retrieves relevant information based on the visually impaired person's inquiries. Filtering conditions can also be set to search for information based on specific criteria.
[0066] (Example of form 2) An embodiment of the present invention provides a system for assisting the visually impaired that utilizes AI technology to easily retrieve past visual information, enabling visually impaired individuals to recall more events. When a visually impaired person asks a question using the device, the AI analyzes the information captured by the camera and provides an audible response. For example, if the user asks, "What was that building I saw on my last trip?", the AI analyzes the camera's image and responds, "It was a historical church." This response information is stored in a database to aid the visually impaired person's memory. Furthermore, if the visually impaired person tells the device the date, time, or event they want to recall, the AI organizes the response information and provides the answer. For example, if the user asks, "What was I doing one morning a week ago?", the AI searches the database and responds, "I was looking out the window to check the weather. It was raining." This system allows visually impaired individuals to retain more memories and reflect on the past. For example, if the user asks, "What was I having for dinner yesterday?", the AI responds, "It was fish and miso soup. The fish was well-grilled saury." Furthermore, when asked, "When did I meet so-and-so?", the AI replies, "I met them a month ago at XX. We had a conversation like this..." In this way, visually impaired individuals can recall everyday events in detail and develop rich memories. Assistance systems for the visually impaired are important tools for improving the quality of life for them. This allows visually impaired individuals to easily recall past visual information and develop rich memories.
[0067] The visually impaired assistance system according to this embodiment comprises a reception unit, an analysis unit, a provision unit, a storage unit, and a search unit. The reception unit receives inquiries from visually impaired persons. Inquiries from visually impaired persons include, but are not limited to, voice input or text input. The reception unit receives inquiries from visually impaired persons using, for example, speech recognition technology. The reception unit can also receive inquiries from visually impaired persons using text input. For example, speech recognition technology recognizes the voice of a visually impaired person with high accuracy and converts it into text data. Text input allows visually impaired persons to input inquiries using a keyboard or touchscreen. The analysis unit analyzes the camera image based on the inquiries received by the reception unit. The analysis unit analyzes the camera image using, for example, image recognition technology. The analysis unit can also analyze the camera image using an object detection algorithm. For example, image recognition technology detects specific objects from the camera image and analyzes their information. An object detection algorithm identifies the position and shape of objects from the camera image and analyzes their information. The provision unit provides the information analyzed by the analysis unit in audio format. The provision unit provides the analyzed information in audio format, for example, using speech synthesis technology. The provision unit can also adjust the audio quality before providing the information. For example, speech synthesis technology converts text data into natural-sounding audio and provides it to visually impaired individuals. The audio quality is adjusted to be easily understood by visually impaired individuals. The storage unit stores the information provided by the provision unit in a database. The storage unit stores the information in the database, for example, by specifying the data format. The storage unit can also store the information in the database with a set retention period. For example, the data format includes text format and audio format. The retention period is set so that the information is stored for a certain period of time. The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. The search unit searches the database, for example, using a search algorithm. The search unit can also search the database with set filtering conditions. For example, the search algorithm searches for relevant information based on the visually impaired person's inquiry. The filtering conditions are set so that the visually impaired person can search for information based on specific conditions.As a result, the visual impairment support system according to this embodiment allows visually impaired individuals to easily recall past visual information and maintain a rich memory.
[0068] The reception desk accepts inquiries from visually impaired individuals. These inquiries may include, but are not limited to, voice input or text input. The reception desk uses, for example, speech recognition technology to accept inquiries from visually impaired individuals. Specifically, the speech recognition technology accurately recognizes the voices of visually impaired individuals and converts them into text data. This speech recognition technology has a noise-canceling function, which removes ambient noise and allows for clear recognition of the voices of visually impaired individuals. Furthermore, the speech recognition technology supports multiple languages, so it can accurately recognize inquiries made by visually impaired individuals in different languages. Text input allows visually impaired individuals to input inquiries using a keyboard or touchscreen. For keyboard input, braille keyboards and keyboards with voice feedback functions are provided to make them easier for visually impaired individuals to use. For touchscreen input, vibration functions and voice guidance are incorporated so that visually impaired individuals can input while receiving tactile feedback. In this way, the reception desk provides an environment in which visually impaired individuals can easily make inquiries using voice or text. Furthermore, when receiving inquiries from visually impaired individuals, the reception desk can refer to the user's profile information and suggest the most suitable input method for each individual user. For example, based on past usage history and user settings, it can recommend voice input to users who are proficient in voice input, and text input to users who are proficient in text input. In this way, the reception desk supports visually impaired individuals in making inquiries in the most user-friendly way.
[0069] The analysis unit analyzes the camera footage based on the questions received by the reception unit. The analysis unit analyzes the camera footage using, for example, image recognition technology. Specifically, image recognition technology detects specific objects from the camera footage and analyzes that information. For example, if a visually impaired person asks, "What is that object in front of me?", the analysis unit analyzes the camera footage and identifies the object in front of them. The object detection algorithm identifies the position and shape of objects from the camera footage and analyzes that information. For example, if a visually impaired person asks, "Is it okay to go down this road?", the analysis unit analyzes the camera footage and identifies the road conditions and the presence or absence of obstacles. The analysis unit performs these analyses using AI. The AI analyzes image data using a deep learning model and recognizes objects with high accuracy. For example, the AI can learn from millions of image data and extract object features to quickly and accurately identify the object that the visually impaired person is asking about. Furthermore, the analysis unit can automatically adjust the camera's field of view in response to the visually impaired person's questions. For example, if a visually impaired person asks, "What is on the right?", the analysis unit will adjust the camera's field of view to the right and analyze the image. This allows the analysis unit to provide appropriate information in response to the visually impaired person's question.
[0070] The information provider unit provides the information analyzed by the analysis unit in audio format. For example, the information provider unit provides the analyzed information in audio format using speech synthesis technology. Specifically, speech synthesis technology converts text data into natural-sounding speech and provides it to visually impaired individuals. Speech synthesis technology can adjust the speed, volume, and sound quality of the speech to make it easier for visually impaired individuals to understand. For example, if a visually impaired person is elderly, the speech speed can be slowed and the volume increased to make it easier to understand. Furthermore, the information provider unit can change the gender and tone of the voice according to the visually impaired person's preferences. For example, if a visually impaired person prefers a female voice, the information provider unit will provide the information in a female voice. In addition, the information provider unit can devise methods of providing information to make it easier for visually impaired individuals to understand. For example, if a visually impaired person requests directions, the information provider unit will explain the route from the current location to the destination step by step, reiterating important points. The information provider unit also provides a repeat function in case the visually impaired person wants to hear the information again. This allows the information provider unit to provide visually impaired individuals with the necessary information accurately and clearly.
[0071] The storage unit stores information provided by the service provider in a database. The storage unit stores information in the database, for example, by specifying the data format. Specifically, data formats include text format and audio format. For example, information received by a visually impaired person can be saved in the database in text format so that it can be reviewed later. It can also be saved in the database in audio format so that the information can be played back by the visually impaired person. The storage period is set so that the information is stored for a certain period of time. For example, the storage period can be set so that a visually impaired person can review past information within a certain period of time. Furthermore, the storage unit records the usage history of visually impaired people and stores data to provide information that is optimal for each individual user. For example, it records what questions a visually impaired person asked and what information they received in the past, so that this can be used as a reference the next time they use the service. In this way, the storage unit supports visually impaired people in easily recalling past visual information and developing rich memories.
[0072] The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. The search unit searches the database using, for example, a search algorithm. Specifically, the search algorithm searches for relevant information based on the visually impaired person's inquiry. For example, if a visually impaired person asks, "Can you tell me yesterday's directions again?", the search unit searches the database for information about yesterday's directions and provides it to the visually impaired person through the provision unit. Filtering conditions are set so that visually impaired people can search for information based on specific conditions. For example, if a visually impaired person asks, "Can you tell me the information for the past week?", the search unit filters the information to include the past week and searches for relevant information. Furthermore, the search unit organizes and provides the search results so that visually impaired people can easily understand them. For example, the search results are organized by category so that visually impaired people can quickly find the information they need. In addition, the search unit provides the search results audibly using speech synthesis technology so that visually impaired people can confirm the search results by voice. In this way, the search unit supports visually impaired people in easily recalling past visual information and developing rich memories.
[0073] The reception area can provide an interface for visually impaired individuals to ask questions. For example, the reception area can provide a voice interface. For example, the reception area can be equipped with a microphone for visually impaired individuals to ask questions by voice. The reception area can also provide a touch interface. For example, the reception area can be equipped with an interface for visually impaired individuals to ask questions using a touchscreen. This makes it easier for visually impaired individuals to ask questions. Some or all of the above processing in the reception area may be performed using AI, for example, or without AI. For example, the reception area can input the voice input of a visually impaired person into the AI and have the AI perform voice recognition.
[0074] The analysis unit can analyze camera footage and identify information useful to visually impaired individuals. For example, the analysis unit can analyze camera footage using image recognition technology. For instance, it can detect specific objects from camera footage and analyze their information. Alternatively, the analysis unit can analyze camera footage using object detection algorithms. For example, it can identify the location and shape of objects from camera footage and analyze their information. This allows for the identification of information useful to visually impaired individuals, enabling the provision of appropriate information. Some or all of the above-described processes in the analysis unit may be performed using AI, or without AI. For example, the analysis unit can input camera footage data into an AI and have the AI perform object detection.
[0075] The service provider can provide the analyzed information in audio format. For example, the service provider can provide the analyzed information in audio format using speech synthesis technology. For example, the service provider can convert text data into natural-sounding audio and provide it to visually impaired individuals. The service provider can also adjust the audio quality before providing information. For example, the service provider can adjust the audio quality so that it is easy for visually impaired individuals to hear. This allows visually impaired individuals to receive information in audio format. Some or all of the above-described processes in the service provider may be performed using, for example, a generative AI, or without a generative AI. For example, the service provider can input text data into a generative AI and have the generative AI generate audio data.
[0076] The storage unit can store the provided information in a database. The storage unit can store information in the database by specifying the data format, for example. For example, the storage unit can store information by specifying the data format, such as text format or audio format. The storage unit can also store information in the database with a set retention period. For example, the storage unit can set a retention period so that the information is stored for a certain period of time. This allows past information of visually impaired individuals to be stored in the database and referenced later. Some or all of the above processing in the storage unit may be performed using AI, for example, or without using AI. For example, the storage unit can input the provided information into AI and have AI perform the storage of the information in the database.
[0077] The search unit can search a database and provide useful information for visually impaired individuals. The search unit can search the database using, for example, a search algorithm. For example, the search unit can search for relevant information based on a visually impaired person's inquiry. The search unit can also search the database by setting filtering conditions. For example, the search unit can set filtering conditions so that a visually impaired person can search for information based on specific criteria. This allows visually impaired individuals to easily search and refer to past information. Some or all of the above-described processes in the search unit may be performed using, for example, AI, or not. For example, the search unit can input the visually impaired person's inquiry data into the AI and have the AI perform the search for relevant information.
[0078] The reception unit can estimate the user's emotions and adjust the way it handles inquiries based on those emotions. For example, if the user is stressed, the reception unit can provide a simple interface and minimize the inquiry process. If the user is relaxed, the reception unit can also provide detailed inquiry options and suggest a customizable reception method. Furthermore, if the user is in a hurry, the reception unit can prioritize voice input to quickly handle inquiries. This allows for the provision of an appropriate reception method tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not. For example, the reception unit can input the user's voice data into a generative AI and have the generative AI perform emotion estimation.
[0079] The reception desk can analyze the user's past inquiry history and select the optimal reception method. For example, the reception desk can automatically display frequently asked questions from the user's past as candidates. The reception desk can also prioritize suggesting reception methods (voice, text, etc.) that the user has used in the past. Furthermore, the reception desk can predict and suggest questions to be used during specific time periods based on the user's past inquiry history. This allows the reception desk to provide the optimal reception method based on past history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past inquiry data into an AI and have the AI select the optimal reception method.
[0080] The reception unit can filter inquiries based on the user's current situation and areas of interest. For example, the reception unit can prioritize inquiries based on the user's current location. It can also filter and prioritize inquiries based on the user's areas of interest. Furthermore, the reception unit can suggest appropriate inquiries based on the user's current situation (e.g., traveling, working). This ensures that the reception unit receives appropriate inquiries tailored to the user's situation and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input the user's current situation data into the AI and have the AI perform the filtering.
[0081] The reception desk can estimate the user's emotions and determine the priority of questions to answer based on the estimated emotions. For example, if the user is nervous, the reception desk may prioritize important questions. If the user is relaxed, the reception desk may also prioritize detailed questions. Furthermore, if the user is in a hurry, the reception desk may prioritize questions that require a quick response. This allows questions to be answered with appropriate priority according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception desk may be performed using AI or not using AI. For example, the reception desk can input the user's voice data into a generative AI and have the generative AI perform emotion estimation.
[0082] The reception desk can prioritize receiving inquiries that are highly relevant, taking into account the user's geographical location. For example, if the user is in a specific location, the reception desk will prioritize inquiries related to that location. Furthermore, if the user is traveling, the reception desk can prioritize inquiries related to their travel destination. Additionally, if the user is at home, the reception desk can prioritize inquiries related to their home. This allows the reception desk to receive appropriate inquiries based on the user's geographical location. Some or all of the above processing in the reception desk may be performed using AI, or not. For example, the reception desk can input the user's geographical location data into an AI and have the AI select highly relevant inquiries.
[0083] The reception desk can analyze the user's social media activity when receiving an inquiry and accept relevant inquiries. For example, the reception desk can prioritize accepting relevant inquiries based on information the user has shared on social media. It can also suggest relevant inquiries based on the accounts the user follows on social media. Furthermore, the reception desk can analyze the user's social media activity history and prioritize accepting relevant inquiries. This allows for the acceptance of appropriate inquiries based on the user's social media activity. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's social media data into an AI and have the AI select relevant inquiries.
[0084] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the user is relaxed, the analysis unit can provide detailed analysis results. If the user is in a hurry, the analysis unit can also provide concise analysis results. Furthermore, if the user is excited, the analysis unit can provide visually stimulating analysis results. This allows for the provision of appropriate analysis results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's voice data into a generative AI and have the generative AI perform emotion estimation.
[0085] The analysis unit can adjust the level of detail of the analysis based on the importance of the video during the analysis. For example, the analysis unit performs a detailed analysis for important video. It can also perform a simplified analysis for general video. Furthermore, if the video relates to a specific event, the analysis unit can perform an analysis that includes details of the event. This allows for the provision of appropriate analysis results according to the importance of the video. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform an analysis based on importance.
[0086] The analysis unit can apply different analysis algorithms depending on the category of the video during analysis. For example, in the case of landscape video, the analysis unit can apply an algorithm that analyzes the features of the landscape. In the case of video of people, the analysis unit can also apply a face recognition algorithm. Furthermore, in the case of video of buildings, the analysis unit can apply an algorithm that analyzes the structure of the building. This allows for the provision of appropriate analysis results according to the category of the video. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform analysis according to the category.
[0087] The analysis unit can estimate the user's emotions and adjust the length of the analysis based on the estimated emotions. For example, if the user is in a hurry, the analysis unit can provide a short, concise analysis. If the user is relaxed, the analysis unit can also provide a detailed analysis. Furthermore, if the user is excited, the analysis unit can provide a visually stimulating analysis. This allows for the provision of appropriate analysis results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the user's voice data into the generative AI and have the generative AI perform emotion estimation.
[0088] The analysis unit can determine the priority of analysis based on the video's shooting date during the analysis. For example, the analysis unit may prioritize analyzing recently shot video. It can also prioritize analyzing video related to a specific event. Furthermore, the analysis unit can prioritize analyzing video within a period specified by the user. This allows for the provision of appropriate analysis results based on the video's shooting date. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform analysis based on the shooting date.
[0089] The analysis unit can adjust the order of analysis based on the relevance of the videos during the analysis process. For example, the analysis unit may prioritize analyzing videos related to a theme specified by the user. It can also prioritize analyzing videos related to the user's past question history. Furthermore, it can prioritize analyzing videos related to the user's current areas of interest. This allows for the provision of appropriate analysis results based on the relevance of the videos. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input video data into AI and have the AI perform relevance-based analysis.
[0090] The service provider can estimate the user's emotions and adjust the presentation of the service based on the estimated emotions. For example, if the user is relaxed, the service provider can provide detailed information. If the user is in a hurry, the service provider can provide concise information. Furthermore, if the user is excited, the service provider can provide visually stimulating information. This enables the provision of appropriate information according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the user's voice data into the generative AI and have the generative AI perform emotion estimation.
[0091] The information delivery unit can adjust the level of detail provided based on the importance of the information at the time of delivery. For example, the delivery unit can provide a detailed explanation for important information. It can also provide a concise explanation for general information. Furthermore, if the information relates to a specific event, the delivery unit can provide an explanation that includes details of the event. This enables the provision of appropriate information according to its importance. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not. For example, the delivery unit can input information data into AI and have the AI perform the delivery based on importance.
[0092] The information provider can apply different information provision algorithms depending on the category of information at the time of provision. For example, in the case of landscape information, the provider can apply an algorithm that describes the characteristics of the landscape. In the case of person information, the provider can also apply an algorithm that describes the characteristics of the person. Furthermore, in the case of building information, the provider can also apply an algorithm that describes the structure of the building. This enables the provision of appropriate information according to the category of information. Some or all of the above processing in the information provider may be performed using AI, for example, or without using AI. For example, the information provider can input information data into AI and have the AI perform provision according to the category.
[0093] The service provider can estimate the user's emotions and adjust the length of the service based on the estimated emotions. For example, if the user is in a hurry, the service provider can provide short, concise information. If the user is relaxed, the service provider can provide detailed information. Furthermore, if the user is excited, the service provider can provide visually stimulating information. This enables the provision of appropriate information according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, or not using AI. For example, the service provider can input the user's voice data into a generative AI and have the generative AI perform emotion estimation.
[0094] The information delivery unit can determine the priority of information delivery based on when the information was acquired. For example, the delivery unit may prioritize the delivery of recently acquired information. It may also prioritize the delivery of information related to a specific event. Furthermore, the delivery unit may prioritize the delivery of information within a period specified by the user. This enables the delivery of appropriate information based on when the information was acquired. Some or all of the above processing in the delivery unit may be performed using AI, for example, or not using AI. For example, the delivery unit can input information data into AI and have the AI perform delivery based on the acquisition time.
[0095] The information provider can adjust the order of information delivery based on its relevance. For example, the provider may prioritize information related to a theme specified by the user. It can also prioritize information based on the user's past inquiry history. Furthermore, it can prioritize information based on the user's current areas of interest. This enables the provision of appropriate information based on its relevance. Some or all of the above processing in the information provider may be performed using AI, for example, or without AI. For example, the provider can input information data into AI and have the AI perform relevance-based delivery.
[0096] The data storage unit can estimate the user's emotions and select data to store based on the estimated emotions. For example, if the user is relaxed, the data storage unit will store detailed data. If the user is in a hurry, the data storage unit can also store concise data. Furthermore, if the user is excited, the data storage unit can store visually stimulating data. This enables the storage of appropriate data according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data storage unit may be performed using AI, or not using AI. For example, the data storage unit can input the user's voice data into the generative AI and have the generative AI perform emotion estimation.
[0097] The storage unit can optimize its storage algorithm by referring to past stored data during storage. For example, the storage unit can analyze past stored data and select the optimal storage algorithm. It can also apply an algorithm to eliminate duplicate data from past stored data. Furthermore, the storage unit can propose an efficient data storage method based on past stored data. This enables optimal data storage based on past data. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input past stored data into AI and have the AI select the optimal storage algorithm.
[0098] The data storage unit can estimate the user's emotions and adjust the frequency of data storage based on the estimated emotions. For example, if the user is relaxed, the data storage unit will store data frequently. If the user is in a hurry, the data storage unit can store only the minimum necessary data. Furthermore, if the user is excited, the data storage unit can prioritize the storage of important data. This enables appropriate data storage according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data storage unit may be performed using AI, or not using AI. For example, the data storage unit can input the user's voice data into the generative AI and have the generative AI perform emotion estimation.
[0099] The storage unit can weight the stored data based on when the information was acquired. For example, the storage unit can assign a higher weight to recently acquired information. It can also assign a higher weight to information related to a specific event. Furthermore, the storage unit can assign a higher weight to information within a period specified by the user. This enables appropriate data storage based on when the information was acquired. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input information data into AI and have the AI perform weighting based on the acquisition time.
[0100] The search unit can estimate the user's emotions and adjust the search method based on the estimated emotions. For example, if the user is relaxed, the search unit can provide detailed search results. If the user is in a hurry, the search unit can also provide concise search results. Furthermore, if the user is excited, the search unit can provide visually stimulating search results. This allows for the provision of appropriate search results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the search unit may be performed using AI, for example, or not using AI. For example, the search unit can input the user's voice data into a generative AI and have the generative AI perform emotion estimation.
[0101] The search unit can select the optimal search method by referring to the user's past search history during a search. For example, the search unit can automatically display keywords that the user has frequently searched for in the past as suggestions. The search unit can also prioritize suggesting search methods (voice, text, etc.) that the user has used in the past. Furthermore, the search unit can predict and suggest search methods to be used at specific times based on the user's past search history. This allows for the provision of optimal search results based on past search history. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's past search data into AI and have the AI select the optimal search method.
[0102] The search unit can customize the search method based on the user's current situation during a search. For example, the search unit can prioritize providing relevant search results based on the user's current location. It can also filter relevant search results based on the user's current areas of interest. Furthermore, the search unit can suggest appropriate search methods based on the user's current situation (e.g., traveling, working, etc.). This allows for the provision of appropriate search results tailored to the user's current situation. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's current situation data into the AI and have the AI perform the customization of the search method.
[0103] The search unit can estimate the user's emotions and determine search priorities based on the estimated emotions. For example, if the user is stressed, the search unit may prioritize important search results. If the user is relaxed, the search unit may also prioritize detailed search results. Furthermore, if the user is in a hurry, the search unit may prioritize search results that require immediate attention. This allows for the provision of appropriate search results tailored to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the search unit may be performed using AI or not. For example, the search unit can input user voice data into a generative AI and have the generative AI perform emotion estimation.
[0104] The search unit can select the optimal search method by considering the user's geographical location information during a search. For example, if the user is in a specific location, the search unit can prioritize providing search results related to that location. Furthermore, if the user is traveling, the search unit can prioritize providing search results related to their travel destination. Additionally, if the user is at home, the search unit can prioritize providing search results related to their home. This allows the system to provide optimal search results based on the user's geographical location information. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's geographical location data into AI and have the AI select the optimal search method.
[0105] The search unit can analyze the user's social media activity during a search and suggest search methods. For example, the search unit can prioritize providing relevant search results based on information the user has shared on social media. It can also suggest relevant search results based on accounts the user follows on social media. Furthermore, the search unit can analyze the user's social media activity history and prioritize providing relevant search results. This allows for the provision of optimal search results based on the user's social media activity. Some or all of the above processing in the search unit may be performed using AI, for example, or without AI. For example, the search unit can input the user's social media data into AI and have the AI perform the search method suggestion.
[0106] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0107] The visual impairment assistance system can also be equipped with a health management unit that monitors the user's health status. For example, the health management unit can measure the user's heart rate and blood pressure and issue a warning if an abnormality is detected. It can also record the user's exercise level and provide appropriate exercise advice. Furthermore, it can record the user's diet and provide advice on nutritional balance. This makes health management easier for visually impaired individuals and improves their quality of life.
[0108] The visual impairment assistance system can also include a relaxation unit that estimates the user's emotions and provides relaxing music or sounds based on those emotions. For example, if the user is feeling stressed, the relaxation unit can play relaxing music. It can also provide reassuring audio guidance if the user is feeling anxious. Furthermore, if the user is relaxed, it can play pleasant ambient sounds. This allows for relaxation tailored to the user's emotions.
[0109] The visually impaired assistance system can also include a navigation unit that acquires the user's location information and provides information about the surrounding environment based on that location. For example, if the user is in a specific location, the navigation unit can provide information about buildings and facilities around that location. It can also provide directions to the destination if the user is on the move. Furthermore, if the user is using public transportation, it can provide transfer information and service status. This makes travel smoother for visually impaired people, allowing them to go out with peace of mind.
[0110] The visual impairment assistance system may also include a communication support unit that estimates the user's emotions and suggests appropriate communication methods based on those emotions. For example, if the user is anxious, the communication support unit can provide concise and clear instructions. If the user is relaxed, it can provide detailed explanations. Furthermore, if the user is agitated, it can offer advice to calm them down. This enables appropriate communication support tailored to the user's emotions.
[0111] The visual impairment assistance system can also include a personalization section that provides customized information based on the user's hobbies and interests. For example, if the user is interested in music, the personalization section can provide the latest music information and recommended artists. If the user is interested in sports, it can provide match results and player information. Furthermore, if the user is interested in cooking, it can provide new recipes and cooking tips. This makes it possible to provide information tailored to the interests and hobbies of visually impaired individuals.
[0112] The visual impairment assistance system can also include a fitness section that estimates the user's emotions and suggests appropriate exercises based on those emotions. For example, if the user is feeling stressed, the fitness section might suggest relaxing yoga or stretching. If the user is energetic, it could suggest cardio exercises. Furthermore, if the user is relaxed, it could suggest light walking or relaxation exercises. This allows for the provision of exercises appropriate to the user's emotions.
[0113] The support system for the visually impaired can also include an educational section to further assist the user's learning. For example, if the user wants to learn a new language, the educational section can provide language learning materials and practice exercises. It can also provide information and materials related to a specific field of interest. Furthermore, if the user wants to prepare for an exam, it can provide mock exams and past exam questions. This effectively supports the learning of visually impaired individuals.
[0114] The visual impairment support system can also include a nutrition management section that estimates the user's emotions and proposes an appropriate meal plan based on those emotions. For example, if the user is stressed, the nutrition management section might suggest recipes using ingredients with relaxing properties. If the user is energetic, it could suggest a nutritionally balanced meal plan. Furthermore, if the user is relaxed, it could suggest a light meal or dessert. This allows for the provision of an appropriate meal plan tailored to the user's emotions.
[0115] The system for assisting the visually impaired can also include a social support section to further support the user's social activities. For example, if the user wants to contact friends or family, the social support section can provide an easy-to-use interface. It can also provide relevant information if the user wants to participate in events or gatherings. Furthermore, if the user wants to make new friends, it can match them with people who share common hobbies and interests. This can stimulate social activity for the visually impaired and reduce feelings of isolation.
[0116] The visual impairment assistance system may also include a reminder unit that estimates the user's emotions and provides appropriate reminders based on those emotions. For example, if the user is feeling stressed, the reminder unit might suggest a break to relax. It could also prioritize reminders of important tasks if the user is busy. Furthermore, if the user is relaxed, it could remind them of lighter tasks or leisure time. This allows for the provision of appropriate reminders tailored to the user's emotions.
[0117] The following briefly describes the processing flow for example form 2.
[0118] Step 1: The reception desk receives inquiries from visually impaired individuals. These inquiries may include voice input or text input. For example, speech recognition technology can be used to accurately recognize the voices of visually impaired individuals and convert them into text data. Text input allows visually impaired individuals to enter their inquiries using a keyboard or touchscreen. Step 2: The analysis unit analyzes the camera footage based on the questions received by the reception unit. For example, it analyzes the camera footage using image recognition technology and object detection algorithms to detect specific objects and analyze their information. Step 3: The delivery unit provides the information analyzed by the analysis unit in audio format. For example, it converts the analyzed information into natural-sounding audio using speech synthesis technology and provides it to visually impaired individuals. The audio quality is adjusted to be easily understood by visually impaired individuals. Step 4: The storage unit stores the information provided by the provision unit in a database. For example, it stores the information in a database by specifying the data format and sets a retention period to store the information for a certain period. Step 5: The search unit searches the information stored in the storage unit and provides it to visually impaired individuals. For example, it searches the database using a search algorithm and retrieves relevant information based on the visually impaired person's inquiries. Filtering conditions can also be set to search for information based on specific criteria.
[0119] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0120] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0121] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0122] Each of the multiple elements described above, including the reception unit, analysis unit, provision unit, storage unit, and search unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart device 14 and receives inquiries from visually impaired persons using speech recognition technology. The analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes camera images using image recognition technology. The provision unit is implemented by the control unit 46A of the smart device 14 and provides the analyzed information in voice using speech synthesis technology. The storage unit is implemented by the specific processing unit 290 of the data processing unit 12 and stores the provided information in the database 24. The search unit is implemented by the specific processing unit 290 of the data processing unit 12 and searches the stored information and provides it to visually impaired persons. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0123] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0124] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0125] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0126] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0127] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0128] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0129] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0130] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0131] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0132] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0133] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0134] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0135] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0136] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0137] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0138] Each of the multiple elements described above, including the reception unit, analysis unit, provision unit, storage unit, and search unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart glasses 214 and receives inquiries from visually impaired persons using speech recognition technology. The analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes camera images using image recognition technology. The provision unit is implemented by the control unit 46A of the smart glasses 214 and provides the analyzed information in voice using speech synthesis technology. The storage unit is implemented by the specific processing unit 290 of the data processing unit 12 and stores the provided information in the database 24. The search unit is implemented by the specific processing unit 290 of the data processing unit 12 and searches the stored information and provides it to visually impaired persons. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0139] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0140] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0141] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0142] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0143] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0144] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0145] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0146] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0147] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0148] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0149] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0150] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0151] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0152] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0153] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0154] Each of the multiple elements described above, including the reception unit, analysis unit, provision unit, storage unit, and search unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the headset terminal 314 and receives inquiries from visually impaired persons using speech recognition technology. The analysis unit is implemented by the specific processing unit 290 of the data processing unit 12 and analyzes camera images using image recognition technology. The provision unit is implemented by the control unit 46A of the headset terminal 314 and provides the analyzed information in voice using speech synthesis technology. The storage unit is implemented by the specific processing unit 290 of the data processing unit 12 and stores the provided information in the database 24. The search unit is implemented by the specific processing unit 290 of the data processing unit 12 and searches the stored information and provides it to visually impaired persons. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0155] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0156] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0157] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0158] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0159] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0160] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0161] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0162] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0163] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0164] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0165] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0166] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0167] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0168] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0169] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0170] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0171] Each of the multiple elements described above, including the reception unit, analysis unit, provision unit, storage unit, and search unit, is implemented by, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the robot 414 and receives inquiries from visually impaired persons using speech recognition technology. The analysis unit is implemented by, for example, the specific processing unit 290 of the data processing unit 12 and analyzes camera images using image recognition technology. The provision unit is implemented by, for example, the control unit 46A of the robot 414 and provides the analyzed information in voice using speech synthesis technology. The storage unit is implemented by, for example, the specific processing unit 290 of the data processing unit 12 and stores the provided information in the database 24. The search unit is implemented by, for example, the specific processing unit 290 of the data processing unit 12 and searches the stored information and provides it to visually impaired persons. The correspondence between each unit and the device or control unit is not limited to the example described above and can be changed in various ways.
[0172] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0173] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0174] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0175] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0176] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0177] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0178] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0179] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0180] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0181] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0182] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0183] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0184] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0185] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0186] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0187] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0188] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0189] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0190] (Note 1) A reception desk that accepts inquiries from visually impaired people, An analysis unit analyzes the camera footage based on the questions received by the reception unit, A providing unit that provides the information analyzed by the aforementioned analysis unit in audio format, A storage unit that stores the information provided by the aforementioned provisioning unit in a database, The system includes a search unit that searches the information stored in the storage unit and provides it to visually impaired persons. A system characterized by the following features. (Note 2) The aforementioned reception unit is Provides an interface for visually impaired individuals to ask questions. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, Analyze camera footage to identify information useful for visually impaired individuals. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned supply unit is, The analyzed information is provided in audio format. The system described in Appendix 1, characterized by the features described herein. (Note 5) The storage unit is Store the provided information in a database. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned search unit, Search the database and provide useful information for visually impaired people. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is The system estimates the user's emotions and adjusts how questions are received based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is Analyze the user's past inquiry history and select the optimal reception method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When a question is submitted, filtering is performed based on the user's current situation and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is It estimates the user's emotions and determines the priority of questions to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When receiving inquiries, the system prioritizes accepting inquiries that are highly relevant, taking into account the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is When receiving a question, the system analyzes the user's social media activity and accepts relevant questions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the representation of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, the level of detail is adjusted based on the importance of the video footage. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of the video. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, It estimates the user's emotions and adjusts the length of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when the video was filmed. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of the video footage. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned supply unit is, We estimate the user's emotions and adjust the way we present the content based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned supply unit is, When providing information, adjust the level of detail based on its importance. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned supply unit is, When providing information, different delivery algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned supply unit is, It estimates the user's emotions and adjusts the length of the service based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned supply unit is, When providing information, the priority of provision will be determined based on when the information was acquired. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned supply unit is, When providing information, the order of provision will be adjusted based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 25) The storage unit is The system estimates the user's emotions and selects stored data based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The storage unit is During data storage, the storage algorithm is optimized by referring to past stored data. The system described in Appendix 1, characterized by the features described herein. (Note 27) The storage unit is It estimates the user's emotions and adjusts the frequency of accumulation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The storage unit is During data storage, the stored data is weighted based on when the information was acquired. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned search unit, It estimates the user's sentiment and adjusts the search method based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned search unit, During a search, the system selects the optimal search method by referring to the user's past search history. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned search unit, When searching, customize the search method based on the user's current situation. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned search unit, It estimates the user's sentiment and determines search priorities based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned search unit, When performing a search, the system selects the optimal search method by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned search unit, When you search, we analyze your social media activity and suggest search methods. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0191] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception desk that accepts inquiries from visually impaired people, An analysis unit analyzes the camera footage based on the questions received by the reception unit, A providing unit that provides the information analyzed by the aforementioned analysis unit in audio format, A storage unit that stores the information provided by the aforementioned provisioning unit in a database, The system includes a search unit that searches the information stored in the storage unit and provides it to visually impaired persons. A system characterized by the following features.
2. The aforementioned reception unit is Provides an interface for visually impaired individuals to ask questions. The system according to feature 1.
3. The aforementioned analysis unit, Analyze camera footage to identify information useful for visually impaired individuals. The system according to feature 1.
4. The aforementioned supply unit is, The analyzed information is provided in audio format. The system according to feature 1.
5. The storage unit is Store the provided information in a database. The system according to feature 1.
6. The aforementioned search unit, Search the database and provide useful information for visually impaired people. The system according to feature 1.
7. The aforementioned reception unit is The system estimates the user's emotions and adjusts how questions are received based on those estimated emotions. The system according to feature 1.
8. The aforementioned reception unit is Analyze the user's past inquiry history and select the optimal reception method. The system according to feature 1.
9. The aforementioned reception unit is When a question is submitted, filtering is performed based on the user's current situation and areas of interest. The system according to feature 1.
10. The aforementioned reception unit is It estimates the user's emotions and determines the priority of questions to accept based on the estimated user emotions. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A