system

A system with a generating AI agent addresses the challenge of visually impaired users by enabling efficient text reading and summarization based on voice triggers, enhancing their ability to utilize smartphones and personal computers effectively.

JP2026073590APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Visually impaired individuals face challenges in accurately understanding information on device screens due to scattered or varied notation methods, making it difficult to quickly grasp necessary information from complex content on smartphones and personal computers.

Method used

A system utilizing a generating AI agent that performs text reading, transcription, and screen information summarization based on voice triggers, employing generative AI and natural language processing to analyze, summarize, and read aloud information according to user requests.

Benefits of technology

Enables visually impaired individuals to efficiently use smartphones and personal computers by providing quick and accurate access to relevant information, improving their quality of life and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026073590000001_ABST
    Figure 2026073590000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to enable visually impaired individuals to understand the information on the device screen more accurately. [Solution] The system according to the embodiment comprises a reception unit, an analysis unit, a summarization unit, and a reading unit. The reception unit receives an audio trigger. The analysis unit analyzes the information based on the audio trigger received by the reception unit. The summarization unit summarizes the information analyzed by the analysis unit. The reading unit reads aloud the information summarized by the summarization unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , , ,

[0005] , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0007] The system according to this embodiment can enable visually impaired individuals to understand the information on the device screen more accurately. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The visual impairment support system according to an embodiment of the present invention is a system that enables visually impaired individuals to use smartphones and personal computers more efficiently. This system uses a generating AI agent to perform text reading and transcription, screen information summarization, and reading information according to the user's requests, based on the user's voice trigger. This system allows visually impaired individuals to utilize smartphones and personal computers more effectively than before, improving their quality of life and expanding their range of work. For example, one challenge for visually impaired individuals using smartphones and personal computers is that, because the text reader reads all the information, it is difficult to accurately understand the information on the screen if the text is scattered or uses various notation methods. For example, if the content of a web page or email is complex, it is difficult for visually impaired individuals to quickly grasp the necessary information. In this system, the user gives instructions to the generating AI agent using a voice trigger. For example, if the user gives a voice instruction such as "Tell me the summary of this page," the generating AI agent analyzes the content of the page, summarizes the important information, and reads it aloud. Also, when specific information needs to be transcribed, the generating AI agent automatically transcribes it based on the voice instruction. Furthermore, it is also possible to read information according to the user's requests. For example, if the user gives an instruction such as "Read the following email," the generating AI agent will read the content of the email aloud. In this way, visually impaired individuals can quickly and accurately obtain the information they need. This system allows visually impaired individuals to use smartphones and computers more efficiently, improving their quality of life. It also broadens their job opportunities, enabling them to contribute to both themselves and society as a whole. For example, when a visually impaired person shops online, a generating AI agent can summarize and read aloud product descriptions, allowing them to select products more smoothly. Furthermore, quickly understanding the content of work emails can improve work efficiency. In short, this support system for the visually impaired enables them to use smartphones and computers more efficiently.

[0029] The visual impairment assistance system according to this embodiment comprises a reception unit, an analysis unit, a summarization unit, and a reading unit. The reception unit receives voice triggers. Voice triggers include, but are not limited to, specific keywords or voice commands. The reception unit receives voice triggers using, for example, speech recognition technology. The reception unit can also transmit the user's voice triggers to the analysis unit. The analysis unit analyzes information based on the voice triggers received by the reception unit. The analysis unit analyzes the voice triggers using, for example, generative AI and extracts the necessary information. The analysis unit can also analyze the voice triggers using, for example, natural language processing technology. The analysis unit can also transcribe specific information based on the voice triggers. The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit summarizes the information using, for example, generative AI. The summarization unit can also summarize based on, for example, the length of the text or the importance of the information being summarized. The summarization unit can also extract and summarize important parts of the information using generative AI. The reading unit reads aloud the information summarized by the summarization unit. The reading unit reads information using, for example, a generative AI. The reading unit can also read information using, for example, speech synthesis technology. Furthermore, the reading unit can read information tailored to the user's requests. For example, the reading unit reads specific information based on the user's voice instructions. This enables the visually impaired to efficiently use smartphones and personal computers as part of the support system. Some or all of the above-described processes in the reception unit, analysis unit, summarization unit, and reading unit may be performed using, for example, AI, or without AI. For example, the reception unit can receive voice triggers using an AI model that receives voice triggers. The analysis unit can analyze information using an AI model that analyzes voice triggers. The summarization unit can summarize information using an AI model that summarizes information. The reading unit can read information using an AI model that reads information aloud.

[0030] The reception unit receives voice triggers. Voice triggers include, but are not limited to, specific keywords or voice commands. The reception unit receives voice triggers using, for example, speech recognition technology. Specifically, deep learning-based speech recognition models are often used as speech recognition technology. By learning from large amounts of voice data, this model can recognize various accents and pronunciation differences with high accuracy. For example, if a user utters a voice command such as "Read the news aloud," the reception unit analyzes this voice in real time and recognizes it as an appropriate trigger. In addition, noise cancellation technology and voice enhancement technology may be used in combination to improve the accuracy of voice trigger recognition. This allows for accurate reception of voice triggers even in environments with a lot of ambient noise and background sound. Furthermore, the reception unit can also send the user's voice triggers to the analysis unit. The data sent includes not only the voice data itself but also the data converted into text by speech recognition. This allows the analysis unit to quickly and accurately analyze the content of the voice triggers. The reception unit can not only receive the user's voice triggers but also provide more personalized services by referring to the user's profile information and past usage history. For example, it's possible to configure the system to prioritize the recognition of commands and keywords frequently used by specific users. This allows users to use the system more smoothly.

[0031] The analysis unit analyzes information based on voice triggers received by the reception unit. For example, the analysis unit uses generative AI to analyze voice triggers and extract necessary information. Specifically, the generative AI utilizes natural language processing technology to understand the content of the voice trigger and extract appropriate information. For example, if a user makes a voice trigger such as "Tell me today's weather," the analysis unit analyzes this voice trigger and generates a query to obtain weather information. The generative AI performs grammatical and semantic analysis to understand the context and intent of the voice trigger. This allows it to accurately grasp the user's intent and provide appropriate information. The analysis unit can also transcribe specific information based on the voice trigger. For example, if a user makes a voice trigger such as "Take notes," the analysis unit analyzes this voice trigger and saves the content of the notes as text. Furthermore, the analysis unit can perform more accurate analysis by referring to past data and user profile information. For example, by referring to what kind of voice triggers the user has made in the past, it can more accurately grasp the intent of the current voice trigger. The analysis unit sends the analysis results of the voice trigger to the summarization unit. The transmitted data includes not only the analysis results themselves, but also the metadata and supplementary information used in the analysis. This allows the summarization unit to efficiently summarize the information based on the analysis results.

[0032] The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit uses, for example, generative AI to summarize the information. Specifically, the generative AI utilizes natural language generation technology to summarize the analyzed information concisely and clearly. For example, if a user issues a voice trigger such as "Tell me the latest news," the summarization unit summarizes the news article acquired by the analysis unit, extracting only the important points and compiling them into short sentences. The generative AI can perform summarization based on the length of the sentence and the importance of the information being summarized. For example, when summarizing a long news article, it prioritizes extracting important paragraphs and keywords and compiling them into short sentences. The summarization unit can also use generative AI to extract and summarize the most important parts of the information. This allows users to efficiently obtain only the information they need. Furthermore, the summarization unit can provide more personalized summaries by referring to the user's profile information and past usage history. For example, it can be set to prioritize summarizing topics and keywords of interest to a particular user. This allows users to quickly obtain more relevant information. The summarization unit then sends the summarized information to the text-to-speech unit. The data transmitted includes not only the summary result itself, but also the metadata and supplementary information used in the summary. This allows the text-to-speech unit to efficiently read the information based on the summary result.

[0033] The reading unit reads aloud the information summarized by the summarizing unit. The reading unit uses, for example, generative AI to read the information. Specifically, the generative AI utilizes speech synthesis technology to read the summarized information in a natural voice. Text-to-speech (TTS) technology is often used as the speech synthesis technology. This technology converts text data into speech data, and deep learning models are used to achieve natural intonation and pronunciation. For example, if a user issues a voice trigger such as "Read the latest news," the reading unit reads the news summary received from the summarizing unit in a natural voice. The reading unit can also read information tailored to the user's requests. For example, if a user issues a voice command such as "Read only specific news," the reading unit will read only that specific news based on that command. Furthermore, the reading unit can provide more personalized readings by referring to the user's profile information and past usage history. For example, it can be set to read information in a voice tone and speed preferred by a particular user. This allows users to acquire information more comfortably. The text-to-speech unit generates audio data in real time during the reading process and provides it to the user. This allows the user to always obtain the latest information in real time.

[0034] The analysis unit can transcribe specific information. For example, the analysis unit can transcribe specific information such as the user's name, address, and telephone number. The analysis unit can also automatically transcribe specific information based on voice triggers, for example. Furthermore, the analysis unit can transcribe specific information based on the user's voice instructions. This improves the work efficiency of visually impaired individuals by automatically transcribing specific information. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can transcribe specific information using an AI model that transcribes specific information.

[0035] The text-to-speech unit can read aloud information that aligns with the user's requests. For example, the text-to-speech unit can read specific information based on the user's voice instructions. The text-to-speech unit can also read specific information based on the user's settings or past usage history. Furthermore, the text-to-speech unit can read information according to the user's requests. This improves convenience for visually impaired individuals by providing information that meets the user's needs. Some or all of the above-described processes in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that reads information that aligns with the user's requests.

[0036] The reception unit can analyze the user's past voice trigger history and select the optimal reception method. For example, the reception unit prioritizes receiving voice triggers that the user has frequently used in the past. The reception unit can also predict triggers to be used during specific time periods based on the user's past voice trigger history and receive them accordingly. Furthermore, the reception unit can analyze the user's past voice trigger history and propose the most efficient reception method. This improves user convenience by selecting the optimal reception method based on past history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can select the optimal reception method using an AI model that analyzes the user's past voice trigger history.

[0037] The reception unit can filter out noise by filtering the user's current ambient sounds when it receives a voice trigger. For example, if the user is in a noisy environment, the reception unit can filter out ambient sounds to accurately receive the voice trigger. For example, if the user is in a quiet environment, the reception unit can also filter out ambient sounds minimally to receive the voice trigger. Furthermore, if the user is moving, the reception unit can filter out ambient sounds in real time to receive the voice trigger. This improves the accuracy of voice trigger reception by filtering out ambient sounds. Some or all of the processing described above in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can remove noise using an AI model that filters ambient sounds.

[0038] The reception unit can prioritize receiving voice triggers based on the user's geographical location information. For example, if the user is in a specific location, the reception unit will prioritize receiving voice triggers related to that location. If the user is on the move, the reception unit can also prioritize receiving the most appropriate voice trigger based on the user's current location. Furthermore, if the user is at home, the reception unit can prioritize receiving voice triggers related to home. This improves user convenience by prioritizing voice triggers based on geographical location information. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can receive voice triggers using an AI model that prioritizes receiving highly relevant triggers while considering the user's geographical location information.

[0039] The reception unit can analyze the user's social media activity and receive relevant triggers when it receives a voice trigger. For example, the reception unit can receive voice triggers based on keywords that the user frequently uses on social media. The reception unit can also prioritize receiving voice triggers related to specific topics from the user's social media activity. Furthermore, the reception unit can analyze the content of the user's social media posts and receive relevant voice triggers. This improves user convenience by receiving voice triggers based on social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can receive relevant triggers using an AI model that analyzes the user's social media activity.

[0040] The analysis unit can adjust the level of detail of the analysis based on the importance of the information during the analysis. For example, the analysis unit can analyze important information in detail and provide it to the user. For example, the analysis unit can also analyze general information simply and provide it to the user. Furthermore, the analysis unit can omit the analysis of unnecessary information and not provide it to the user. In this way, by adjusting the level of detail of the analysis based on the importance of the information, important information can be provided in detail. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that adjusts the level of detail of the analysis based on the importance of the information.

[0041] The analysis unit can apply different analysis algorithms depending on the category of information during analysis. For example, the analysis unit can apply a natural language processing algorithm to text information. For example, the analysis unit can also apply an image analysis algorithm to image information. Furthermore, the analysis unit can apply a speech analysis algorithm to speech information. By applying an appropriate analysis algorithm according to the category of information, the accuracy of the analysis is improved. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that applies different analysis algorithms depending on the category of information.

[0042] The analysis unit can determine the priority of analysis based on the timing of information submission during the analysis process. For example, the analysis unit may prioritize the analysis of the latest information and provide it to the user. The analysis unit may also postpone the analysis of older information. Furthermore, the analysis unit may prioritize the analysis of information where the timing of submission is important and provide it to the user. This allows for the rapid provision of the latest information by determining the priority of analysis based on the timing of information submission. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that determines the priority of analysis based on the timing of information submission.

[0043] The analysis unit can adjust the order of analysis based on the relevance of the information during the analysis process. For example, the analysis unit can prioritize the analysis of highly relevant information and provide it to the user. For example, the analysis unit can also postpone the analysis of less relevant information. Furthermore, the analysis unit can analyze the relevance of the information and perform the analysis in the optimal order. This allows for the priority provision of highly relevant information by adjusting the order of analysis based on the relevance of the information. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that adjusts the order of analysis based on the relevance of the information.

[0044] The summarization unit can adjust the level of detail in the summary based on the importance of the information during summary generation. For example, the summarization unit can summarize important information in detail and provide it to the user. For example, the summarization unit can also summarize general information simply and provide it to the user. Furthermore, the summarization unit can omit summarizing unnecessary information and not provide it to the user. In this way, by adjusting the level of detail in the summary based on the importance of the information, important information can be provided in detail. Some or all of the above processing in the summarization unit may be performed using AI, for example, or not using AI. For example, the summarization unit can summarize information using an AI model that adjusts the level of detail in the summary based on the importance of the information.

[0045] The summarization unit can apply different summarization algorithms depending on the information category when generating summaries. For example, the summarization unit can apply a natural language processing algorithm to text information. For example, it can also apply an image analysis algorithm to image information. Furthermore, it can apply a speech analysis algorithm to audio information. This improves summarization accuracy by applying an appropriate summarization algorithm according to the information category. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI. For example, the summarization unit can summarize information using an AI model that applies different summarization algorithms depending on the information category.

[0046] The summarization unit can determine the priority of summaries based on the information submission timing when generating summaries. For example, the summarization unit can prioritize summarizing the most recent information and provide it to the user. For example, the summarization unit can also postpone summarizing older information. Furthermore, the summarization unit can prioritize summarizing information where the submission timing is important and provide it to the user. This allows for the rapid provision of the latest information by determining the priority of summaries based on the information submission timing. Some or all of the above processing in the summarization unit may be performed using AI, for example, or not using AI. For example, the summarization unit can summarize information using an AI model that determines the priority of summaries based on the information submission timing.

[0047] The summarization unit can adjust the order of summaries based on the relevance of the information during summarization. For example, the summarization unit can prioritize summarizing highly relevant information and provide it to the user. For example, the summarization unit can also postpone summarizing less relevant information. Furthermore, the summarization unit can analyze the relevance of the information and summarize it in the optimal order. This allows for the priority provision of highly relevant information by adjusting the order of summaries based on the relevance of the information. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI. For example, the summarization unit can summarize information using an AI model that adjusts the order of summaries based on the relevance of the information.

[0048] The text-to-speech unit can adjust the level of detail in its reading based on the importance of the information. For example, it can read important information in detail. For example, it can read general information simply. It can also omit reading unnecessary information. In this way, by adjusting the level of detail in the reading based on the importance of the information, it can provide important information in detail. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that adjusts the level of detail in the reading based on the importance of the information.

[0049] The text-to-speech unit can apply different reading algorithms depending on the category of information during reading. For example, the text-to-speech unit can apply a natural language processing algorithm to text information. For example, the text-to-speech unit can also apply an image analysis algorithm to image information. Furthermore, the text-to-speech unit can apply a speech analysis algorithm to speech information. This improves reading accuracy by applying an appropriate reading algorithm according to the category of information. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that applies different reading algorithms depending on the category of information.

[0050] The reading unit can determine the reading priority based on the timing of information submission when reading aloud. For example, the reading unit may prioritize reading the most recent information. It may also postpone reading older information. Furthermore, it may prioritize reading information where the timing of submission is important. This allows for the rapid provision of the latest information by determining the reading priority based on the timing of information submission. Some or all of the above processing in the reading unit may be performed using AI, for example, or without AI. For example, the reading unit can read information using an AI model that determines the reading priority based on the timing of information submission.

[0051] The reading unit can adjust the reading order based on the relevance of the information during reading. For example, the reading unit can prioritize reading highly relevant information. For example, the reading unit can also postpone reading less relevant information. Furthermore, the reading unit can analyze the relevance of the information and read it in the optimal order. This allows for the prioritization of highly relevant information by adjusting the reading order based on the relevance of the information. Some or all of the above processing in the reading unit may be performed using AI, for example, or without AI. For example, the reading unit can read information using an AI model that adjusts the reading order based on the relevance of the information.

[0052] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0053] The visual impairment assistance system may further include a behavioral analysis unit that analyzes the user's past behavioral history and selects the optimal method of information delivery. For example, the behavioral analysis unit may prioritize providing information that the user has frequently accessed in the past. The behavioral analysis unit may also predict and provide information needed at a specific time period based on the user's past behavioral history. Furthermore, the behavioral analysis unit may analyze the user's past behavioral history and propose the most efficient method of information delivery. This improves user convenience by selecting the optimal method of information delivery based on past history. Some or all of the above processing in the behavioral analysis unit may be performed using AI, for example, or without AI. For example, the behavioral analysis unit may select the optimal method of information delivery using an AI model that analyzes the user's past behavioral history.

[0054] The visually impaired assistance system may further include an ambient sound analysis unit that analyzes the user's current ambient sounds and selects the optimal audio output method. For example, the ambient sound analysis unit may increase the volume of the audio output if the user is in a noisy place. It may also decrease the volume of the audio output if the user is in a quiet place. Furthermore, if the user is moving, the ambient sound analysis unit may adjust the audio output method according to the ambient sounds. This makes it possible to optimize the audio output by analyzing the ambient sounds. Some or all of the above processing in the ambient sound analysis unit may be performed using AI, for example, or without AI. For example, the ambient sound analysis unit may select the audio output method using an AI model that analyzes ambient sounds.

[0055] The visually impaired assistance system may further include a location information analysis unit that provides highly relevant information based on the user's geographical location. For example, if the user is in a specific location, the location information analysis unit may prioritize providing information related to that location. For example, if the user is on the move, the location information analysis unit may also provide optimal information based on the user's current location. Furthermore, if the user is at home, the location information analysis unit may prioritize providing information related to home. This improves user convenience by providing information based on geographical location. Some or all of the above processing in the location information analysis unit may be performed using AI, for example, or without AI. For example, the location information analysis unit may provide information using an AI model that provides highly relevant information considering the user's geographical location.

[0056] The visually impaired support system may further include a social media analysis unit that analyzes the user's social media activity and provides relevant information. The social media analysis unit may, for example, provide relevant information based on keywords that the user frequently uses on social media. The social media analysis unit may also, for example, prioritize providing information related to specific topics from the user's social media activity. Furthermore, the social media analysis unit may analyze the content of the user's social media posts and provide relevant information. This improves user convenience by providing information based on social media activity. Some or all of the above processing in the social media analysis unit may be performed using AI, for example, or without AI. For example, the social media analysis unit may provide relevant information using an AI model that analyzes the user's social media activity.

[0057] The following briefly describes the processing flow for example form 1.

[0058] Step 1: The reception unit receives voice triggers. Voice triggers include specific keywords or voice commands. The reception unit uses speech recognition technology to receive voice triggers and transmits the user's voice triggers to the analysis unit. Step 2: The analysis unit analyzes the information based on the voice triggers received by the reception unit. The analysis unit analyzes the voice triggers using generative AI and natural language processing technology and extracts the necessary information. Step 3: The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit uses a generative AI to summarize the information, performing the summary based on the length of the text and the importance of the information being summarized. Step 4: The reading unit reads aloud the information summarized by the summarizing unit. The reading unit uses generation AI and speech synthesis technology to read the information and provide information that meets the user's needs.

[0059] (Example of form 2) The visual impairment support system according to an embodiment of the present invention is a system that enables visually impaired individuals to use smartphones and personal computers more efficiently. This system uses a generating AI agent to perform text reading and transcription, screen information summarization, and reading information according to the user's requests, based on the user's voice trigger. This system allows visually impaired individuals to utilize smartphones and personal computers more effectively than before, improving their quality of life and expanding their range of work. For example, one challenge for visually impaired individuals using smartphones and personal computers is that, because the text reader reads all the information, it is difficult to accurately understand the information on the screen if the text is scattered or uses various notation methods. For example, if the content of a web page or email is complex, it is difficult for visually impaired individuals to quickly grasp the necessary information. In this system, the user gives instructions to the generating AI agent using a voice trigger. For example, if the user gives a voice instruction such as "Tell me the summary of this page," the generating AI agent analyzes the content of the page, summarizes the important information, and reads it aloud. Also, when specific information needs to be transcribed, the generating AI agent automatically transcribes it based on the voice instruction. Furthermore, it is also possible to read information according to the user's requests. For example, if the user gives an instruction such as "Read the following email," the generating AI agent will read the content of the email aloud. In this way, visually impaired individuals can quickly and accurately obtain the information they need. This system allows visually impaired individuals to use smartphones and computers more efficiently, improving their quality of life. It also broadens their job opportunities, enabling them to contribute to both themselves and society as a whole. For example, when a visually impaired person shops online, a generating AI agent can summarize and read aloud product descriptions, allowing them to select products more smoothly. Furthermore, quickly understanding the content of work emails can improve work efficiency. In short, this support system for the visually impaired enables them to use smartphones and computers more efficiently.

[0060] The visual impairment assistance system according to this embodiment comprises a reception unit, an analysis unit, a summarization unit, and a reading unit. The reception unit receives voice triggers. Voice triggers include, but are not limited to, specific keywords or voice commands. The reception unit receives voice triggers using, for example, speech recognition technology. The reception unit can also transmit the user's voice triggers to the analysis unit. The analysis unit analyzes information based on the voice triggers received by the reception unit. The analysis unit analyzes the voice triggers using, for example, generative AI and extracts the necessary information. The analysis unit can also analyze the voice triggers using, for example, natural language processing technology. The analysis unit can also transcribe specific information based on the voice triggers. The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit summarizes the information using, for example, generative AI. The summarization unit can also summarize based on, for example, the length of the text or the importance of the information being summarized. The summarization unit can also extract and summarize important parts of the information using generative AI. The reading unit reads aloud the information summarized by the summarization unit. The reading unit reads information using, for example, a generative AI. The reading unit can also read information using, for example, speech synthesis technology. Furthermore, the reading unit can read information tailored to the user's requests. For example, the reading unit reads specific information based on the user's voice instructions. This enables the visually impaired to efficiently use smartphones and personal computers as part of the support system. Some or all of the above-described processes in the reception unit, analysis unit, summarization unit, and reading unit may be performed using, for example, AI, or without AI. For example, the reception unit can receive voice triggers using an AI model that receives voice triggers. The analysis unit can analyze information using an AI model that analyzes voice triggers. The summarization unit can summarize information using an AI model that summarizes information. The reading unit can read information using an AI model that reads information aloud.

[0061] The reception unit receives voice triggers. Voice triggers include, but are not limited to, specific keywords or voice commands. The reception unit receives voice triggers using, for example, speech recognition technology. Specifically, deep learning-based speech recognition models are often used as speech recognition technology. By learning from large amounts of voice data, this model can recognize various accents and pronunciation differences with high accuracy. For example, if a user utters a voice command such as "Read the news aloud," the reception unit analyzes this voice in real time and recognizes it as an appropriate trigger. In addition, noise cancellation technology and voice enhancement technology may be used in combination to improve the accuracy of voice trigger recognition. This allows for accurate reception of voice triggers even in environments with a lot of ambient noise and background sound. Furthermore, the reception unit can also send the user's voice triggers to the analysis unit. The data sent includes not only the voice data itself but also the data converted into text by speech recognition. This allows the analysis unit to quickly and accurately analyze the content of the voice triggers. The reception unit can not only receive the user's voice triggers but also provide more personalized services by referring to the user's profile information and past usage history. For example, it's possible to configure the system to prioritize the recognition of commands and keywords frequently used by specific users. This allows users to use the system more smoothly.

[0062] The analysis unit analyzes information based on voice triggers received by the reception unit. For example, the analysis unit uses generative AI to analyze voice triggers and extract necessary information. Specifically, the generative AI utilizes natural language processing technology to understand the content of the voice trigger and extract appropriate information. For example, if a user makes a voice trigger such as "Tell me today's weather," the analysis unit analyzes this voice trigger and generates a query to obtain weather information. The generative AI performs grammatical and semantic analysis to understand the context and intent of the voice trigger. This allows it to accurately grasp the user's intent and provide appropriate information. The analysis unit can also transcribe specific information based on the voice trigger. For example, if a user makes a voice trigger such as "Take notes," the analysis unit analyzes this voice trigger and saves the content of the notes as text. Furthermore, the analysis unit can perform more accurate analysis by referring to past data and user profile information. For example, by referring to what kind of voice triggers the user has made in the past, it can more accurately grasp the intent of the current voice trigger. The analysis unit sends the analysis results of the voice trigger to the summarization unit. The transmitted data includes not only the analysis results themselves, but also the metadata and supplementary information used in the analysis. This allows the summarization unit to efficiently summarize the information based on the analysis results.

[0063] The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit uses, for example, generative AI to summarize the information. Specifically, the generative AI utilizes natural language generation technology to summarize the analyzed information concisely and clearly. For example, if a user issues a voice trigger such as "Tell me the latest news," the summarization unit summarizes the news article acquired by the analysis unit, extracting only the important points and compiling them into short sentences. The generative AI can perform summarization based on the length of the sentence and the importance of the information being summarized. For example, when summarizing a long news article, it prioritizes extracting important paragraphs and keywords and compiling them into short sentences. The summarization unit can also use generative AI to extract and summarize the most important parts of the information. This allows users to efficiently obtain only the information they need. Furthermore, the summarization unit can provide more personalized summaries by referring to the user's profile information and past usage history. For example, it can be set to prioritize summarizing topics and keywords of interest to a particular user. This allows users to quickly obtain more relevant information. The summarization unit then sends the summarized information to the text-to-speech unit. The data transmitted includes not only the summary result itself, but also the metadata and supplementary information used in the summary. This allows the text-to-speech unit to efficiently read the information based on the summary result.

[0064] The reading unit reads aloud the information summarized by the summarizing unit. The reading unit uses, for example, generative AI to read the information. Specifically, the generative AI utilizes speech synthesis technology to read the summarized information in a natural voice. Text-to-speech (TTS) technology is often used as the speech synthesis technology. This technology converts text data into speech data, and deep learning models are used to achieve natural intonation and pronunciation. For example, if a user issues a voice trigger such as "Read the latest news," the reading unit reads the news summary received from the summarizing unit in a natural voice. The reading unit can also read information tailored to the user's requests. For example, if a user issues a voice command such as "Read only specific news," the reading unit will read only that specific news based on that command. Furthermore, the reading unit can provide more personalized readings by referring to the user's profile information and past usage history. For example, it can be set to read information in a voice tone and speed preferred by a particular user. This allows users to acquire information more comfortably. The text-to-speech unit generates audio data in real time during the reading process and provides it to the user. This allows the user to always obtain the latest information in real time.

[0065] The analysis unit can transcribe specific information. For example, the analysis unit can transcribe specific information such as the user's name, address, and telephone number. The analysis unit can also automatically transcribe specific information based on voice triggers, for example. Furthermore, the analysis unit can transcribe specific information based on the user's voice instructions. This improves the work efficiency of visually impaired individuals by automatically transcribing specific information. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can transcribe specific information using an AI model that transcribes specific information.

[0066] The text-to-speech unit can read aloud information that aligns with the user's requests. For example, the text-to-speech unit can read specific information based on the user's voice instructions. The text-to-speech unit can also read specific information based on the user's settings or past usage history. Furthermore, the text-to-speech unit can read information according to the user's requests. This improves convenience for visually impaired individuals by providing information that meets the user's needs. Some or all of the above-described processes in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that reads information that aligns with the user's requests.

[0067] The reception unit can estimate the user's emotions and adjust the timing of voice trigger reception based on the estimated emotions. For example, if the user is stressed, the reception unit can quickly receive the voice trigger to reduce the user's burden. For example, if the user is relaxed, the reception unit can slightly delay receiving the voice trigger to facilitate a natural conversation. Also, if the user is in a hurry, the reception unit can immediately receive the voice trigger to enable a quick response. This allows for a more natural conversation by adjusting the timing of voice trigger reception according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI, for example, or not using AI. For example, the reception unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the timing of voice trigger reception.

[0068] The reception unit can analyze the user's past voice trigger history and select the optimal reception method. For example, the reception unit prioritizes receiving voice triggers that the user has frequently used in the past. The reception unit can also predict triggers to be used during specific time periods based on the user's past voice trigger history and receive them accordingly. Furthermore, the reception unit can analyze the user's past voice trigger history and propose the most efficient reception method. This improves user convenience by selecting the optimal reception method based on past history. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can select the optimal reception method using an AI model that analyzes the user's past voice trigger history.

[0069] The reception unit can filter out noise by filtering the user's current ambient sounds when it receives a voice trigger. For example, if the user is in a noisy environment, the reception unit can filter out ambient sounds to accurately receive the voice trigger. For example, if the user is in a quiet environment, the reception unit can also filter out ambient sounds minimally to receive the voice trigger. Furthermore, if the user is moving, the reception unit can filter out ambient sounds in real time to receive the voice trigger. This improves the accuracy of voice trigger reception by filtering out ambient sounds. Some or all of the processing described above in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can remove noise using an AI model that filters ambient sounds.

[0070] The reception unit can estimate the user's emotions and determine the priority of voice triggers to receive based on the estimated emotions. For example, if the user is nervous, the reception unit will prioritize important voice triggers. If the user is relaxed, the reception unit may also prioritize all voice triggers equally. If the user is in a hurry, the reception unit may also prioritize voice triggers with high urgency. This allows for the rapid acquisition of important information by prioritizing voice triggers according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not. For example, the reception unit can estimate the user's emotions using an AI model that estimates user emotions and determine the priority of voice triggers.

[0071] The reception unit can prioritize receiving voice triggers based on the user's geographical location information. For example, if the user is in a specific location, the reception unit will prioritize receiving voice triggers related to that location. If the user is on the move, the reception unit can also prioritize receiving the most appropriate voice trigger based on the user's current location. Furthermore, if the user is at home, the reception unit can prioritize receiving voice triggers related to home. This improves user convenience by prioritizing voice triggers based on geographical location information. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can receive voice triggers using an AI model that prioritizes receiving highly relevant triggers while considering the user's geographical location information.

[0072] The reception unit can analyze the user's social media activity and receive relevant triggers when it receives a voice trigger. For example, the reception unit can receive voice triggers based on keywords that the user frequently uses on social media. The reception unit can also prioritize receiving voice triggers related to specific topics from the user's social media activity. Furthermore, the reception unit can analyze the content of the user's social media posts and receive relevant voice triggers. This improves user convenience by receiving voice triggers based on social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can receive relevant triggers using an AI model that analyzes the user's social media activity.

[0073] The analysis unit can estimate the user's emotions and adjust the information analysis method based on the estimated user emotions. For example, if the user is relaxed, the analysis unit can perform detailed information analysis. If the user is in a hurry, the analysis unit can also perform concise information analysis. Furthermore, if the user is excited, the analysis unit can perform visually stimulating information analysis. By adjusting the information analysis method according to the user's emotions, more appropriate information can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the information analysis method accordingly.

[0074] The analysis unit can adjust the level of detail of the analysis based on the importance of the information during the analysis. For example, the analysis unit can analyze important information in detail and provide it to the user. For example, the analysis unit can also analyze general information simply and provide it to the user. Furthermore, the analysis unit can omit the analysis of unnecessary information and not provide it to the user. In this way, by adjusting the level of detail of the analysis based on the importance of the information, important information can be provided in detail. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that adjusts the level of detail of the analysis based on the importance of the information.

[0075] The analysis unit can apply different analysis algorithms depending on the category of information during analysis. For example, the analysis unit can apply a natural language processing algorithm to text information. For example, the analysis unit can also apply an image analysis algorithm to image information. Furthermore, the analysis unit can apply a speech analysis algorithm to speech information. By applying an appropriate analysis algorithm according to the category of information, the accuracy of the analysis is improved. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that applies different analysis algorithms depending on the category of information.

[0076] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the analysis unit can provide a simple and highly visible display method. For example, if the user is relaxed, the analysis unit can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the analysis unit can provide a concise display method. By adjusting the display method of the analysis results according to the user's emotions, visibility is improved. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the display method of the analysis results.

[0077] The analysis unit can determine the priority of analysis based on the timing of information submission during the analysis process. For example, the analysis unit may prioritize the analysis of the latest information and provide it to the user. The analysis unit may also postpone the analysis of older information. Furthermore, the analysis unit may prioritize the analysis of information where the timing of submission is important and provide it to the user. This allows for the rapid provision of the latest information by determining the priority of analysis based on the timing of information submission. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that determines the priority of analysis based on the timing of information submission.

[0078] The analysis unit can adjust the order of analysis based on the relevance of the information during the analysis process. For example, the analysis unit can prioritize the analysis of highly relevant information and provide it to the user. For example, the analysis unit can also postpone the analysis of less relevant information. Furthermore, the analysis unit can analyze the relevance of the information and perform the analysis in the optimal order. This allows for the priority provision of highly relevant information by adjusting the order of analysis based on the relevance of the information. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can analyze information using an AI model that adjusts the order of analysis based on the relevance of the information.

[0079] The summarization unit can estimate the user's emotions and adjust the way the summary is presented based on the estimated emotions. For example, if the user is relaxed, the summarization unit can provide a detailed summary. If the user is in a hurry, the summarization unit can also provide a concise summary that gets straight to the point. Furthermore, if the user is excited, the summarization unit can provide a visually stimulating summary. This allows for the provision of a more appropriate summary by adjusting the way the summary is presented according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI. For example, the summarization unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the way the summary is presented.

[0080] The summarization unit can adjust the level of detail in the summary based on the importance of the information during summary generation. For example, the summarization unit can summarize important information in detail and provide it to the user. For example, the summarization unit can also summarize general information simply and provide it to the user. Furthermore, the summarization unit can omit summarizing unnecessary information and not provide it to the user. In this way, by adjusting the level of detail in the summary based on the importance of the information, important information can be provided in detail. Some or all of the above processing in the summarization unit may be performed using AI, for example, or not using AI. For example, the summarization unit can summarize information using an AI model that adjusts the level of detail in the summary based on the importance of the information.

[0081] The summarization unit can apply different summarization algorithms depending on the information category when generating summaries. For example, the summarization unit can apply a natural language processing algorithm to text information. For example, it can also apply an image analysis algorithm to image information. Furthermore, it can apply a speech analysis algorithm to audio information. This improves summarization accuracy by applying an appropriate summarization algorithm according to the information category. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI. For example, the summarization unit can summarize information using an AI model that applies different summarization algorithms depending on the information category.

[0082] The summarization unit can estimate the user's emotions and adjust the length of the summary based on the estimated emotions. For example, if the user is in a hurry, the summarization unit can provide a short, concise summary. If the user is relaxed, the summarization unit can provide a longer summary with more detailed explanations. If the user is excited, the summarization unit can also provide a visually stimulating summary. By adjusting the length of the summary according to the user's emotions, a more appropriate summary can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the summarization unit may be performed using AI, for example, or not using AI. For example, the summarization unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the length of the summary accordingly.

[0083] The summarization unit can determine the priority of summaries based on the information submission timing when generating summaries. For example, the summarization unit can prioritize summarizing the most recent information and provide it to the user. For example, the summarization unit can also postpone summarizing older information. Furthermore, the summarization unit can prioritize summarizing information where the submission timing is important and provide it to the user. This allows for the rapid provision of the latest information by determining the priority of summaries based on the information submission timing. Some or all of the above processing in the summarization unit may be performed using AI, for example, or not using AI. For example, the summarization unit can summarize information using an AI model that determines the priority of summaries based on the information submission timing.

[0084] The summarization unit can adjust the order of summaries based on the relevance of the information during summarization. For example, the summarization unit can prioritize summarizing highly relevant information and provide it to the user. For example, the summarization unit can also postpone summarizing less relevant information. Furthermore, the summarization unit can analyze the relevance of the information and summarize it in the optimal order. This allows for the priority provision of highly relevant information by adjusting the order of summaries based on the relevance of the information. Some or all of the above processing in the summarization unit may be performed using AI, for example, or without AI. For example, the summarization unit can summarize information using an AI model that adjusts the order of summaries based on the relevance of the information.

[0085] The text-to-speech unit can estimate the user's emotions and adjust the tone and speed of the reading based on the estimated emotions. For example, if the user is relaxed, the text-to-speech unit will read in a relaxed tone and speed. If the user is in a hurry, the text-to-speech unit can also read in a fast tone and speed. Furthermore, if the user is excited, the text-to-speech unit can read in a visually stimulating tone. By adjusting the tone and speed of the reading according to the user's emotions, it becomes possible to provide more appropriate information. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or not using AI. For example, the text-to-speech unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the tone and speed of the reading accordingly.

[0086] The text-to-speech unit can adjust the level of detail in its reading based on the importance of the information. For example, it can read important information in detail. For example, it can read general information simply. It can also omit reading unnecessary information. In this way, by adjusting the level of detail in the reading based on the importance of the information, it can provide important information in detail. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that adjusts the level of detail in the reading based on the importance of the information.

[0087] The text-to-speech unit can apply different reading algorithms depending on the category of information during reading. For example, the text-to-speech unit can apply a natural language processing algorithm to text information. For example, the text-to-speech unit can also apply an image analysis algorithm to image information. Furthermore, the text-to-speech unit can apply a speech analysis algorithm to speech information. This improves reading accuracy by applying an appropriate reading algorithm according to the category of information. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can read information using an AI model that applies different reading algorithms depending on the category of information.

[0088] The text-to-speech unit can estimate the user's emotions and adjust the reading order based on the estimated emotions. For example, if the user is nervous, the text-to-speech unit can prioritize reading important information. If the user is relaxed, the text-to-speech unit can also read all information equally. Furthermore, if the user is in a hurry, the text-to-speech unit can prioritize reading information of high urgency. This allows for the rapid delivery of important information by adjusting the reading order according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the text-to-speech unit may be performed using AI, for example, or without AI. For example, the text-to-speech unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the reading order accordingly.

[0089] The reading unit can determine the reading priority based on the timing of information submission when reading aloud. For example, the reading unit may prioritize reading the most recent information. It may also postpone reading older information. Furthermore, it may prioritize reading information where the timing of submission is important. This allows for the rapid provision of the latest information by determining the reading priority based on the timing of information submission. Some or all of the above processing in the reading unit may be performed using AI, for example, or without AI. For example, the reading unit can read information using an AI model that determines the reading priority based on the timing of information submission.

[0090] The reading unit can adjust the reading order based on the relevance of the information during reading. For example, the reading unit can prioritize reading highly relevant information. For example, the reading unit can also postpone reading less relevant information. Furthermore, the reading unit can analyze the relevance of the information and read it in the optimal order. This allows for the prioritization of highly relevant information by adjusting the reading order based on the relevance of the information. Some or all of the above processing in the reading unit may be performed using AI, for example, or without AI. For example, the reading unit can read information using an AI model that adjusts the reading order based on the relevance of the information.

[0091] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0092] The visual impairment assistance system may further include a filtering unit that estimates the user's emotions and filters information based on the estimated emotions. For example, if the user is stressed, the filtering unit may reduce the amount of information and provide only the essential information. If the user is relaxed, the filtering unit may also provide detailed information. Furthermore, if the user is in a hurry, the filtering unit may provide quickly understandable summary information. This allows for more appropriate information to be provided by filtering information according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above processing in the filtering unit may be performed using AI, for example, or without AI. For example, the filtering unit may estimate the user's emotions using an AI model that estimates user emotions and then filter the information.

[0093] The visual impairment assistance system may further include a behavioral analysis unit that analyzes the user's past behavioral history and selects the optimal method of information delivery. For example, the behavioral analysis unit may prioritize providing information that the user has frequently accessed in the past. The behavioral analysis unit may also predict and provide information needed at a specific time period based on the user's past behavioral history. Furthermore, the behavioral analysis unit may analyze the user's past behavioral history and propose the most efficient method of information delivery. This improves user convenience by selecting the optimal method of information delivery based on past history. Some or all of the above processing in the behavioral analysis unit may be performed using AI, for example, or without AI. For example, the behavioral analysis unit may select the optimal method of information delivery using an AI model that analyzes the user's past behavioral history.

[0094] The visually impaired assistance system may further include an ambient sound analysis unit that analyzes the user's current ambient sounds and selects the optimal audio output method. For example, the ambient sound analysis unit may increase the volume of the audio output if the user is in a noisy place. It may also decrease the volume of the audio output if the user is in a quiet place. Furthermore, if the user is moving, the ambient sound analysis unit may adjust the audio output method according to the ambient sounds. This makes it possible to optimize the audio output by analyzing the ambient sounds. Some or all of the above processing in the ambient sound analysis unit may be performed using AI, for example, or without AI. For example, the ambient sound analysis unit may select the audio output method using an AI model that analyzes ambient sounds.

[0095] The visual impairment assistance system may further include a display adjustment unit that estimates the user's emotions and adjusts the way information is displayed based on the estimated emotions. For example, if the user is tense, the display adjustment unit may provide a simple and highly visible display method. If the user is relaxed, the display adjustment unit may also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the display adjustment unit may provide a display method that gets straight to the point. This improves visibility by adjusting the way information is displayed according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above processing in the display adjustment unit may be performed using AI, for example, or without AI. For example, the display adjustment unit may estimate the user's emotions using an AI model that estimates user emotions and adjust the way information is displayed.

[0096] The visually impaired assistance system may further include a location information analysis unit that provides highly relevant information based on the user's geographical location. For example, if the user is in a specific location, the location information analysis unit may prioritize providing information related to that location. For example, if the user is on the move, the location information analysis unit may also provide optimal information based on the user's current location. Furthermore, if the user is at home, the location information analysis unit may prioritize providing information related to home. This improves user convenience by providing information based on geographical location. Some or all of the above processing in the location information analysis unit may be performed using AI, for example, or without AI. For example, the location information analysis unit may provide information using an AI model that provides highly relevant information considering the user's geographical location.

[0097] The visual impairment assistance system may further include a priority determination unit that estimates the user's emotions and determines the priority of information based on the estimated emotions. For example, if the user is tense, the priority determination unit may prioritize providing important information. If the user is relaxed, the priority determination unit may also provide all information equally. Furthermore, if the user is in a hurry, the priority determination unit may prioritize providing information of high urgency. This allows for the rapid acquisition of important information by prioritizing information according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above-described processing in the priority determination unit may be performed using AI, for example, or without AI. For example, the priority determination unit may estimate the user's emotions using an AI model that estimates user emotions and then determine the priority of information.

[0098] The visually impaired support system may further include a social media analysis unit that analyzes the user's social media activity and provides relevant information. The social media analysis unit may, for example, provide relevant information based on keywords that the user frequently uses on social media. The social media analysis unit may also, for example, prioritize providing information related to specific topics from the user's social media activity. Furthermore, the social media analysis unit may analyze the content of the user's social media posts and provide relevant information. This improves user convenience by providing information based on social media activity. Some or all of the above processing in the social media analysis unit may be performed using AI, for example, or without AI. For example, the social media analysis unit may provide relevant information using an AI model that analyzes the user's social media activity.

[0099] The visual impairment assistance system may further include an analysis adjustment unit that estimates the user's emotions and adjusts the information analysis method based on the estimated emotions. For example, the analysis adjustment unit may perform detailed information analysis when the user is relaxed. For example, it may perform concise information analysis when the user is in a hurry. It may also perform visually stimulating information analysis when the user is excited. By adjusting the information analysis method according to the user's emotions, more appropriate information can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above processing in the analysis adjustment unit may be performed using AI, for example, or without AI. For example, the analysis adjustment unit may estimate the user's emotions using an AI model that estimates user emotions and adjust the information analysis method accordingly.

[0100] The visual impairment assistance system may further include a summary adjustment unit that estimates the user's emotions and adjusts the way the summary is presented based on the estimated emotions. For example, the summary adjustment unit may provide a detailed summary if the user is relaxed. For example, it may also provide a concise summary that gets to the point if the user is in a hurry. Furthermore, if the user is excited, the summary adjustment unit may provide a visually stimulating summary. This allows for the provision of a more appropriate summary by adjusting the way the summary is presented according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above-described processing in the summary adjustment unit may be performed using AI, for example, or without AI. For example, the summary adjustment unit may estimate the user's emotions using an AI model that estimates user emotions and adjust the way the summary is presented.

[0101] The visually impaired assistance system may further include a reading adjustment unit that estimates the user's emotions and adjusts the tone and speed of reading based on the estimated emotions. For example, if the user is relaxed, the reading adjustment unit will read in a relaxed tone and speed. If the user is in a hurry, the reading adjustment unit can also read in a rapid tone and speed. Furthermore, if the user is excited, the reading adjustment unit can read in a visually stimulating tone. This allows for the provision of more appropriate information by adjusting the tone and speed of reading according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Some or all of the above processing in the reading adjustment unit may be performed using AI, for example, or without AI. For example, the reading adjustment unit can estimate the user's emotions using an AI model that estimates user emotions and adjust the tone and speed of reading accordingly.

[0102] The following briefly describes the processing flow for example form 2.

[0103] Step 1: The reception unit receives voice triggers. Voice triggers include specific keywords or voice commands. The reception unit uses speech recognition technology to receive voice triggers and transmits the user's voice triggers to the analysis unit. Step 2: The analysis unit analyzes the information based on the voice triggers received by the reception unit. The analysis unit analyzes the voice triggers using generative AI and natural language processing technology and extracts the necessary information. Step 3: The summarization unit summarizes the information analyzed by the analysis unit. The summarization unit uses a generative AI to summarize the information, performing the summary based on the length of the text and the importance of the information being summarized. Step 4: The reading unit reads aloud the information summarized by the summarizing unit. The reading unit uses generation AI and speech synthesis technology to read the information and provide information that meets the user's needs.

[0104] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0105] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0106] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0107] Each of the multiple elements described above, including the reception unit, analysis unit, summarization unit, and reading unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit receives a voice trigger using the microphone 38B of the smart device 14 and transmits it to the data processing unit 12 via the control unit 46A. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and analyzes the voice trigger and extracts the necessary information. The summarization unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and summarizes the analyzed information. The reading unit reads the summarized information aloud using the speaker 40B of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0108] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0109] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0110] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0111] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0112] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0113] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0114] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0115] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0116] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0117] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0118] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0119] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0120] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0121] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0122] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0123] Each of the multiple elements described above, including the reception unit, analysis unit, summarization unit, and reading unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit receives a voice trigger using the microphone 238 of the smart glasses 214 and transmits it to the data processing unit 12 via the control unit 46A. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the voice trigger and extracts the necessary information. The summarization unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which summarizes the analyzed information. The reading unit reads the summarized information aloud using, for example, the speaker 240 of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0124] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0125] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0126] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0127] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0128] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0129] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0130] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0131] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0132] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0133] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0134] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0135] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0136] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0137] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0138] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0139] Each of the multiple elements described above, including the reception unit, analysis unit, summarization unit, and reading unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit receives an audio trigger using the microphone 238 of the headset terminal 314 and transmits it to the data processing unit 12 via the control unit 46A. The analysis unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which analyzes the audio trigger and extracts the necessary information. The summarization unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, which summarizes the analyzed information. The reading unit reads the summarized information aloud using, for example, the speaker 240 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0140] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0141] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0142] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0143] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0144] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0145] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0146] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0147] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0148] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0149] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0150] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0151] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0152] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0153] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0154] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0155] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0156] Each of the multiple elements described above, including the reception unit, analysis unit, summarization unit, and reading unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the reception unit receives a voice trigger using the microphone 238 of the robot 414 and transmits it to the data processing unit 12 via the control unit 46A. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the voice trigger and extracts the necessary information. The summarization unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which summarizes the analyzed information. The reading unit reads the summarized information aloud using, for example, the speaker 240 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0157] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0158] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0159] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0160] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0161] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0162] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0163] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0164] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0165] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0166] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0167] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0168] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0169] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0170] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0171] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0172] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0173] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0174] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0175] (Note 1) A reception unit that accepts voice triggers, An analysis unit analyzes information based on the voice trigger received by the reception unit, A summarization unit that summarizes the information analyzed by the aforementioned analysis unit, The system includes a reading unit that reads out the information summarized by the summarization unit. A system characterized by the following features. (Note 2) The aforementioned analysis unit, Transcribing specific information The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned reading unit, Reads out information that matches the user's request. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned reception unit is It estimates the user's emotions and adjusts the timing of voice trigger acceptance based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned reception unit is The system analyzes the user's past voice trigger history and selects the optimal triggering method. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned reception unit is When an audio trigger is received, the system filters the user's current ambient noise to remove noise. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned reception unit is It estimates the user's emotions and determines the priority of voice triggers to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned reception unit is When an audio trigger is received, the system prioritizes receiving triggers that are more relevant based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is When an audio trigger is received, the system analyzes the user's social media activity and accepts relevant triggers. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, We estimate the user's emotions and adjust the information analysis method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, During analysis, adjust the level of detail based on the importance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During the analysis, the priority of the analysis is determined based on when the information was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 16) The summary section above is, It estimates the user's emotions and adjusts the way the summary is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The summary section above is, When generating a summary, adjust the level of detail in the summary based on the importance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 18) The summary section above is, When generating summaries, different summarization algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 19) The summary section above is, It estimates the user's sentiment and adjusts the length of the summary based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 20) The summary section above is, When generating summaries, prioritize summaries based on when the information was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 21) The summary section above is, When generating summaries, adjust the order of the summaries based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned reading unit, It estimates the user's emotions and adjusts the tone and speed of the text-to-speech based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned reading unit, When reading aloud, adjust the level of detail based on the importance of the information. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned reading unit, When reading aloud, different reading algorithms are applied depending on the category of information. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned reading unit, It estimates the user's emotions and adjusts the reading order based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned reading unit, When reading aloud, the priority of reading is determined based on when the information was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned reading unit, When reading aloud, the reading order is adjusted based on the relevance of the information. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0176] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A reception unit that accepts voice triggers, An analysis unit analyzes information based on the voice trigger received by the reception unit, A summarization unit that summarizes the information analyzed by the aforementioned analysis unit, The system includes a reading unit that reads out the information summarized by the summarization unit. A system characterized by the following features.

2. The aforementioned analysis unit, Transcribing specific information The system according to feature 1.

3. The aforementioned reading unit, Reads out information that matches the user's request. The system according to feature 1.

4. The aforementioned reception unit is It estimates the user's emotions and adjusts the timing of voice trigger acceptance based on the estimated user emotions. The system according to feature 1.

5. The aforementioned reception unit is The system analyzes the user's past voice trigger history and selects the optimal triggering method. The system according to feature 1.

6. The aforementioned reception unit is When an audio trigger is received, the system filters the user's current ambient noise to remove noise. The system according to feature 1.

7. The aforementioned reception unit is It estimates the user's emotions and determines the priority of voice triggers to accept based on the estimated user emotions. The system according to feature 1.

8. The aforementioned reception unit is When an audio trigger is received, the system prioritizes receiving triggers that are more relevant based on the user's geographical location. The system according to feature 1.

9. The aforementioned reception unit is When an audio trigger is received, the system analyzes the user's social media activity and accepts relevant triggers. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A