system

The system efficiently extracts information from images and provides tailored advice using image recognition, OCR, and AI to address the challenge of understanding complex documents, enhancing user accessibility.

JP2026072456APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional systems struggle to efficiently extract useful information from images and provide appropriate advice, particularly in complex documents like insurance clauses or service terms.

Method used

A system comprising an analysis unit, generation unit, and advice unit that uses image recognition, OCR, deep learning, natural language processing, and generative AI to analyze images, convert characters into text data, and provide tailored advice based on user questions.

Benefits of technology

Enables quick and accurate extraction of information from images and generation of appropriate advice, supporting users with declining eyesight or cognitive abilities, and providing personalized assistance in understanding complex documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072456000001_ABST
    Figure 2026072456000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to extract useful information from images taken by the user and provide appropriate advice. [Solution] The system according to the embodiment comprises an analysis unit, a generation unit, and an advice unit. The analysis unit analyzes images taken by the user. The generation unit generates text data from the images analyzed by the analysis unit. The advice unit analyzes the text data generated by the generation unit and provides advice in response to the user's questions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, there is a problem that it is difficult to efficiently extract useful information from an image taken by a user and provide appropriate advice.

[0005] The system according to the embodiment aims to extract useful information from an image taken by a user and provide appropriate advice.

Means for Solving the Problems

[0006] The system according to this embodiment comprises an analysis unit, a generation unit, and an advice unit. The analysis unit analyzes images taken by the user. The generation unit generates text data from the images analyzed by the analysis unit. The advice unit analyzes the text data generated by the generation unit and provides advice in response to the user's questions. [Effects of the Invention]

[0007] The system according to this embodiment can extract useful information from images taken by the user and provide appropriate advice. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The advice provision system according to an embodiment of the present invention is a system that uses a generating AI to interpret detailed terms and conditions such as "service usage terms," ​​"insurance clauses," and "specifications" when purchasing goods, and provides appropriate advice to the user. The advice provision system works by having the user take a photograph of the terms and conditions document and feed the image into the generating AI. The generating AI analyzes the image and converts the document's content into text data. The generating AI then analyzes the converted text data and provides appropriate advice in response to the user's question. For example, if a user asks, "I've been injured. Are there any good hospitals nearby?", the generating AI analyzes the insurance clause photographed by the user and provides advice such as, "You have a bicycle liability insurance rider!" This mechanism allows users to avoid the trouble of reading detailed terms and conditions and to quickly obtain the necessary information. For example, even people with declining eyesight or cognitive abilities can easily understand detailed contract information and receive appropriate advice by using the generating AI. Furthermore, the generating AI is a useful support tool for people who are too busy to read the terms and conditions or who cannot remember contract clauses. This allows the advice system to support users' lives and eliminate the need to read detailed terms and conditions.

[0029] The advice provision system according to this embodiment comprises an analysis unit, a generation unit, and an advice unit. The analysis unit analyzes an image taken by the user. The analysis unit extracts characters from the image using, for example, image recognition technology. The analysis unit can recognize characters from the image using, for example, OCR technology and convert them into text data. The analysis unit can also extract characters from the image with high accuracy using, for example, deep learning technology. The generation unit generates text data from the image analyzed by the analysis unit. The generation unit converts the extracted characters into text data. The generation unit can generate text data using, for example, natural language processing technology. The generation unit can also generate text data using, for example, generative AI. The advice unit analyzes the text data generated by the generation unit and provides advice to the user in response to their questions. The advice unit provides, for example, appropriate advice based on the content of the user's questions. The advice unit can, for example, analyze insurance clauses and provide the user with advice regarding insurance riders. The advice unit can also suggest nearby hospitals based on the user's questions. As a result, the advice provision system according to this embodiment can analyze images taken by the user, generate text data, and provide appropriate advice in response to the user's questions.

[0030] The analysis unit analyzes images taken by the user. For example, the analysis unit extracts characters from images using image recognition technology. Specifically, when a user uploads an image taken with a smartphone or digital camera to the system, the analysis unit receives the image and first performs preprocessing. Preprocessing includes denoising, contrast adjustment, and rotation correction. This makes the characters in the image more clearly recognizable. Next, the analysis unit recognizes the characters in the image using OCR (Optical Character Recognition) technology and converts them into text data. OCR technology analyzes characters in an image at the pixel level and identifies the shape of the characters using pattern matching and machine learning algorithms. Furthermore, by using deep learning technology, it is possible to extract characters with high accuracy even from handwritten characters and complex fonts. Deep learning models are pre-trained on a large dataset of character images and learn the shape and features of characters with high accuracy, thus achieving higher recognition accuracy than conventional OCR technology. By combining these technologies, the analysis unit can quickly and accurately extract characters from images taken by the user and output them as text data.

[0031] The generation unit generates text data from images analyzed by the analysis unit. For example, the generation unit converts extracted characters into text data. Specifically, it receives character data sent from the analysis unit, formats it, and outputs it as consistent text data. The generation unit uses natural language processing (NLP) techniques to analyze the grammar and structure of the text data and arrange it into meaningful sentences. For example, it performs tasks such as sentence segmentation, punctuation insertion, and correction of typos and grammatical errors. Furthermore, the generation unit can generate even more sophisticated text data using a generation AI. The generation AI is pre-trained with a large amount of text data and has the ability to understand context and generate natural-sounding sentences. For example, based on character data extracted from an image of an insurance contract taken by a user, the generation AI can summarize the contract contents or highlight important clauses. The generation unit utilizes these techniques to generate high-quality text data from the character data obtained from the user's images and hands it over to the next advice unit.

[0032] The advice unit analyzes the text data generated by the generation unit and provides advice in response to user questions. Specifically, it analyzes the content of the questions entered by the user into the system and extracts information related to those questions from the text data received from the generation unit. For example, the advice unit can analyze insurance clauses and provide the user with advice on insurance riders. For analyzing insurance clauses, natural language processing technology is used to extract specific keywords and phrases within the contract and provide appropriate advice to the user based on that. The advice unit can also suggest nearby hospitals based on the user's questions. For example, if a user asks, "Can you tell me about nearby hospitals?", the advice unit uses location information services to determine the user's current location and provides information on hospitals in that area. Furthermore, the advice unit can also generate answers to user questions using generative AI. Generative AI has the ability to understand the content of the user's questions and generate appropriate answers. For example, if a user asks, "Under what circumstances does this insurance apply?", the generative AI analyzes the content of the contract and generates an answer that clearly explains the application conditions. This allows the advice section to provide quick and accurate advice to users' questions, resolving their doubts and concerns.

[0033] The analysis unit can extract characters from an image using image recognition technology. For example, the analysis unit can recognize characters in an image using OCR technology and convert them into text data. The analysis unit can also extract characters from an image with high accuracy using deep learning technology. Furthermore, the analysis unit can recognize handwritten characters using image recognition technology and convert them into text data. This allows for accurate extraction of characters from an image using image recognition technology.

[0034] The generation unit can convert extracted characters into text data. The generation unit can generate text data using, for example, natural language processing techniques. The generation unit can also generate text data using, for example, generative AI. The generation unit can also convert extracted characters into text data based on context. This allows the analysis results to be used in text format by converting the extracted characters into text data.

[0035] The advice section can provide appropriate advice in response to user questions. For example, the advice section can provide appropriate advice based on the content of the user's question. For example, the advice section can also provide relevant information in response to the user's question. For example, the advice section can propose specific solutions in response to the user's question. In this way, by providing appropriate advice in response to user questions, the user's doubts can be resolved.

[0036] The advisory unit can analyze insurance clauses and provide users with advice on insurance riders. For example, the advisory unit can analyze insurance clauses and provide users with appropriate advice on insurance riders. The advisory unit can also provide users with appropriate advice based on insurance riders. For example, the advisory unit can analyze insurance clauses and provide users with advice on the scope of insurance coverage. Thus, by analyzing insurance clauses, it can provide users with appropriate advice on insurance riders.

[0037] The advice unit can suggest nearby hospitals based on the user's questions. For example, it can suggest nearby hospitals based on the content of the user's questions. It can also suggest nearby hospitals based on the user's location information. Furthermore, it can suggest appropriate hospitals based on the user's symptoms. This allows the system to provide advice tailored to the user's needs by suggesting nearby hospitals based on the user's questions.

[0038] The analysis unit can apply different analysis algorithms depending on the type of document during image analysis. For example, in the case of terms of service, the generating AI can apply an algorithm specialized in analyzing contract clauses. For example, in the case of insurance clauses, the generating AI can apply an algorithm specialized in analyzing insurance special provisions. For example, in the case of specifications, the generating AI can apply an algorithm specialized in analyzing technical content. By applying an analysis algorithm appropriate to the type of document, the analysis accuracy can be improved.

[0039] The analysis unit can improve analysis accuracy by considering the document's layout and format during image analysis. For example, if a document is divided into multiple columns, the generating AI can analyze each column individually. The analysis unit can also ignore images and diagrams in a document and analyze only the text portion. For example, if a document is handwritten, the generating AI can use handwriting recognition technology for analysis. This improves analysis accuracy by considering the document's layout and format.

[0040] The analysis unit can select the optimal analysis method by referring to the user's past analysis history during image analysis. For example, the analysis unit can use the types of documents the user has previously analyzed to select the optimal analysis method for the generating AI. The analysis unit can also use the user's past analysis history to make adjustments to the generating AI to reduce misrecognition. For example, the analysis unit can also prioritize the application of analysis algorithms the user has used in the past. This allows the system to select the optimal analysis method and improve analysis accuracy by referring to the user's past analysis history.

[0041] The analysis unit can apply the optimal analysis algorithm during image analysis, taking into account the user's device information. For example, if the user is using a smartphone, the generating AI will apply an analysis algorithm optimized for smartphones. Similarly, if the user is using a tablet, the generating AI can apply an analysis algorithm optimized for tablets. Furthermore, if the user is using a desktop computer, the generating AI can apply an analysis algorithm optimized for desktops. This allows the system to apply the optimal analysis algorithm by considering the user's device information, thereby improving analysis accuracy.

[0042] The generation unit can adjust the level of detail generated based on the importance of the document during text generation. For example, in the case of an important document, the generation AI will generate text with detailed explanations. For example, in the case of a general document, the generation AI can also generate text with a standard level of detail. For example, in the case of a simple document, the generation AI can also generate concise text. By adjusting the level of detail generated based on the importance of the document, it is possible to generate text with an appropriate level of detail.

[0043] The generation unit can apply different generation algorithms depending on the document category when generating text. For example, in the case of a contract, the generation AI can generate text using legal terminology. For example, in the case of insurance clauses, the generation AI can generate text using insurance-specific terminology. For example, in the case of a specification document, the generation AI can generate text using technical terminology. By applying a generation algorithm appropriate to the document category, it is possible to generate text with appropriate expression.

[0044] The generation unit can determine the priority of text generation based on the document's submission date. For example, in the case of an urgent document, the generation AI will prioritize text generation. For example, in the case of a document with an approaching submission deadline, the generation AI can also generate text earlier. For example, in the case of a document with a far-off submission deadline, the generation AI can also postpone text generation. This allows for text generation tailored to the urgency of the document by determining the generation priority based on its submission date.

[0045] The generation unit can adjust the generation order based on the relevance of documents during text generation. For example, the generation unit will prioritize generating text from highly relevant documents. The generation unit can also postpone generating text from less relevant documents. For example, if multiple documents are related, the generation unit can generate text from them all at once. This allows the generation unit to prioritize the generation of highly relevant documents by adjusting the generation order based on the relevance of documents.

[0046] The advice unit can select the most appropriate advice by referring to the user's past question history when providing advice. For example, the advice unit's generating AI can select the most appropriate advice based on the content of questions the user has asked in the past. The advice unit can also provide relevant advice based on the user's past question history. For example, the advice unit can provide supplementary advice based on advice the user has received in the past. In this way, the system can provide the most appropriate advice by referring to the user's past question history.

[0047] The advice function can customize the content of advice based on the user's current situation and needs when providing it. For example, if the user is injured, the AI ​​can suggest a nearby hospital. If the user asks a question about insurance, the AI ​​can also provide advice on insurance riders. If the user is traveling, the AI ​​can also provide advice for their travel destination. This allows for more appropriate support by providing advice tailored to the user's current situation and needs.

[0048] The advice function can provide optimal advice by considering the user's geographical location. For example, the AI ​​can suggest nearby hospitals based on the user's current location. If the user is traveling, the AI ​​can also provide advice for their travel destination. If the user is in a specific region, the AI ​​can provide region-specific advice. This allows the system to provide appropriate advice based on the user's current location by considering their geographical location.

[0049] The advice function can analyze a user's social media activity and provide relevant advice when offering it. For example, the advice function's generative AI can provide relevant advice based on information shared by the user on social media. The advice function's generative AI can also provide advice based on information about accounts the user follows on social media. The advice function's generative AI can also provide advice based on information about groups the user participates in on social media. In this way, by analyzing the user's social media activity, it can provide relevant advice.

[0050] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0051] The advice system can provide more personalized advice by referring to the user's past behavior history. For example, if a user has asked many questions about a particular insurance policy in the past, the advice system can prioritize providing detailed information about that policy. It can also suggest hospitals again based on information about hospitals the user has visited in the past. Furthermore, it can provide supplementary advice by referring to the user's history of advice received in the past. In this way, it can provide more appropriate advice by leveraging the user's past behavior history.

[0052] An advice system can monitor a user's current health status and provide advice based on that information. For example, if a user is using a health management app, the system can refer to that data and provide appropriate health advice. If a user is using a wearable device, the system can also provide advice on exercise and diet based on that data. Furthermore, it can suggest appropriate treatments and preventative measures based on the user's diagnosis from a medical institution. This allows for more effective support by providing advice tailored to the user's health condition.

[0053] The advice system can leverage the user's geographical location to provide region-specific advice. For example, if a user is in a specific area, it can provide information about local medical facilities and services. If the user is traveling, it can provide emergency contact information and medical facilities in their destination. Furthermore, if the user is planning to move, it can provide information useful for living in their new area. This allows for more appropriate advice to be provided by utilizing the user's geographical location.

[0054] The advice-providing system can analyze a user's social media activity and provide relevant advice. For example, based on information a user shares on social media, the generative AI can provide relevant advice. It can also provide advice based on information about accounts a user follows. Furthermore, it can provide advice based on information about groups a user participates in. This allows for the provision of more appropriate advice by analyzing the user's social media activity.

[0055] The advice-providing system can offer optimal advice by taking into account the user's device information. For example, if the user is using a smartphone, the generating AI can provide advice optimized for smartphones. Similarly, if the user is using a tablet, the generating AI can provide advice optimized for tablets. Furthermore, if the user is using a desktop computer, the generating AI can provide advice optimized for desktops. This allows for the provision of more appropriate advice by considering the user's device information.

[0056] The following briefly describes the processing flow for example form 1.

[0057] Step 1: The analysis unit analyzes the image taken by the user. The analysis unit extracts characters from the image using, for example, image recognition technology. Furthermore, it can recognize characters in the image with high accuracy using OCR technology and deep learning technology and convert them into text data. Step 2: The generation unit generates text data from the image analyzed by the analysis unit. The generation unit can, for example, convert extracted characters into text data and generate text data using natural language processing technology or generative AI. Step 3: The advice unit analyzes the text data generated by the generation unit and provides advice in response to the user's question. For example, the advice unit can provide appropriate advice based on the user's question, analyze insurance clauses to provide advice on insurance riders, or suggest nearby hospitals.

[0058] (Example of form 2) The advice provision system according to an embodiment of the present invention is a system that uses a generating AI to interpret detailed terms and conditions such as "service usage terms," ​​"insurance clauses," and "specifications" when purchasing goods, and provides appropriate advice to the user. The advice provision system works by having the user take a photograph of the terms and conditions document and feed the image into the generating AI. The generating AI analyzes the image and converts the document's content into text data. The generating AI then analyzes the converted text data and provides appropriate advice in response to the user's question. For example, if a user asks, "I've been injured. Are there any good hospitals nearby?", the generating AI analyzes the insurance clause photographed by the user and provides advice such as, "You have a bicycle liability insurance rider!" This mechanism allows users to avoid the trouble of reading detailed terms and conditions and to quickly obtain the necessary information. For example, even people with declining eyesight or cognitive abilities can easily understand detailed contract information and receive appropriate advice by using the generating AI. Furthermore, the generating AI is a useful support tool for people who are too busy to read the terms and conditions or who cannot remember contract clauses. This allows the advice system to support users' lives and eliminate the need to read detailed terms and conditions.

[0059] The advice provision system according to this embodiment comprises an analysis unit, a generation unit, and an advice unit. The analysis unit analyzes an image taken by the user. The analysis unit extracts characters from the image using, for example, image recognition technology. The analysis unit can recognize characters from the image using, for example, OCR technology and convert them into text data. The analysis unit can also extract characters from the image with high accuracy using, for example, deep learning technology. The generation unit generates text data from the image analyzed by the analysis unit. The generation unit converts the extracted characters into text data. The generation unit can generate text data using, for example, natural language processing technology. The generation unit can also generate text data using, for example, generative AI. The advice unit analyzes the text data generated by the generation unit and provides advice to the user in response to their questions. The advice unit provides, for example, appropriate advice based on the content of the user's questions. The advice unit can, for example, analyze insurance clauses and provide the user with advice regarding insurance riders. The advice unit can also suggest nearby hospitals based on the user's questions. As a result, the advice provision system according to this embodiment can analyze images taken by the user, generate text data, and provide appropriate advice in response to the user's questions.

[0060] The analysis unit analyzes images taken by the user. For example, the analysis unit extracts characters from images using image recognition technology. Specifically, when a user uploads an image taken with a smartphone or digital camera to the system, the analysis unit receives the image and first performs preprocessing. Preprocessing includes denoising, contrast adjustment, and rotation correction. This makes the characters in the image more clearly recognizable. Next, the analysis unit recognizes the characters in the image using OCR (Optical Character Recognition) technology and converts them into text data. OCR technology analyzes characters in an image at the pixel level and identifies the shape of the characters using pattern matching and machine learning algorithms. Furthermore, by using deep learning technology, it is possible to extract characters with high accuracy even from handwritten characters and complex fonts. Deep learning models are pre-trained on a large dataset of character images and learn the shape and features of characters with high accuracy, thus achieving higher recognition accuracy than conventional OCR technology. By combining these technologies, the analysis unit can quickly and accurately extract characters from images taken by the user and output them as text data.

[0061] The generation unit generates text data from images analyzed by the analysis unit. For example, the generation unit converts extracted characters into text data. Specifically, it receives character data sent from the analysis unit, formats it, and outputs it as consistent text data. The generation unit uses natural language processing (NLP) techniques to analyze the grammar and structure of the text data and arrange it into meaningful sentences. For example, it performs tasks such as sentence segmentation, punctuation insertion, and correction of typos and grammatical errors. Furthermore, the generation unit can generate even more sophisticated text data using a generation AI. The generation AI is pre-trained with a large amount of text data and has the ability to understand context and generate natural-sounding sentences. For example, based on character data extracted from an image of an insurance contract taken by a user, the generation AI can summarize the contract contents or highlight important clauses. The generation unit utilizes these techniques to generate high-quality text data from the character data obtained from the user's images and hands it over to the next advice unit.

[0062] The advice unit analyzes the text data generated by the generation unit and provides advice in response to user questions. Specifically, it analyzes the content of the questions entered by the user into the system and extracts information related to those questions from the text data received from the generation unit. For example, the advice unit can analyze insurance clauses and provide the user with advice on insurance riders. For analyzing insurance clauses, natural language processing technology is used to extract specific keywords and phrases within the contract and provide appropriate advice to the user based on that. The advice unit can also suggest nearby hospitals based on the user's questions. For example, if a user asks, "Can you tell me about nearby hospitals?", the advice unit uses location information services to determine the user's current location and provides information on hospitals in that area. Furthermore, the advice unit can also generate answers to user questions using generative AI. Generative AI has the ability to understand the content of the user's questions and generate appropriate answers. For example, if a user asks, "Under what circumstances does this insurance apply?", the generative AI analyzes the content of the contract and generates an answer that clearly explains the application conditions. This allows the advice section to provide quick and accurate advice to users' questions, resolving their doubts and concerns.

[0063] The analysis unit can extract characters from an image using image recognition technology. For example, the analysis unit can recognize characters in an image using OCR technology and convert them into text data. The analysis unit can also extract characters from an image with high accuracy using deep learning technology. Furthermore, the analysis unit can recognize handwritten characters using image recognition technology and convert them into text data. This allows for accurate extraction of characters from an image using image recognition technology.

[0064] The generation unit can convert extracted characters into text data. The generation unit can generate text data using, for example, natural language processing techniques. The generation unit can also generate text data using, for example, generative AI. The generation unit can also convert extracted characters into text data based on context. This allows the analysis results to be used in text format by converting the extracted characters into text data.

[0065] The advice section can provide appropriate advice in response to user questions. For example, the advice section can provide appropriate advice based on the content of the user's question. For example, the advice section can also provide relevant information in response to the user's question. For example, the advice section can propose specific solutions in response to the user's question. In this way, by providing appropriate advice in response to user questions, the user's doubts can be resolved.

[0066] The advisory unit can analyze insurance clauses and provide users with advice on insurance riders. For example, the advisory unit can analyze insurance clauses and provide users with appropriate advice on insurance riders. The advisory unit can also provide users with appropriate advice based on insurance riders. For example, the advisory unit can analyze insurance clauses and provide users with advice on the scope of insurance coverage. Thus, by analyzing insurance clauses, it can provide users with appropriate advice on insurance riders.

[0067] The advice unit can suggest nearby hospitals based on the user's questions. For example, it can suggest nearby hospitals based on the content of the user's questions. It can also suggest nearby hospitals based on the user's location information. Furthermore, it can suggest appropriate hospitals based on the user's symptoms. This allows the system to provide advice tailored to the user's needs by suggesting nearby hospitals based on the user's questions.

[0068] The analysis unit can estimate the user's emotions and adjust the accuracy of image analysis based on the estimated emotions. For example, if the user is stressed, the analysis unit can improve the accuracy of image analysis and reduce misrecognition using the generative AI. If the user is relaxed, the analysis unit can also have the generative AI perform image analysis with normal accuracy. If the user is in a hurry, the analysis unit can have the generative AI perform image analysis quickly and provide results sooner. This allows for more accurate analysis results by adjusting the accuracy of image analysis according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0069] The analysis unit can apply different analysis algorithms depending on the type of document during image analysis. For example, in the case of terms of service, the generating AI can apply an algorithm specialized in analyzing contract clauses. For example, in the case of insurance clauses, the generating AI can apply an algorithm specialized in analyzing insurance special provisions. For example, in the case of specifications, the generating AI can apply an algorithm specialized in analyzing technical content. By applying an analysis algorithm appropriate to the type of document, the analysis accuracy can be improved.

[0070] The analysis unit can improve analysis accuracy by considering the document's layout and format during image analysis. For example, if a document is divided into multiple columns, the generating AI can analyze each column individually. The analysis unit can also ignore images and diagrams in a document and analyze only the text portion. For example, if a document is handwritten, the generating AI can use handwriting recognition technology for analysis. This improves analysis accuracy by considering the document's layout and format.

[0071] The analysis unit can estimate the user's emotions and adjust the display method of the analysis results based on the estimated emotions. For example, if the user is nervous, the generation AI can provide a simple and highly visible display method. If the user is relaxed, the generation AI can also provide a display method that includes detailed information. If the user is in a hurry, the generation AI can also provide a display method that gets straight to the point. In this way, by adjusting the display method of the analysis results according to the user's emotions, a display that is easy for the user to understand can be provided. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples.

[0072] The analysis unit can select the optimal analysis method by referring to the user's past analysis history during image analysis. For example, the analysis unit can use the types of documents the user has previously analyzed to select the optimal analysis method for the generating AI. The analysis unit can also use the user's past analysis history to make adjustments to the generating AI to reduce misrecognition. For example, the analysis unit can also prioritize the application of analysis algorithms the user has used in the past. This allows the system to select the optimal analysis method and improve analysis accuracy by referring to the user's past analysis history.

[0073] The analysis unit can apply the optimal analysis algorithm during image analysis, taking into account the user's device information. For example, if the user is using a smartphone, the generating AI will apply an analysis algorithm optimized for smartphones. Similarly, if the user is using a tablet, the generating AI can apply an analysis algorithm optimized for tablets. Furthermore, if the user is using a desktop computer, the generating AI can apply an analysis algorithm optimized for desktops. This allows the system to apply the optimal analysis algorithm by considering the user's device information, thereby improving analysis accuracy.

[0074] The generation unit can estimate the user's emotions and adjust the text generation method based on the estimated emotions. For example, if the user is relaxed, the generation AI will generate text using detailed and polite language. If the user is in a hurry, the generation AI can generate text using concise and to-the-point language. If the user is excited, the generation AI can generate text using visually stimulating language. By adjusting the text generation method according to the user's emotions, it is possible to generate text that is easy for the user to understand. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples.

[0075] The generation unit can adjust the level of detail generated based on the importance of the document during text generation. For example, in the case of an important document, the generation AI will generate text with detailed explanations. For example, in the case of a general document, the generation AI can also generate text with a standard level of detail. For example, in the case of a simple document, the generation AI can also generate concise text. By adjusting the level of detail generated based on the importance of the document, it is possible to generate text with an appropriate level of detail.

[0076] The generation unit can apply different generation algorithms depending on the document category when generating text. For example, in the case of a contract, the generation AI can generate text using legal terminology. For example, in the case of insurance clauses, the generation AI can generate text using insurance-specific terminology. For example, in the case of a specification document, the generation AI can generate text using technical terminology. By applying a generation algorithm appropriate to the document category, it is possible to generate text with appropriate expression.

[0077] The generation unit can estimate the user's emotions and adjust the length of the generated text based on the estimated emotions. For example, if the user is in a hurry, the generation AI can generate short, concise text. If the user is relaxed, the generation AI can also generate longer text with detailed explanations. If the user is excited, the generation AI can also generate text with visually stimulating expressions. By adjusting the text length according to the user's emotions, it is possible to generate text of an appropriate length for the user. Emotion estimation is achieved using emotion estimation functions, such as emotion engines or generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples.

[0078] The generation unit can determine the priority of text generation based on the document's submission date. For example, in the case of an urgent document, the generation AI will prioritize text generation. For example, in the case of a document with an approaching submission deadline, the generation AI can also generate text earlier. For example, in the case of a document with a far-off submission deadline, the generation AI can also postpone text generation. This allows for text generation tailored to the urgency of the document by determining the generation priority based on its submission date.

[0079] The generation unit can adjust the generation order based on the relevance of documents during text generation. For example, the generation unit will prioritize generating text from highly relevant documents. The generation unit can also postpone generating text from less relevant documents. For example, if multiple documents are related, the generation unit can generate text from them all at once. This allows the generation unit to prioritize the generation of highly relevant documents by adjusting the generation order based on the relevance of documents.

[0080] The advice unit can estimate the user's emotions and adjust the way advice is expressed based on those emotions. For example, if the user is nervous, the generating AI will provide advice in a calm manner. If the user is relaxed, the generating AI may provide advice that includes detailed information. If the user is in a hurry, the generating AI may provide concise and to-the-point advice. By adjusting the way advice is expressed according to the user's emotions, it is possible to provide advice that is easy for the user to understand. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generating AI. The generating AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples.

[0081] The advice unit can select the most appropriate advice by referring to the user's past question history when providing advice. For example, the advice unit's generating AI can select the most appropriate advice based on the content of questions the user has asked in the past. The advice unit can also provide relevant advice based on the user's past question history. For example, the advice unit can provide supplementary advice based on advice the user has received in the past. In this way, the system can provide the most appropriate advice by referring to the user's past question history.

[0082] The advice function can customize the content of advice based on the user's current situation and needs when providing it. For example, if the user is injured, the AI ​​can suggest a nearby hospital. If the user asks a question about insurance, the AI ​​can also provide advice on insurance riders. If the user is traveling, the AI ​​can also provide advice for their travel destination. This allows for more appropriate support by providing advice tailored to the user's current situation and needs.

[0083] The advice unit can estimate the user's emotions and prioritize advice based on those emotions. For example, if the user is stressed, the generative AI will prioritize providing important advice. If the user is relaxed, the generative AI can provide detailed advice. If the user is in a hurry, the generative AI can prioritize providing concise and to-the-point advice. This allows for the prioritization of important advice based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI.

[0084] The advice function can provide optimal advice by considering the user's geographical location. For example, the AI ​​can suggest nearby hospitals based on the user's current location. If the user is traveling, the AI ​​can also provide advice for their travel destination. If the user is in a specific region, the AI ​​can provide region-specific advice. This allows the system to provide appropriate advice based on the user's current location by considering their geographical location.

[0085] The advice function can analyze a user's social media activity and provide relevant advice when offering it. For example, the advice function's generative AI can provide relevant advice based on information shared by the user on social media. The advice function's generative AI can also provide advice based on information about accounts the user follows on social media. The advice function's generative AI can also provide advice based on information about groups the user participates in on social media. In this way, by analyzing the user's social media activity, it can provide relevant advice.

[0086] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0087] The advice-providing system can estimate the user's emotions and customize the advice based on those emotions. For example, if the user is feeling anxious, the advice system can provide reassuring language. If the user is agitated, the advice system can provide calming advice. Furthermore, if the user is tired, the advice system can provide concise and easy-to-understand advice. This allows for more effective support by providing advice tailored to the user's emotions.

[0088] The advice system can provide more personalized advice by referring to the user's past behavior history. For example, if a user has asked many questions about a particular insurance policy in the past, the advice system can prioritize providing detailed information about that policy. It can also suggest hospitals again based on information about hospitals the user has visited in the past. Furthermore, it can provide supplementary advice by referring to the user's history of advice received in the past. In this way, it can provide more appropriate advice by leveraging the user's past behavior history.

[0089] An advice system can monitor a user's current health status and provide advice based on that information. For example, if a user is using a health management app, the system can refer to that data and provide appropriate health advice. If a user is using a wearable device, the system can also provide advice on exercise and diet based on that data. Furthermore, it can suggest appropriate treatments and preventative measures based on the user's diagnosis from a medical institution. This allows for more effective support by providing advice tailored to the user's health condition.

[0090] The advice-providing system can estimate the user's emotions and adjust the timing of advice based on those emotions. For example, if the user is feeling stressed, the advice system can provide advice at a time when the user can relax. If the user is concentrating, the system can refrain from providing advice to avoid interrupting their concentration. Furthermore, if the user is relaxed, the system can provide advice that includes detailed information. This allows for more effective support by providing advice at a time that matches the user's emotions.

[0091] The advice system can leverage the user's geographical location to provide region-specific advice. For example, if a user is in a specific area, it can provide information about local medical facilities and services. If the user is traveling, it can provide emergency contact information and medical facilities in their destination. Furthermore, if the user is planning to move, it can provide information useful for living in their new area. This allows for more appropriate advice to be provided by utilizing the user's geographical location.

[0092] The advice-providing system can estimate the user's emotions and adjust the format of the advice based on those emotions. For example, if the user is stressed, the advice system can provide advice in a simple and visually easy-to-understand format. If the user is relaxed, the advice system can provide advice in a format that includes detailed explanations. Furthermore, if the user is in a hurry, the advice system can provide advice in a concise and to-the-point format. This allows for more effective support by providing advice in a format that matches the user's emotions.

[0093] The advice-providing system can analyze a user's social media activity and provide relevant advice. For example, based on information a user shares on social media, the generative AI can provide relevant advice. It can also provide advice based on information about accounts a user follows. Furthermore, it can provide advice based on information about groups a user participates in. This allows for the provision of more appropriate advice by analyzing the user's social media activity.

[0094] The advice-providing system can estimate the user's emotions and adjust the content of the advice based on those emotions. For example, if the user is sad, the advice system can provide advice that includes words of encouragement and comfort. If the user is happy, the advice system can provide advice that shares that joy. Furthermore, if the user is angry, the advice system can provide advice that encourages calmness. In this way, by providing advice that is tailored to the user's emotions, more effective support can be provided.

[0095] The advice-providing system can offer optimal advice by taking into account the user's device information. For example, if the user is using a smartphone, the generating AI can provide advice optimized for smartphones. Similarly, if the user is using a tablet, the generating AI can provide advice optimized for tablets. Furthermore, if the user is using a desktop computer, the generating AI can provide advice optimized for desktops. This allows for the provision of more appropriate advice by considering the user's device information.

[0096] The advice-providing system can estimate the user's emotions and prioritize advice based on those emotions. For example, if the user is stressed, the advice system can prioritize important advice. If the user is relaxed, the advice system can provide more detailed advice. Furthermore, if the user is in a hurry, the advice system can prioritize concise and to-the-point advice. In this way, by prioritizing advice according to the user's emotions, important advice can be delivered preferentially.

[0097] The following briefly describes the processing flow for example form 2.

[0098] Step 1: The analysis unit analyzes the image taken by the user. The analysis unit extracts characters from the image using, for example, image recognition technology. Furthermore, it can recognize characters in the image with high accuracy using OCR technology and deep learning technology and convert them into text data. Step 2: The generation unit generates text data from the image analyzed by the analysis unit. The generation unit can, for example, convert extracted characters into text data and generate text data using natural language processing technology or generative AI. Step 3: The advice unit analyzes the text data generated by the generation unit and provides advice in response to the user's question. For example, the advice unit can provide appropriate advice based on the user's question, analyze insurance clauses to provide advice on insurance riders, or suggest nearby hospitals.

[0099] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0100] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0101] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0102] Each of the multiple elements described above, including the analysis unit, generation unit, and advice unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the analysis unit acquires an image taken by the user using the camera 42 of the smart device 14 and extracts characters from the image using the control unit 46A. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and converts the extracted characters into text data. The advice unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the generated text data to provide appropriate advice to the user's questions. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0103] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0104] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0105] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0106] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0107] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0108] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0109] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0110] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0111] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0112] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0113] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0114] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0115] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0116] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0117] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0118] Each of the multiple elements described above, including the analysis unit, generation unit, and advice unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the analysis unit acquires an image taken by the user using the camera 42 of the smart glasses 214 and extracts characters from the image using the control unit 46A. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and converts the extracted characters into text data. The advice unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the generated text data to provide appropriate advice to the user's questions. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0119] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0120] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0121] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0122] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0123] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0124] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0125] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0126] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0127] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0128] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0129] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0130] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0131] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0132] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0133] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0134] Each of the multiple elements described above, including the analysis unit, generation unit, and advice unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the analysis unit acquires an image taken by the user using the camera 42 of the headset terminal 314 and extracts characters from the image using the control unit 46A. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and converts the extracted characters into text data. The advice unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the generated text data to provide appropriate advice in response to the user's questions. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0135] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0136] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0137] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0138] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0139] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0140] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0141] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0142] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0143] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0144] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0145] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0146] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0147] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0148] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0149] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0150] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0151] Each of the multiple elements described above, including the analysis unit, generation unit, and advice unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the analysis unit acquires an image taken by the user using the camera 42 of the robot 414 and extracts characters from the image using the control unit 46A. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12 and converts the extracted characters into text data. The advice unit is implemented in the specific processing unit 290 of the data processing unit 12 and analyzes the generated text data to provide appropriate advice in response to the user's questions. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.

[0152] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0153] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0154] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0155] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0156] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0157] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0158] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0159] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0160] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0161] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0162] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0163] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0164] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0165] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0166] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0167] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0168] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0169] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0170] (Note 1) An analysis unit that analyzes images taken by the user, A generation unit generates text data from the image analyzed by the analysis unit, The system includes an advice unit that analyzes the text data generated by the generation unit and provides advice in response to the user's questions. A system characterized by the following features. (Note 2) The aforementioned analysis unit, Extract text from an image using image recognition technology. The system described in Appendix 1, characterized by the features described herein. (Note 3) The generating unit is Convert the extracted characters into text data. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned advice section, Provide appropriate advice to users' questions. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned advice section, Analyze insurance clauses and provide users with advice on insurance riders. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned advice section, Based on the user's questions, we suggest nearby hospitals. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned analysis unit, It estimates the user's emotions and adjusts the accuracy of image analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned analysis unit, When analyzing images, different analysis algorithms are applied depending on the type of document. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned analysis unit, When analyzing images, consider the document's layout and format to improve analysis accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned analysis unit, It estimates the user's emotions and adjusts how the analysis results are displayed based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned analysis unit, During image analysis, the system selects the optimal analysis method by referring to the user's past analysis history. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned analysis unit, When analyzing images, the optimal analysis algorithm is applied, taking into account the user's device information. The system described in Appendix 1, characterized by the features described herein. (Note 13) The generating unit is It estimates the user's emotions and adjusts the text generation methodology based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The generating unit is When generating text, adjust the level of detail based on the importance of the document. The system described in Appendix 1, characterized by the features described herein. (Note 15) The generating unit is When generating text, different generation algorithms are applied depending on the document category. The system described in Appendix 1, characterized by the features described herein. (Note 16) The generating unit is It estimates the user's emotions and adjusts the length of the generated text based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The generating unit is When generating text, the generation priority is determined based on the document submission date. The system described in Appendix 1, characterized by the features described herein. (Note 18) The generating unit is When generating text, adjust the generation order based on the relevance of the documents. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned advice section, It estimates the user's emotions and adjusts the way advice is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned advice section, When providing advice, the system selects the most appropriate advice by referring to the user's past question history. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned advice section, When providing advice, customize the content of the advice based on the user's current situation and needs. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned advice section, It estimates the user's emotions and prioritizes advice based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned advice section, When providing advice, we take the user's geographical location into consideration to provide the most appropriate advice. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned advice section, When providing advice, we analyze the user's social media activity to provide relevant advice. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]

[0171] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. An analysis unit that analyzes images taken by the user, A generation unit generates text data from the image analyzed by the analysis unit, The system includes an advice unit that analyzes the text data generated by the generation unit and provides advice in response to the user's questions. A system characterized by the following features.

2. The aforementioned analysis unit, Extract text from an image using image recognition technology. The system according to feature 1.

3. The generating unit is Convert the extracted characters into text data. The system according to feature 1.

4. The aforementioned advice section, Provide appropriate advice to users' questions. The system according to feature 1.

5. The aforementioned advice section, Analyze insurance clauses and provide users with advice on insurance riders. The system according to feature 1.

6. The aforementioned advice section, Based on the user's questions, we suggest nearby hospitals. The system according to feature 1.

7. The aforementioned analysis unit, It estimates the user's emotions and adjusts the accuracy of image analysis based on the estimated user emotions. The system according to feature 1.

8. The aforementioned analysis unit, When analyzing images, different analysis algorithms are applied depending on the type of document. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A