system
The system addresses the challenge of visually impaired individuals accessing book translations by providing AI-driven audio transcriptions with intonation and sound effects, enhancing the reading experience and enabling user interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Visually impaired individuals face difficulties in easily obtaining phonetic translations of books they like.
A system comprising a reception unit, authorization unit, notification unit, analysis unit, and generation unit, which enables visually impaired users to request and receive AI-powered audio transcriptions of books, complete with intonation and sound effects, through a dedicated app.
Facilitates easy access to immersive audio versions of books for visually impaired individuals, allowing them to enjoy books with enhanced realism and the ability to share their impressions and learn about reader opinions.
Smart Images

Figure 2026073188000001_ABST
Abstract
Description
Technical Field
[0006] , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there was a problem that it was difficult for visually impaired people to easily obtain a phonetic translation of a book they liked.
[0005] The system according to the embodiment aims to enable visually impaired people to easily request and use the phonetic translation of a book.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a reception unit, an authorization unit, a notification unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives a request from a visually impaired person to have a book transcribed into audio from the library. The authorization unit grants permission for AI-powered audio transcription based on the request received by the reception unit. The notification unit notifies the visually impaired person's app of the information authorized by the authorization unit. The analysis unit analyzes the contents of the book based on the information notified by the notification unit. The generation unit generates audio effects based on the content analyzed by the analysis unit. The provision unit provides the audio generated by the generation unit to the visually impaired person. [Effects of the Invention]
[0007] The system according to this embodiment can make it easy for visually impaired people to request and use audio versions of books. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9]This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The Audio Bloom app for the visually impaired, according to an embodiment of the present invention, is a service in which a visually impaired person requests an audio version of a book from a library, and when the library grants permission for AI audio translation, a notification is sent to the user's app, and the AI reads the book aloud within the app. The Audio Bloom app for the visually impaired allows a visually impaired person to request an audio version of a book from a library, and when the library grants permission for AI audio translation, a notification is sent to the user's app, and the AI reads the book aloud within the app. The AI not only reads the text aloud, but can also add intonation and create and add sound effects according to the content to create more immersive data. The user can change the speaking speed, pitch, gender, and language within the app. Users can easily listen to many books without needing volunteers. Furthermore, users can share their impressions of the book with the AI after finishing it, and the AI can understand the content of the book and give its own impressions, or summarize and convey reader opinions found on the internet. For example, a visually impaired person requests an audio version of a book from a library. This request is sent to the library's system. Once the library approves the AI-generated audio, a notification is sent to the visually impaired user's app. This notification informs the visually impaired user that the requested book can now be audio-generated. Upon receiving the notification, the visually impaired user can begin the AI-generated audio reading within the app. The AI analyzes the content of the requested book, adding intonation and creating sound effects to create an immersive reading experience. For example, in a scene where it rains, the AI can create and play rain sounds. Furthermore, the user can adjust the speaking speed, pitch, and gender within the app, allowing for a reading tailored to their preferences. Users can also share their thoughts with the AI after finishing the book. The AI can understand the book's content and offer its own opinions, or summarize and convey reader feedback found online. For example, if a user asks, "What do you think of the ending of this book?", the AI might respond, "The ending of this book seems to have moved many readers." This system allows visually impaired individuals to easily listen to many books without the need for volunteers.Furthermore, the reading experience is enriched by AI-powered, immersive readings and the sharing of impressions. For example, an AI can read a book aloud that a visually impaired person has requested at the library, and the user can enjoy the reading. In addition, after finishing the book, the user can share their impressions with the AI and learn about the opinions of other readers. This means that the audio translation app for the visually impaired, "Audio Bloom," can enable visually impaired people to request audio translations of books from libraries, have the AI create the audio translations, and provide them to visually impaired people.
[0029] The audio translation app for the visually impaired, "Audio Bloom," according to this embodiment, comprises a reception unit, an authorization unit, a notification unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives requests from the library for audio translation of books from the visually impaired. When a visually impaired person requests audio translation of books from the library, they can, for example, send the request through the app. The request is sent to the library's system. The authorization unit grants permission for AI audio translation based on the request received by the reception unit. The authorization unit, for example, checks the content of the request and determines whether AI audio translation is possible. The notification unit notifies the visually impaired person's app of the information authorized by the authorization unit. The notification unit, for example, sends a notification to the app to inform the visually impaired person that the request has been authorized. The analysis unit analyzes the content of the book based on the information notified by the notification unit. The analysis unit, for example, analyzes the content of the requested book and extracts information for generating intonation and sound effects. The generation unit generates intonation and sound effects based on the content analyzed by the analysis unit. The generation unit, for example, uses AI to add intonation and create and add sound effects. The provisioning unit provides the audio generated by the generation unit to visually impaired individuals. The provisioning unit provides the generated audio to visually impaired individuals through an app, for example. As a result, the audio translation app for visually impaired individuals, "Audio Bloom," according to this embodiment, allows visually impaired individuals to request audio translations of books from libraries, have the AI perform the translations, and provide them to visually impaired individuals.
[0030] The reception desk handles requests from visually impaired individuals to have books transcribed into audio format. When a visually impaired person requests an audio transcription from the library, they can, for example, submit the request through an app. The request is sent to the library's system. Specifically, the visually impaired person launches the app and enters information such as the title, author, and ISBN code of the book they wish to have transcribed. The app automatically formats this information and accesses the library's database to verify the book's details. Furthermore, the app allows the visually impaired person to enter detailed requests, such as specific sections they wish to have transcribed, a specific reading speed, and the gender of the voice. This allows the visually impaired person to request an audio transcription tailored to their needs. Once a request is submitted, the library's system receives it, and the reception desk reviews its contents. The reception desk checks if the request is accurate and complete, and if any information is missing, it notifies the visually impaired person and requests additional information. This allows the reception desk to confirm the accuracy and completeness of the request and prepare to proceed to the next step.
[0031] The authorization department grants permission for AI audio translation based on requests received by the reception department. For example, the authorization department reviews the content of the request and determines whether AI audio translation is possible. Specifically, the authorization department verifies whether the requested book can be translated without copyright issues. To verify copyright, it refers to the library's database and the databases of external copyright management organizations to determine whether the book is eligible for translation. If the requested book has already been translated, or if there are similar requests from other visually impaired individuals, the authorization department adjusts the request based on that information. The authorization department also verifies whether there are sufficient resources for AI audio translation. For example, it considers the current operational status of the AI audio translation system and the time required for translation to determine whether the request can be processed appropriately. This allows the authorization department to manage requests appropriately and efficiently, enabling it to provide a fast and accurate service to visually impaired individuals.
[0032] The notification unit notifies the visually impaired user's app of information authorized by the permission unit. For example, the notification unit sends a notification to the app to inform the visually impaired user that their request has been approved. Specifically, when a request is approved, the notification unit sends a push notification to the visually impaired user's app to inform them of the request's progress. The notification includes information such as that the request has been approved, that audio transcription has started, and that audio transcription is complete. The notification unit also informs the visually impaired user of the reason if the request is not approved. For example, it explains the reason specifically, such as copyright issues or lack of resources, so that the visually impaired user can understand the next steps. The notification unit also has a notification receipt confirmation function to confirm that the visually impaired user has received the notification. This allows the notification unit to quickly and accurately communicate the progress of requests to the visually impaired user, enabling them to use the service with confidence.
[0033] The analysis unit analyzes the content of a book based on information notified by the notification unit. For example, the analysis unit analyzes the content of a requested book and extracts information for generating intonation and sound effects. Specifically, the analysis unit obtains the text data of the requested book and analyzes the structure and meaning of the text using natural language processing technology. For example, it identifies sentence breaks, paragraph structure, important keywords and phrases, and extracts the information necessary for audio translation. The analysis unit also analyzes the emotion and tone of the text and generates instructions for generating intonation and sound effects. For example, in emotionally charged scenes or important moments, it emphasizes intonation and adds sound effects to allow visually impaired people to enjoy the book's content with greater realism. The analysis unit uses AI to automatically perform these analyses, quickly and accurately preparing the book for audio translation. This allows the analysis unit to build a foundation for providing high-quality audio translation services to visually impaired people.
[0034] The generation unit generates intonation and sound effects based on the content analyzed by the analysis unit. For example, the generation unit uses AI to add intonation and create and add sound effects. Specifically, the generation unit converts text into speech using speech synthesis technology based on instructions provided by the analysis unit. For speech synthesis, it is important to give the speech natural intonation and rhythm so that it is easy for visually impaired people to hear. The generation unit uses AI to generate intonation that corresponds to the emotion and tone of the text, enabling visually impaired people to understand the content of the book more deeply. In addition, the generation unit adds sound effects so that visually impaired people can enjoy the content of the book with a greater sense of realism. For example, by adding the sound of swords and explosions in battle scenes and background music in emotional scenes, visually impaired people can become immersed in the world of the book. The generation unit combines this audio data into a single file and provides it in a format that visually impaired people can easily play. In this way, the generation unit provides a high-quality audio translation service to visually impaired people, enabling them to enjoy the content of books.
[0035] The service provider will provide the audio generated by the generation unit to visually impaired individuals. For example, the service provider will provide the generated audio to visually impaired individuals through an app. Specifically, the service provider will provide the generated audio files in a format that can be downloaded to the visually impaired person's app. The visually impaired person will be able to download the audio files through the app and play them offline. The service provider can also provide audio in streaming format. This will allow the visually impaired person to play the audio in real time wherever there is an internet connection. The service provider will make the audio playback interface intuitive and easy to use to simplify the operation for visually impaired people when playing the audio. For example, basic operations such as play, pause, rewind, and fast forward will be made easy, so that visually impaired people can enjoy the audio without stress. Furthermore, the service provider will collect feedback from visually impaired people and use it to improve the service. For example, the service provider will continuously improve the service based on feedback on audio quality, the appropriateness of sound effects, and ease of use. In this way, the service provider will provide a high-quality audio translation service to visually impaired people, enabling them to enjoy the content of books.
[0036] The service provider may have functions to change the speaking speed, voice pitch, and gender. For example, the service provider may allow the user to adjust the speaking speed within the app, for instance, by making it faster or slower. The service provider may also adjust the voice pitch, for example, by making it higher or lower. Furthermore, the service provider may change the gender, for example, by allowing the user to select a male or female voice. This enables reading aloud tailored to the user's preferences. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider may use AI to adjust the voice speed, pitch, and gender based on the user's settings.
[0037] The service provider can have a language selection function. For example, the service provider can allow the user to select a language within the app. For example, the user can choose from multiple languages, such as English, Japanese, and French. The service provider can also perform reading aloud based on the selected language. For example, it can perform reading aloud in the language chosen by the user. This makes it possible to read aloud in the language chosen by the user. Some or all of the above processing in the service provider may be performed using AI, or not using AI. For example, the service provider can have AI generate speech in the selected language based on the user's settings.
[0038] The analysis unit can generate sound effects according to the scene in the story. For example, if there is a rain scene in the story, the AI can create and play the sound of rain. Also, if there is a battle scene in the story, the AI can create and play the sound of battle. Furthermore, if there is a quiet scene in the story, the AI can create and play a quiet ambient sound. This allows for a more immersive reading experience by generating sound effects appropriate to the scene in the story. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can generate sound effects based on the scene in the story using AI.
[0039] The generation unit can perform realistic readings based on the generated audio. For example, the generation unit can use AI-generated audio to add intonation and sound effects to create a more immersive reading experience. The generation unit can also perform readings that are appropriate to the scene of the story based on the generated audio. For example, the generation unit can perform emotionally charged readings during moving scenes in the story. Furthermore, the generation unit can perform readings tailored to the user's preferences based on the generated audio. For example, the generation unit can perform readings based on the speaking speed, pitch, and gender set by the user. This enables a more immersive reading experience. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can perform realistic readings using AI-generated audio.
[0040] The service provider can include a function that allows users to share their opinions on a book they have finished reading with the AI. For example, the service provider can allow users to input their thoughts on a book they have finished reading within the app. For example, if a user asks, "What do you think of the ending of this book?", the AI can respond with a comment such as, "The ending of this book seems to have moved many readers." The service provider can also have the AI analyze the comments entered by the user and provide appropriate feedback. For example, the AI can provide feedback to the user saying, "That comment is shared by many readers." This allows users to share their thoughts on a book they have finished reading with the AI. Some or all of the above processing in the service provider may be performed using the AI, or not. For example, the service provider can have the AI analyze the comments entered by the user and provide feedback.
[0041] The service provider can have the function of having AI understand the content of a book and express its opinion. For example, the service provider can have AI analyze the content of a book and give its opinion to the user. For example, the AI may say, "The ending of this book seems to have moved many readers." The service provider can also have AI summarize and convey reader opinions found on the internet. For example, the AI may provide a summary such as, "Many readers empathize with the characters in this book." This allows the AI to understand the content of the book and express its opinion. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can have AI analyze the content of a book and express its opinion.
[0042] The delivery unit can have a function to summarize and convey reader feedback from the internet. For example, the delivery unit can use AI to collect reader opinions from the internet, summarize them, and convey them to the user. For example, the AI could provide a summary such as, "Many readers were moved by the ending of this book." The delivery unit can also use AI to analyze reader opinions and provide appropriate feedback to the user. For example, the AI could provide feedback such as, "Many readers are discussing the theme of this book." This allows for the summarization and conveyance of reader opinions from the internet. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can use AI to collect reader opinions from the internet, summarize them, and convey them.
[0043] The reception desk can analyze the user's past request history and select the most appropriate reception method. For example, the reception desk can prioritize requests for book genres that the user has frequently requested in the past. It can also prioritize suggesting request methods (voice, text, etc.) that the user has used in the past. Furthermore, based on the user's past request history, the reception desk can select the most suitable reception method for a particular time period if there are many requests during that time. This allows the reception desk to select the optimal reception method based on the user's past request history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past request history into a generating AI and have the generating AI select the optimal reception method.
[0044] The reception unit can filter requests based on the user's current reading habits and areas of interest. For example, it can prioritize requests based on the genre of books the user has recently read. It can also filter requests based on the user's areas of interest, prioritizing highly relevant books. Furthermore, it can analyze the user's reading history and prioritize requests for books that the user might be interested in. This allows for filtering requests based on the user's current reading habits and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input the user's reading habits and areas of interest into a generating AI and have the generating AI perform the filtering.
[0045] The reception unit can prioritize requests based on the user's geographical location when receiving requests. For example, if the user is near a library, the reception unit can prioritize requests for that library. It can also prioritize requests for books related to a specific region if the user is in that region. Furthermore, if the user is traveling, the reception unit can prioritize requests for books related to their travel destination. This allows for priority processing of requests based on the user's geographical location. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For instance, the reception unit can input the user's geographical location into a generating AI and have the generating AI select the most relevant requests.
[0046] The reception unit can analyze the user's social media activity when receiving a request and accept relevant requests. For example, the reception unit can prioritize requests for books that the user is talking about on social media. It can also prioritize requests for books that are frequently requested by the user's social media followers. Furthermore, the reception unit can accept relevant requests based on the content of the user's social media posts. This allows requests to be accepted based on the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's social media activity into a generating AI and have the generating AI select relevant requests.
[0047] The authorization unit can adjust the details of the authorization based on the importance of the request when granting authorization. For example, the authorization unit can perform a detailed authorization procedure for high-importance requests. It can also perform a simplified authorization procedure for low-importance requests. Furthermore, the authorization unit can adjust the level of authorization detail in stages according to the importance. This allows the level of authorization detail to be adjusted based on the importance of the request. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of authorization detail.
[0048] The permission unit can apply different permission methods depending on the category of the request when granting permission. For example, the permission unit can apply a specific permission algorithm to requests for novels. It can also apply a different permission algorithm to requests for academic books. Furthermore, it can apply yet another permission algorithm to requests for children's books. This allows different permission algorithms to be applied depending on the category of the request. Some or all of the above processing in the permission unit may be performed using AI, for example, or not using AI. For example, the permission unit can input the category of the request into a generating AI and have the generating AI execute the application of different permission algorithms.
[0049] The authorization unit can determine the order of authorizations based on when the requests were submitted. For example, the authorization unit can prioritize requests submitted early. It can also prioritize requests of high urgency. Furthermore, the authorization unit can adjust the priority of authorizations in stages according to the submission time. This allows the authorization unit to determine the priority of authorizations based on when the requests were submitted. Some or all of the above processing in the authorization unit may be performed using AI, for example, or not using AI. For example, the authorization unit can input the submission time of the requests into a generating AI and have the generating AI determine the order of authorizations.
[0050] The authorization unit can adjust the order of authorizations based on the relevance of the requests. For example, the authorization unit can prioritize authorizations if the content of the request aligns with the library's policies. It can also prioritize authorizations if the content of the request is highly relevant to other requests. Furthermore, the authorization unit can adjust the order of authorizations in stages according to the relevance of the requests. This allows the order of authorizations to be adjusted based on the relevance of the requests. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the authorization order.
[0051] The notification unit can adjust the details of the notification based on the importance of the request when it sends a notification. For example, the notification unit can send a detailed notification for high-importance requests. It can also send a simplified notification for low-importance requests. Furthermore, the notification unit can adjust the level of detail of the notification in stages according to the importance. This allows the level of detail of the notification to be adjusted based on the importance of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the notification.
[0052] The notification unit can apply different notification methods depending on the category of the request when sending a notification. For example, the notification unit can apply a specific notification algorithm to requests for novels. It can also apply a different notification algorithm to requests for academic books. Furthermore, it can apply yet another different notification algorithm to requests for children's books. This allows different notification algorithms to be applied depending on the category of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the category of the request into a generating AI and have the generating AI execute the application of different notification algorithms.
[0053] The notification unit can determine the order of notifications based on when the requests were submitted. For example, the notification unit can prioritize notifications for requests submitted early. It can also prioritize notifications for urgent requests. Furthermore, the notification unit can adjust the priority of notifications in stages according to the submission time. This allows the notification unit to determine the priority of notifications based on when the requests were submitted. Some or all of the above processing in the notification unit may be performed using AI, for example, or not using AI. For example, the notification unit can input the submission time of the requests into a generating AI and have the generating AI determine the order of notifications.
[0054] The notification unit can adjust the order of notifications based on the relevance of the requests when it sends notifications. For example, the notification unit can prioritize notifications if the content of the request is consistent with the library's policies. It can also prioritize notifications if the content of the request is highly relevant to other requests. Furthermore, the notification unit can adjust the order of notifications in stages according to the relevance of the requests. This allows the order of notifications to be adjusted based on the relevance of the requests. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the order of notifications.
[0055] The analysis unit can adjust the level of analysis based on the importance of the request during the analysis. For example, the analysis unit can perform a detailed analysis for high-importance requests. It can also perform a simplified analysis for low-importance requests. Furthermore, the analysis unit can adjust the level of detail of the analysis in stages according to the importance. This allows the level of detail of the analysis to be adjusted based on the importance of the request. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the analysis.
[0056] The analysis unit can apply different analysis methods depending on the category of the request during analysis. For example, the analysis unit can apply a specific analysis algorithm to a request for a novel. It can also apply a different analysis algorithm to a request for an academic book. Furthermore, it can apply yet another analysis algorithm to a request for a children's book. This allows for the application of different analysis algorithms depending on the category of the request. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the request category into a generating AI and have the generating AI execute the application of different analysis algorithms.
[0057] The analysis unit can determine the order of analysis based on the submission timing of requests during the analysis process. For example, the analysis unit can prioritize the analysis of requests submitted early. It can also prioritize the analysis of requests of high urgency. Furthermore, the analysis unit can adjust the priority of analysis in stages according to the submission timing. This allows the analysis priority to be determined based on the submission timing of requests. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the submission timing of requests into a generating AI and have the generating AI determine the order of analysis.
[0058] The analysis unit can adjust the order of analysis based on the relevance of requests during the analysis process. For example, the analysis unit can prioritize the analysis of requests if their content aligns with the library's policies. It can also prioritize the analysis of requests if their content is highly relevant to other requests. Furthermore, the analysis unit can adjust the order of analysis in stages according to the relevance of the requests. This allows the order of analysis to be adjusted based on the relevance of the requests. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance of requests into a generating AI and have the generating AI perform the adjustment of the analysis order.
[0059] The generation unit can adjust the details of the generated audio based on the importance of the request during generation. For example, the generation unit can generate detailed audio for high-importance requests. It can also generate simplified audio for low-importance requests. Furthermore, the generation unit can adjust the level of detail of the generated audio in stages according to the importance. This allows the level of detail of the generated audio to be adjusted based on the importance of the request. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the request into the generation AI and have the generation AI perform the adjustment of the level of detail of the generated audio.
[0060] The generation unit can apply different generation methods depending on the category of the request during generation. For example, the generation unit can apply a specific generation algorithm to a request for a novel. It can also apply a different generation algorithm to a request for an academic book. Furthermore, it can apply yet another generation algorithm to a request for a children's book. This allows for the application of different generation algorithms depending on the category of the request. Some or all of the above-described processes in the generation unit may be performed using, for example, AI, or without AI. For example, the generation unit can input the category of the request into a generation AI and cause the generation AI to apply different generation algorithms.
[0061] The generation unit can determine the order of generation based on the submission timing of requests during generation. For example, the generation unit can prioritize generating audio for requests submitted early. It can also prioritize generating audio for urgent requests. Furthermore, the generation unit can adjust the generation priority in stages according to the submission timing. This allows the generation priority to be determined based on the submission timing of requests. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the submission timing of requests into a generation AI and have the generation AI determine the order of generation.
[0062] The generation unit can adjust the generation order based on the relevance of the requests during generation. For example, the generation unit can prioritize generation if the content of the request matches the library's policies. It can also prioritize generation if the content of the request is highly relevant to other requests. Furthermore, the generation unit can adjust the generation order in stages according to the relevance of the requests. This allows the generation order to be adjusted based on the relevance of the requests. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the relevance of the requests into a generation AI and have the generation AI perform the adjustment of the generation order.
[0063] The service provider can adjust the details of the service provided based on the importance of the request. For example, the service provider can provide detailed audio for high-priority requests. It can also provide simplified audio for low-priority requests. Furthermore, the service provider can adjust the level of detail of the service in stages according to the importance. This allows the level of detail of the service to be adjusted based on the importance of the request. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the service.
[0064] The delivery unit can apply different delivery methods depending on the category of the request. For example, the delivery unit can apply a specific delivery algorithm to requests for novels. It can also apply a different delivery algorithm to requests for academic books. Furthermore, it can apply yet another delivery algorithm to requests for children's books. This allows different delivery algorithms to be applied depending on the category of the request. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can input the category of the request into a generating AI and have the generating AI execute the application of different delivery algorithms.
[0065] The service provider can determine the order of service provision based on the submission date of the requests. For example, the service provider can prioritize providing audio for requests submitted early. It can also prioritize providing audio for urgent requests. Furthermore, the service provider can adjust the priority of service provision in stages according to the submission date. This allows the service provider to determine the priority of service provision based on the submission date of the requests. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the submission date of the requests into a generating AI and have the generating AI determine the order of service provision.
[0066] The service provider can adjust the order of deliveries based on the relevance of the requests at the time of delivery. For example, the service provider can prioritize deliveries if the content of the request matches the library's policy. It can also prioritize deliveries if the content of the request is highly relevant to other requests. Furthermore, the service provider can adjust the order of deliveries in stages according to the relevance of the requests. This allows the order of deliveries to be adjusted based on the relevance of the requests. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the order of deliveries.
[0067] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0068] The reception desk can analyze the user's past request history and select the most appropriate reception method. For example, the reception desk can prioritize requests for book genres that the user has frequently requested in the past. It can also prioritize suggesting request methods (voice, text, etc.) that the user has used in the past. Furthermore, based on the user's past request history, the reception desk can select the most suitable reception method for a particular time period if there are many requests during that time. This allows the reception desk to select the optimal reception method based on the user's past request history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past request history into a generating AI and have the generating AI select the optimal reception method.
[0069] The authorization unit can adjust the details of the authorization based on the importance of the request when granting authorization. For example, the authorization unit can perform a detailed authorization procedure for high-importance requests. It can also perform a simplified authorization procedure for low-importance requests. Furthermore, the authorization unit can adjust the level of authorization detail in stages according to the importance. This allows the level of authorization detail to be adjusted based on the importance of the request. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of authorization detail.
[0070] The notification unit can adjust the details of the notification based on the importance of the request when it sends a notification. For example, the notification unit can send a detailed notification for high-importance requests. It can also send a simplified notification for low-importance requests. Furthermore, the notification unit can adjust the level of detail of the notification in stages according to the importance. This allows the level of detail of the notification to be adjusted based on the importance of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the notification.
[0071] The analysis unit can adjust the level of analysis based on the importance of the request during the analysis. For example, the analysis unit can perform a detailed analysis for high-importance requests. It can also perform a simplified analysis for low-importance requests. Furthermore, the analysis unit can adjust the level of detail of the analysis in stages according to the importance. This allows the level of detail of the analysis to be adjusted based on the importance of the request. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the analysis.
[0072] The generation unit can adjust the details of the generated audio based on the importance of the request during generation. For example, the generation unit can generate detailed audio for high-importance requests. It can also generate simplified audio for low-importance requests. Furthermore, the generation unit can adjust the level of detail of the generated audio in stages according to the importance. This allows the level of detail of the generated audio to be adjusted based on the importance of the request. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the request into the generation AI and have the generation AI perform the adjustment of the level of detail of the generated audio.
[0073] The following briefly describes the processing flow for example form 1.
[0074] Step 1: The reception desk receives requests from visually impaired individuals to have books converted into audiobooks. Visually impaired individuals submit requests, for example, through an app, and the requests are sent to the library's system. Step 2: The authorization department grants permission for AI voice translation based on the request received by the reception department. The authorization department reviews the content of the request and determines whether AI voice translation is possible. Step 3: The notification unit notifies the visually impaired person's app of the information permitted by the permission unit. The notification unit sends a notification to the app to inform the visually impaired person that the request has been permitted. Step 4: The analysis unit analyzes the book's content based on the information notified by the notification unit. The analysis unit analyzes the requested book's content and extracts information for generating intonation and sound effects. Step 5: The generation unit generates intonation and sound effects based on the analysis performed by the analysis unit. The generation unit uses AI to add intonation and create and apply sound effects. Step 6: The providing unit provides the audio generated by the generating unit to visually impaired individuals. The providing unit provides the generated audio to visually impaired individuals through the app.
[0075] (Example of form 2) The Audio Bloom app for the visually impaired, according to an embodiment of the present invention, is a service in which a visually impaired person requests an audio version of a book from a library, and when the library grants permission for AI audio translation, a notification is sent to the user's app, and the AI reads the book aloud within the app. The Audio Bloom app for the visually impaired allows a visually impaired person to request an audio version of a book from a library, and when the library grants permission for AI audio translation, a notification is sent to the user's app, and the AI reads the book aloud within the app. The AI not only reads the text aloud, but can also add intonation and create and add sound effects according to the content to create more immersive data. The user can change the speaking speed, pitch, gender, and language within the app. Users can easily listen to many books without needing volunteers. Furthermore, users can share their impressions of the book with the AI after finishing it, and the AI can understand the content of the book and give its own impressions, or summarize and convey reader opinions found on the internet. For example, a visually impaired person requests an audio version of a book from a library. This request is sent to the library's system. Once the library approves the AI-generated audio, a notification is sent to the visually impaired user's app. This notification informs the visually impaired user that the requested book can now be audio-generated. Upon receiving the notification, the visually impaired user can begin the AI-generated audio reading within the app. The AI analyzes the content of the requested book, adding intonation and creating sound effects to create an immersive reading experience. For example, in a scene where it rains, the AI can create and play rain sounds. Furthermore, the user can adjust the speaking speed, pitch, and gender within the app, allowing for a reading tailored to their preferences. Users can also share their thoughts with the AI after finishing the book. The AI can understand the book's content and offer its own opinions, or summarize and convey reader feedback found online. For example, if a user asks, "What do you think of the ending of this book?", the AI might respond, "The ending of this book seems to have moved many readers." This system allows visually impaired individuals to easily listen to many books without the need for volunteers.Furthermore, the reading experience is enriched by AI-powered, immersive readings and the sharing of impressions. For example, an AI can read a book aloud that a visually impaired person has requested at the library, and the user can enjoy the reading. In addition, after finishing the book, the user can share their impressions with the AI and learn about the opinions of other readers. This means that the audio translation app for the visually impaired, "Audio Bloom," can enable visually impaired people to request audio translations of books from libraries, have the AI create the audio translations, and provide them to visually impaired people.
[0076] The audio translation app for the visually impaired, "Audio Bloom," according to this embodiment, comprises a reception unit, an authorization unit, a notification unit, an analysis unit, a generation unit, and a provision unit. The reception unit receives requests from the library for audio translation of books from the visually impaired. When a visually impaired person requests audio translation of books from the library, they can, for example, send the request through the app. The request is sent to the library's system. The authorization unit grants permission for AI audio translation based on the request received by the reception unit. The authorization unit, for example, checks the content of the request and determines whether AI audio translation is possible. The notification unit notifies the visually impaired person's app of the information authorized by the authorization unit. The notification unit, for example, sends a notification to the app to inform the visually impaired person that the request has been authorized. The analysis unit analyzes the content of the book based on the information notified by the notification unit. The analysis unit, for example, analyzes the content of the requested book and extracts information for generating intonation and sound effects. The generation unit generates intonation and sound effects based on the content analyzed by the analysis unit. The generation unit, for example, uses AI to add intonation and create and add sound effects. The provisioning unit provides the audio generated by the generation unit to visually impaired individuals. The provisioning unit provides the generated audio to visually impaired individuals through an app, for example. As a result, the audio translation app for visually impaired individuals, "Audio Bloom," according to this embodiment, allows visually impaired individuals to request audio translations of books from libraries, have the AI perform the translations, and provide them to visually impaired individuals.
[0077] The reception desk handles requests from visually impaired individuals to have books transcribed into audio format. When a visually impaired person requests an audio transcription from the library, they can, for example, submit the request through an app. The request is sent to the library's system. Specifically, the visually impaired person launches the app and enters information such as the title, author, and ISBN code of the book they wish to have transcribed. The app automatically formats this information and accesses the library's database to verify the book's details. Furthermore, the app allows the visually impaired person to enter detailed requests, such as specific sections they wish to have transcribed, a specific reading speed, and the gender of the voice. This allows the visually impaired person to request an audio transcription tailored to their needs. Once a request is submitted, the library's system receives it, and the reception desk reviews its contents. The reception desk checks if the request is accurate and complete, and if any information is missing, it notifies the visually impaired person and requests additional information. This allows the reception desk to confirm the accuracy and completeness of the request and prepare to proceed to the next step.
[0078] The authorization department grants permission for AI audio translation based on requests received by the reception department. For example, the authorization department reviews the content of the request and determines whether AI audio translation is possible. Specifically, the authorization department verifies whether the requested book can be translated without copyright issues. To verify copyright, it refers to the library's database and the databases of external copyright management organizations to determine whether the book is eligible for translation. If the requested book has already been translated, or if there are similar requests from other visually impaired individuals, the authorization department adjusts the request based on that information. The authorization department also verifies whether there are sufficient resources for AI audio translation. For example, it considers the current operational status of the AI audio translation system and the time required for translation to determine whether the request can be processed appropriately. This allows the authorization department to manage requests appropriately and efficiently, enabling it to provide a fast and accurate service to visually impaired individuals.
[0079] The notification unit notifies the visually impaired user's app of information authorized by the permission unit. For example, the notification unit sends a notification to the app to inform the visually impaired user that their request has been approved. Specifically, when a request is approved, the notification unit sends a push notification to the visually impaired user's app to inform them of the request's progress. The notification includes information such as that the request has been approved, that audio transcription has started, and that audio transcription is complete. The notification unit also informs the visually impaired user of the reason if the request is not approved. For example, it explains the reason specifically, such as copyright issues or lack of resources, so that the visually impaired user can understand the next steps. The notification unit also has a notification receipt confirmation function to confirm that the visually impaired user has received the notification. This allows the notification unit to quickly and accurately communicate the progress of requests to the visually impaired user, enabling them to use the service with confidence.
[0080] The analysis unit analyzes the content of a book based on information notified by the notification unit. For example, the analysis unit analyzes the content of a requested book and extracts information for generating intonation and sound effects. Specifically, the analysis unit obtains the text data of the requested book and analyzes the structure and meaning of the text using natural language processing technology. For example, it identifies sentence breaks, paragraph structure, important keywords and phrases, and extracts the information necessary for audio translation. The analysis unit also analyzes the emotion and tone of the text and generates instructions for generating intonation and sound effects. For example, in emotionally charged scenes or important moments, it emphasizes intonation and adds sound effects to allow visually impaired people to enjoy the book's content with greater realism. The analysis unit uses AI to automatically perform these analyses, quickly and accurately preparing the book for audio translation. This allows the analysis unit to build a foundation for providing high-quality audio translation services to visually impaired people.
[0081] The generation unit generates intonation and sound effects based on the content analyzed by the analysis unit. For example, the generation unit uses AI to add intonation and create and add sound effects. Specifically, the generation unit converts text into speech using speech synthesis technology based on instructions provided by the analysis unit. For speech synthesis, it is important to give the speech natural intonation and rhythm so that it is easy for visually impaired people to hear. The generation unit uses AI to generate intonation that corresponds to the emotion and tone of the text, enabling visually impaired people to understand the content of the book more deeply. In addition, the generation unit adds sound effects so that visually impaired people can enjoy the content of the book with a greater sense of realism. For example, by adding the sound of swords and explosions in battle scenes and background music in emotional scenes, visually impaired people can become immersed in the world of the book. The generation unit combines this audio data into a single file and provides it in a format that visually impaired people can easily play. In this way, the generation unit provides a high-quality audio translation service to visually impaired people, enabling them to enjoy the content of books.
[0082] The service provider will provide the audio generated by the generation unit to visually impaired individuals. For example, the service provider will provide the generated audio to visually impaired individuals through an app. Specifically, the service provider will provide the generated audio files in a format that can be downloaded to the visually impaired person's app. The visually impaired person will be able to download the audio files through the app and play them offline. The service provider can also provide audio in streaming format. This will allow the visually impaired person to play the audio in real time wherever there is an internet connection. The service provider will make the audio playback interface intuitive and easy to use to simplify the operation for visually impaired people when playing the audio. For example, basic operations such as play, pause, rewind, and fast forward will be made easy, so that visually impaired people can enjoy the audio without stress. Furthermore, the service provider will collect feedback from visually impaired people and use it to improve the service. For example, the service provider will continuously improve the service based on feedback on audio quality, the appropriateness of sound effects, and ease of use. In this way, the service provider will provide a high-quality audio translation service to visually impaired people, enabling them to enjoy the content of books.
[0083] The service provider may have functions to change the speaking speed, voice pitch, and gender. For example, the service provider may allow the user to adjust the speaking speed within the app, for instance, by making it faster or slower. The service provider may also adjust the voice pitch, for example, by making it higher or lower. Furthermore, the service provider may change the gender, for example, by allowing the user to select a male or female voice. This enables reading aloud tailored to the user's preferences. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider may use AI to adjust the voice speed, pitch, and gender based on the user's settings.
[0084] The service provider can have a language selection function. For example, the service provider can allow the user to select a language within the app. For example, the user can choose from multiple languages, such as English, Japanese, and French. The service provider can also perform reading aloud based on the selected language. For example, it can perform reading aloud in the language chosen by the user. This makes it possible to read aloud in the language chosen by the user. Some or all of the above processing in the service provider may be performed using AI, or not using AI. For example, the service provider can have AI generate speech in the selected language based on the user's settings.
[0085] The analysis unit can generate sound effects according to the scene in the story. For example, if there is a rain scene in the story, the AI can create and play the sound of rain. Also, if there is a battle scene in the story, the AI can create and play the sound of battle. Furthermore, if there is a quiet scene in the story, the AI can create and play a quiet ambient sound. This allows for a more immersive reading experience by generating sound effects appropriate to the scene in the story. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can generate sound effects based on the scene in the story using AI.
[0086] The generation unit can perform realistic readings based on the generated audio. For example, the generation unit can use AI-generated audio to add intonation and sound effects to create a more immersive reading experience. The generation unit can also perform readings that are appropriate to the scene of the story based on the generated audio. For example, the generation unit can perform emotionally charged readings during moving scenes in the story. Furthermore, the generation unit can perform readings tailored to the user's preferences based on the generated audio. For example, the generation unit can perform readings based on the speaking speed, pitch, and gender set by the user. This enables a more immersive reading experience. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can perform realistic readings using AI-generated audio.
[0087] The service provider can include a function that allows users to share their opinions on a book they have finished reading with the AI. For example, the service provider can allow users to input their thoughts on a book they have finished reading within the app. For example, if a user asks, "What do you think of the ending of this book?", the AI can respond with a comment such as, "The ending of this book seems to have moved many readers." The service provider can also have the AI analyze the comments entered by the user and provide appropriate feedback. For example, the AI can provide feedback to the user saying, "That comment is shared by many readers." This allows users to share their thoughts on a book they have finished reading with the AI. Some or all of the above processing in the service provider may be performed using the AI, or not. For example, the service provider can have the AI analyze the comments entered by the user and provide feedback.
[0088] The service provider can have the function of having AI understand the content of a book and express its opinion. For example, the service provider can have AI analyze the content of a book and give its opinion to the user. For example, the AI may say, "The ending of this book seems to have moved many readers." The service provider can also have AI summarize and convey reader opinions found on the internet. For example, the AI may provide a summary such as, "Many readers empathize with the characters in this book." This allows the AI to understand the content of the book and express its opinion. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can have AI analyze the content of a book and express its opinion.
[0089] The delivery unit can have a function to summarize and convey reader feedback from the internet. For example, the delivery unit can use AI to collect reader opinions from the internet, summarize them, and convey them to the user. For example, the AI could provide a summary such as, "Many readers were moved by the ending of this book." The delivery unit can also use AI to analyze reader opinions and provide appropriate feedback to the user. For example, the AI could provide feedback such as, "Many readers are discussing the theme of this book." This allows for the summarization and conveyance of reader opinions from the internet. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can use AI to collect reader opinions from the internet, summarize them, and convey them.
[0090] The reception unit can estimate the user's emotions and adjust the timing of request acceptance based on the estimated emotions. For example, if the user is relaxed, the reception unit can accept the request immediately. If the user is stressed, the reception unit can slightly delay accepting the request to give the user time to calm down. Furthermore, if the user is in a hurry, the reception unit can accept the request quickly and start processing immediately. This allows the timing of request acceptance to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not using AI. For example, the reception unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0091] The reception desk can analyze the user's past request history and select the most appropriate reception method. For example, the reception desk can prioritize requests for book genres that the user has frequently requested in the past. It can also prioritize suggesting request methods (voice, text, etc.) that the user has used in the past. Furthermore, based on the user's past request history, the reception desk can select the most suitable reception method for a particular time period if there are many requests during that time. This allows the reception desk to select the optimal reception method based on the user's past request history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past request history into a generating AI and have the generating AI select the optimal reception method.
[0092] The reception unit can filter requests based on the user's current reading habits and areas of interest. For example, it can prioritize requests based on the genre of books the user has recently read. It can also filter requests based on the user's areas of interest, prioritizing highly relevant books. Furthermore, it can analyze the user's reading history and prioritize requests for books that the user might be interested in. This allows for filtering requests based on the user's current reading habits and areas of interest. Some or all of the above processing in the reception unit may be performed using AI, for example, or not. For example, the reception unit can input the user's reading habits and areas of interest into a generating AI and have the generating AI perform the filtering.
[0093] The reception unit can estimate the user's emotions and determine the priority of requests based on the estimated emotions. For example, if the user is excited, the reception unit can set a higher priority for the request. If the user is relaxed, the reception unit can set the priority to normal. Furthermore, if the user is stressed, the reception unit can set a lower priority for the request and prioritize other requests. This allows the system to determine the priority of requests according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI, or not using AI. For example, the reception unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0094] The reception unit can prioritize requests based on the user's geographical location when receiving requests. For example, if the user is near a library, the reception unit can prioritize requests for that library. It can also prioritize requests for books related to a specific region if the user is in that region. Furthermore, if the user is traveling, the reception unit can prioritize requests for books related to their travel destination. This allows for priority processing of requests based on the user's geographical location. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For instance, the reception unit can input the user's geographical location into a generating AI and have the generating AI select the most relevant requests.
[0095] The reception unit can analyze the user's social media activity when receiving a request and accept relevant requests. For example, the reception unit can prioritize requests for books that the user is talking about on social media. It can also prioritize requests for books that are frequently requested by the user's social media followers. Furthermore, the reception unit can accept relevant requests based on the content of the user's social media posts. This allows requests to be accepted based on the user's social media activity. Some or all of the above processing in the reception unit may be performed using AI, for example, or without AI. For example, the reception unit can input the user's social media activity into a generating AI and have the generating AI select relevant requests.
[0096] The permission unit can estimate the user's emotions and adjust the permission criteria based on the estimated emotions. For example, if the user is relaxed, the permission unit can relax the permission criteria. Conversely, if the user is stressed, the permission unit can tighten the permission criteria. Furthermore, if the user is in a hurry, the permission unit can make a quick decision on the permission criteria. This allows the permission criteria to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the permission unit may be performed using AI, for example, or not using AI. For example, the permission unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0097] The authorization unit can adjust the details of the authorization based on the importance of the request when granting authorization. For example, the authorization unit can perform a detailed authorization procedure for high-importance requests. It can also perform a simplified authorization procedure for low-importance requests. Furthermore, the authorization unit can adjust the level of authorization detail in stages according to the importance. This allows the level of authorization detail to be adjusted based on the importance of the request. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of authorization detail.
[0098] The permission unit can apply different permission methods depending on the category of the request when granting permission. For example, the permission unit can apply a specific permission algorithm to requests for novels. It can also apply a different permission algorithm to requests for academic books. Furthermore, it can apply yet another permission algorithm to requests for children's books. This allows different permission algorithms to be applied depending on the category of the request. Some or all of the above processing in the permission unit may be performed using AI, for example, or not using AI. For example, the permission unit can input the category of the request into a generating AI and have the generating AI execute the application of different permission algorithms.
[0099] The permission unit can estimate the user's emotions and determine the priority of permissions based on the estimated emotions. For example, if the user is excited, the permission unit can set a higher priority for permission. If the user is relaxed, the permission unit can set the priority to normal. Furthermore, if the user is stressed, the permission unit can set a lower priority for permission and prioritize other requests. This allows the permission priority to be determined according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the permission unit may be performed using AI or not using AI. For example, the permission unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0100] The authorization unit can determine the order of authorizations based on when the requests were submitted. For example, the authorization unit can prioritize requests submitted early. It can also prioritize requests of high urgency. Furthermore, the authorization unit can adjust the priority of authorizations in stages according to the submission time. This allows the authorization unit to determine the priority of authorizations based on when the requests were submitted. Some or all of the above processing in the authorization unit may be performed using AI, for example, or not using AI. For example, the authorization unit can input the submission time of the requests into a generating AI and have the generating AI determine the order of authorizations.
[0101] The authorization unit can adjust the order of authorizations based on the relevance of the requests. For example, the authorization unit can prioritize authorizations if the content of the request aligns with the library's policies. It can also prioritize authorizations if the content of the request is highly relevant to other requests. Furthermore, the authorization unit can adjust the order of authorizations in stages according to the relevance of the requests. This allows the order of authorizations to be adjusted based on the relevance of the requests. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the authorization order.
[0102] The notification unit can estimate the user's emotions and adjust the notification method based on the estimated emotions. For example, if the user is relaxed, the notification unit can use calm language for notifications. If the user is excited, the notification unit can use lively language for notifications. Furthermore, if the user is stressed, the notification unit can use calm language for notifications. This allows the notification method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0103] The notification unit can adjust the details of the notification based on the importance of the request when it sends a notification. For example, the notification unit can send a detailed notification for high-importance requests. It can also send a simplified notification for low-importance requests. Furthermore, the notification unit can adjust the level of detail of the notification in stages according to the importance. This allows the level of detail of the notification to be adjusted based on the importance of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the notification.
[0104] The notification unit can apply different notification methods depending on the category of the request when sending a notification. For example, the notification unit can apply a specific notification algorithm to requests for novels. It can also apply a different notification algorithm to requests for academic books. Furthermore, it can apply yet another different notification algorithm to requests for children's books. This allows different notification algorithms to be applied depending on the category of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the category of the request into a generating AI and have the generating AI execute the application of different notification algorithms.
[0105] The notification unit can estimate the user's emotions and adjust the length of the notification based on the estimated emotions. For example, if the user is relaxed, the notification unit can provide a detailed notification. If the user is in a hurry, the notification unit can provide a concise notification. Furthermore, if the user is stressed, the notification unit can provide a notification of appropriate length. This allows the length of the notification to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using AI or not using AI. For example, the notification unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0106] The notification unit can determine the order of notifications based on when the requests were submitted. For example, the notification unit can prioritize notifications for requests submitted early. It can also prioritize notifications for urgent requests. Furthermore, the notification unit can adjust the priority of notifications in stages according to the submission time. This allows the notification unit to determine the priority of notifications based on when the requests were submitted. Some or all of the above processing in the notification unit may be performed using AI, for example, or not using AI. For example, the notification unit can input the submission time of the requests into a generating AI and have the generating AI determine the order of notifications.
[0107] The notification unit can adjust the order of notifications based on the relevance of the requests when it sends notifications. For example, the notification unit can prioritize notifications if the content of the request is consistent with the library's policies. It can also prioritize notifications if the content of the request is highly relevant to other requests. Furthermore, the notification unit can adjust the order of notifications in stages according to the relevance of the requests. This allows the order of notifications to be adjusted based on the relevance of the requests. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the order of notifications.
[0108] The analysis unit can estimate the user's emotions and adjust the analysis method based on the estimated emotions. For example, if the user is relaxed, the analysis unit can perform a detailed analysis. If the user is in a hurry, the analysis unit can perform a simplified analysis. Furthermore, if the user is stressed, the analysis unit can perform an analysis with an appropriate level of detail. This allows the analysis method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or not using AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0109] The analysis unit can adjust the level of analysis based on the importance of the request during the analysis. For example, the analysis unit can perform a detailed analysis for high-importance requests. It can also perform a simplified analysis for low-importance requests. Furthermore, the analysis unit can adjust the level of detail of the analysis in stages according to the importance. This allows the level of detail of the analysis to be adjusted based on the importance of the request. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the analysis.
[0110] The analysis unit can apply different analysis methods depending on the category of the request during analysis. For example, the analysis unit can apply a specific analysis algorithm to a request for a novel. It can also apply a different analysis algorithm to a request for an academic book. Furthermore, it can apply yet another analysis algorithm to a request for a children's book. This allows for the application of different analysis algorithms depending on the category of the request. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the request category into a generating AI and have the generating AI execute the application of different analysis algorithms.
[0111] The analysis unit can estimate the user's emotions and determine the priority of analysis based on the estimated emotions. For example, if the user is excited, the analysis unit can set a higher priority for analysis. If the user is relaxed, the analysis unit can set the priority to normal. Furthermore, if the user is stressed, the analysis unit can set a lower priority for analysis and prioritize other requests. This allows the analysis priority to be determined according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or not using AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0112] The analysis unit can determine the order of analysis based on the submission timing of requests during the analysis process. For example, the analysis unit can prioritize the analysis of requests submitted early. It can also prioritize the analysis of requests of high urgency. Furthermore, the analysis unit can adjust the priority of analysis in stages according to the submission timing. This allows the analysis priority to be determined based on the submission timing of requests. Some or all of the above processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the submission timing of requests into a generating AI and have the generating AI determine the order of analysis.
[0113] The analysis unit can adjust the order of analysis based on the relevance of requests during the analysis process. For example, the analysis unit can prioritize the analysis of requests if their content aligns with the library's policies. It can also prioritize the analysis of requests if their content is highly relevant to other requests. Furthermore, the analysis unit can adjust the order of analysis in stages according to the relevance of the requests. This allows the order of analysis to be adjusted based on the relevance of the requests. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the relevance of requests into a generating AI and have the generating AI perform the adjustment of the analysis order.
[0114] The generation unit can estimate the user's emotions and adjust the way it expresses the generated voice based on the estimated emotions. For example, if the user is relaxed, the generation unit can generate a calm voice. If the user is excited, the generation unit can generate a lively voice. Furthermore, if the user is stressed, the generation unit can generate a calm voice. This allows the generation unit to adjust the way it expresses the generated voice according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input user emotion data into the generation AI and have the generation AI perform emotion estimation.
[0115] The generation unit can adjust the details of the generated audio based on the importance of the request during generation. For example, the generation unit can generate detailed audio for high-importance requests. It can also generate simplified audio for low-importance requests. Furthermore, the generation unit can adjust the level of detail of the generated audio in stages according to the importance. This allows the level of detail of the generated audio to be adjusted based on the importance of the request. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the request into the generation AI and have the generation AI perform the adjustment of the level of detail of the generated audio.
[0116] The generation unit can apply different generation methods depending on the category of the request during generation. For example, the generation unit can apply a specific generation algorithm to a request for a novel. It can also apply a different generation algorithm to a request for an academic book. Furthermore, it can apply yet another generation algorithm to a request for a children's book. This allows for the application of different generation algorithms depending on the category of the request. Some or all of the above-described processes in the generation unit may be performed using, for example, AI, or without AI. For example, the generation unit can input the category of the request into a generation AI and cause the generation AI to apply different generation algorithms.
[0117] The generation unit can estimate the user's emotions and determine the priority of the voices to be generated based on the estimated emotions. For example, if the user is excited, the generation unit can set a higher priority for the voices to be generated. If the user is relaxed, the generation unit can set the priority of the voices to be generated as usual. Furthermore, if the user is stressed, the generation unit can set a lower priority for the voices to be generated and prioritize other requests. This allows the priority of the voices to be generated to be determined according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input user emotion data into the generation AI and have the generation AI perform emotion estimation.
[0118] The generation unit can determine the order of generation based on the submission timing of requests during generation. For example, the generation unit can prioritize generating audio for requests submitted early. It can also prioritize generating audio for urgent requests. Furthermore, the generation unit can adjust the generation priority in stages according to the submission timing. This allows the generation priority to be determined based on the submission timing of requests. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the submission timing of requests into a generation AI and have the generation AI determine the order of generation.
[0119] The generation unit can adjust the generation order based on the relevance of the requests during generation. For example, the generation unit can prioritize generation if the content of the request matches the library's policies. It can also prioritize generation if the content of the request is highly relevant to other requests. Furthermore, the generation unit can adjust the generation order in stages according to the relevance of the requests. This allows the generation order to be adjusted based on the relevance of the requests. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the relevance of the requests into a generation AI and have the generation AI perform the adjustment of the generation order.
[0120] The service provider can estimate the user's emotions and adjust the way it delivers audio based on the estimated emotions. For example, if the user is relaxed, the service provider can deliver audio in a calm voice. If the user is excited, the service provider can deliver audio in an energetic voice. Furthermore, if the user is stressed, the service provider can deliver audio in a calm voice. This allows the service provider to adjust the way it delivers audio according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0121] The service provider can adjust the details of the service provided based on the importance of the request. For example, the service provider can provide detailed audio for high-priority requests. It can also provide simplified audio for low-priority requests. Furthermore, the service provider can adjust the level of detail of the service in stages according to the importance. This allows the level of detail of the service to be adjusted based on the importance of the request. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the service.
[0122] The delivery unit can apply different delivery methods depending on the category of the request. For example, the delivery unit can apply a specific delivery algorithm to requests for novels. It can also apply a different delivery algorithm to requests for academic books. Furthermore, it can apply yet another delivery algorithm to requests for children's books. This allows different delivery algorithms to be applied depending on the category of the request. Some or all of the above processing in the delivery unit may be performed using AI, for example, or without AI. For example, the delivery unit can input the category of the request into a generating AI and have the generating AI execute the application of different delivery algorithms.
[0123] The service provider can estimate the user's emotions and determine the priority of the audio to be provided based on the estimated emotions. For example, if the user is excited, the service provider can set a higher priority for the audio to be provided. If the user is relaxed, the service provider can set the priority of the audio to be provided as usual. Furthermore, if the user is stressed, the service provider can set a lower priority for the audio to be provided and prioritize other requests. This allows the service provider to determine the priority of the audio to be provided according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI or not using AI. For example, the service provider can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0124] The service provider can determine the order of service provision based on the submission date of the requests. For example, the service provider can prioritize providing audio for requests submitted early. It can also prioritize providing audio for urgent requests. Furthermore, the service provider can adjust the priority of service provision in stages according to the submission date. This allows the service provider to determine the priority of service provision based on the submission date of the requests. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the submission date of the requests into a generating AI and have the generating AI determine the order of service provision.
[0125] The service provider can adjust the order of deliveries based on the relevance of the requests at the time of delivery. For example, the service provider can prioritize deliveries if the content of the request matches the library's policy. It can also prioritize deliveries if the content of the request is highly relevant to other requests. Furthermore, the service provider can adjust the order of deliveries in stages according to the relevance of the requests. This allows the order of deliveries to be adjusted based on the relevance of the requests. Some or all of the above processing in the service provider may be performed using AI, for example, or not using AI. For example, the service provider can input the relevance of the requests into a generating AI and have the generating AI perform the adjustment of the order of deliveries.
[0126] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0127] The reception unit can estimate the user's emotions and adjust the timing of request acceptance based on the estimated emotions. For example, if the user is relaxed, the reception unit can accept the request immediately. If the user is stressed, the reception unit can slightly delay accepting the request to give the user time to calm down. Furthermore, if the user is in a hurry, the reception unit can accept the request quickly and start processing immediately. This allows the timing of request acceptance to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the reception unit may be performed using AI or not using AI. For example, the reception unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0128] The service provider can estimate the user's emotions and adjust the way it delivers audio based on the estimated emotions. For example, if the user is relaxed, the service provider can deliver audio in a calm voice. If the user is excited, the service provider can deliver audio in an energetic voice. Furthermore, if the user is stressed, the service provider can deliver audio in a calm voice. This allows the service provider to adjust the way it delivers audio according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the service provider may be performed using AI, for example, or without AI. For example, the service provider can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0129] The notification unit can estimate the user's emotions and adjust the notification method based on the estimated emotions. For example, if the user is relaxed, the notification unit can use calm language for notifications. If the user is excited, the notification unit can use lively language for notifications. Furthermore, if the user is stressed, the notification unit can use calm language for notifications. This allows the notification method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0130] The analysis unit can estimate the user's emotions and adjust the analysis method based on the estimated emotions. For example, if the user is relaxed, the analysis unit can perform a detailed analysis. If the user is in a hurry, the analysis unit can perform a simplified analysis. Furthermore, if the user is stressed, the analysis unit can perform an analysis with an appropriate level of detail. This allows the analysis method to be adjusted according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or not using AI. For example, the analysis unit can input user emotion data into a generative AI and have the generative AI perform emotion estimation.
[0131] The generation unit can estimate the user's emotions and adjust the way it expresses the generated voice based on the estimated emotions. For example, if the user is relaxed, the generation unit can generate a calm voice. If the user is excited, the generation unit can generate a lively voice. Furthermore, if the user is stressed, the generation unit can generate a calm voice. This allows the generation unit to adjust the way it expresses the generated voice according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the generation unit may be performed using AI, for example, or not using AI. For example, the generation unit can input user emotion data into the generation AI and have the generation AI perform emotion estimation.
[0132] The reception desk can analyze the user's past request history and select the most appropriate reception method. For example, the reception desk can prioritize requests for book genres that the user has frequently requested in the past. It can also prioritize suggesting request methods (voice, text, etc.) that the user has used in the past. Furthermore, based on the user's past request history, the reception desk can select the most suitable reception method for a particular time period if there are many requests during that time. This allows the reception desk to select the optimal reception method based on the user's past request history. Some or all of the above processing in the reception desk may be performed using AI, for example, or not. For example, the reception desk can input the user's past request history into a generating AI and have the generating AI select the optimal reception method.
[0133] The authorization unit can adjust the details of the authorization based on the importance of the request when granting authorization. For example, the authorization unit can perform a detailed authorization procedure for high-importance requests. It can also perform a simplified authorization procedure for low-importance requests. Furthermore, the authorization unit can adjust the level of authorization detail in stages according to the importance. This allows the level of authorization detail to be adjusted based on the importance of the request. Some or all of the above processing in the authorization unit may be performed using AI, for example, or without AI. For example, the authorization unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of authorization detail.
[0134] The notification unit can adjust the details of the notification based on the importance of the request when it sends a notification. For example, the notification unit can send a detailed notification for high-importance requests. It can also send a simplified notification for low-importance requests. Furthermore, the notification unit can adjust the level of detail of the notification in stages according to the importance. This allows the level of detail of the notification to be adjusted based on the importance of the request. Some or all of the above processing in the notification unit may be performed using AI, for example, or without AI. For example, the notification unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the notification.
[0135] The analysis unit can adjust the level of analysis based on the importance of the request during the analysis. For example, the analysis unit can perform a detailed analysis for high-importance requests. It can also perform a simplified analysis for low-importance requests. Furthermore, the analysis unit can adjust the level of detail of the analysis in stages according to the importance. This allows the level of detail of the analysis to be adjusted based on the importance of the request. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can input the importance of the request into a generating AI and have the generating AI perform the adjustment of the level of detail of the analysis.
[0136] The generation unit can adjust the details of the generated audio based on the importance of the request during generation. For example, the generation unit can generate detailed audio for high-importance requests. It can also generate simplified audio for low-importance requests. Furthermore, the generation unit can adjust the level of detail of the generated audio in stages according to the importance. This allows the level of detail of the generated audio to be adjusted based on the importance of the request. Some or all of the above processing in the generation unit may be performed using AI, for example, or without AI. For example, the generation unit can input the importance of the request into the generation AI and have the generation AI perform the adjustment of the level of detail of the generated audio.
[0137] The following briefly describes the processing flow for example form 2.
[0138] Step 1: The reception desk receives requests from visually impaired individuals to have books converted into audiobooks. Visually impaired individuals submit requests, for example, through an app, and the requests are sent to the library's system. Step 2: The authorization department grants permission for AI voice translation based on the request received by the reception department. The authorization department reviews the content of the request and determines whether AI voice translation is possible. Step 3: The notification unit notifies the visually impaired person's app of the information permitted by the permission unit. The notification unit sends a notification to the app to inform the visually impaired person that the request has been permitted. Step 4: The analysis unit analyzes the book's content based on the information notified by the notification unit. The analysis unit analyzes the requested book's content and extracts information for generating intonation and sound effects. Step 5: The generation unit generates intonation and sound effects based on the analysis performed by the analysis unit. The generation unit uses AI to add intonation and create and apply sound effects. Step 6: The providing unit provides the audio generated by the generating unit to visually impaired individuals. The providing unit provides the generated audio to visually impaired individuals through the app.
[0139] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0140] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0141] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0142] Each of the multiple elements described above, including the reception unit, permission unit, notification unit, analysis unit, generation unit, and provision unit, is implemented, for example, in at least one of the smart device 14 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart device 14 and is used when a visually impaired person requests an audio transcription of a book from the library. The permission unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which verifies the content of the request and determines whether AI audio transcription is possible. The notification unit is implemented, for example, by the control unit 46A of the smart device 14, which sends a notification to inform the visually impaired person that the request has been permitted. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the content of the requested book. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which generates intonation and sound effects. The provision unit is implemented, for example, by the control unit 46A of the smart device 14, which provides the generated audio to the visually impaired person. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0143] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0144] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0145] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0146] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0147] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0148] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0149] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0150] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0151] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0152] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0153] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0154] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0155] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0156] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0157] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0158] Each of the multiple elements described above, including the reception unit, permission unit, notification unit, analysis unit, generation unit, and provision unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the smart glasses 214 and is used when a visually impaired person requests an audio transcription of a book from the library. The permission unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which verifies the content of the request and determines whether AI audio transcription is possible. The notification unit is implemented, for example, by the control unit 46A of the smart glasses 214, which sends a notification to inform the visually impaired person that the request has been permitted. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the content of the requested book. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which generates intonation and sound effects. The provision unit is implemented, for example, by the control unit 46A of the smart glasses 214, which provides the generated audio to the visually impaired person. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0159] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0160] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0161] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0162] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0163] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0164] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0165] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0166] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0167] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0168] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0169] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0170] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0171] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0172] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0173] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0174] Each of the multiple elements described above, including the reception unit, authorization unit, notification unit, analysis unit, generation unit, and provision unit, is implemented, for example, in at least one of the headset terminal 314 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the headset terminal 314 and is used when a visually impaired person requests the library to transcribe a book into audio. The authorization unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which verifies the content of the request and determines whether AI transcription is possible. The notification unit is implemented, for example, by the control unit 46A of the headset terminal 314, which sends a notification to inform the visually impaired person that the request has been authorized. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the content of the requested book. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which generates intonation and sound effects. The provision unit is implemented, for example, by the control unit 46A of the headset terminal 314, which provides the generated audio to the visually impaired person. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0175] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0176] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0177] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0178] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0179] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0180] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0181] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0182] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0183] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0184] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0185] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0186] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0187] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0188] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0189] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0190] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0191] Each of the multiple elements described above, including the reception unit, authorization unit, notification unit, analysis unit, generation unit, and provision unit, is implemented, for example, by at least one of the robot 414 and the data processing unit 12. For example, the reception unit is implemented by the control unit 46A of the robot 414 and is used when a visually impaired person requests the library to transcribe a book into audio. The authorization unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which verifies the content of the request and determines whether AI transcription is possible. The notification unit is implemented, for example, by the control unit 46A of the robot 414, which sends a notification to inform the visually impaired person that the request has been authorized. The analysis unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which analyzes the content of the requested book. The generation unit is implemented, for example, by the identification processing unit 290 of the data processing unit 12, which generates intonation and sound effects. The provision unit is implemented, for example, by the control unit 46A of the robot 414, which provides the generated audio to the visually impaired person. The correspondence between each part and the device or control unit is not limited to the examples described above, and various modifications are possible.
[0192] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0193] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0194] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0195] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0196] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0197] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0198] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0199] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0200] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0201] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0202] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0203] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0204] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0205] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0206] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0207] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0208] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0209] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0210] (Note 1) A reception desk where visually impaired people request audio versions of books from the library, An authorization unit grants permission for AI voice translation based on the request received by the aforementioned reception unit, A notification unit that notifies the app of a visually impaired person of the information authorized by the aforementioned authorization unit, An analysis unit that analyzes the contents of the book based on the information notified by the notification unit, A generation unit that generates sound effects based on the content analyzed by the analysis unit, The system includes a providing unit that provides the sound generated by the generation unit to a visually impaired person. A system characterized by the following features. (Note 2) The aforementioned supply unit is, It features functions to change speaking speed, voice pitch, and gender. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned supply unit is, Features a language selection function. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned analysis unit, Generate sound effects according to the scene in the story. The system described in Appendix 1, characterized by the features described herein. (Note 5) The generating unit is Performs realistic reading aloud based on generated audio. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned supply unit is, It features a function that allows users to share their opinions on books they have finished reading with AI. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned supply unit is, The AI has the ability to understand the content of a book and offer its opinion. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned supply unit is, It has a function to summarize and convey reader feedback found on the internet. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned reception unit is It estimates the user's emotions and adjusts the timing of request acceptance based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned reception unit is Analyze the user's past request history and select the most appropriate method of processing the request. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned reception unit is When a request is received, filtering is performed based on the user's current reading habits and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned reception unit is It estimates the user's emotions and determines the priority of requests to accept based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned reception unit is When receiving a request, the system prioritizes requests that are highly relevant based on the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned reception unit is When a request is received, the system analyzes the user's social media activity and accepts relevant requests. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned authorization unit is, It estimates the user's emotions and adjusts the permission criteria based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned authorization unit is, When granting permission, adjust the permission details based on the importance of the request. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned authorization unit is, When granting permission, different permission methods are applied depending on the category of the request. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned authorization unit is, It estimates the user's emotions and determines permission priorities based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned authorization unit is, When granting permission, the order of approvals will be determined based on when the request was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned authorization unit is, When granting permissions, adjust the order of permissions based on the relevance of the requests. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned notification unit, It estimates the user's emotions and adjusts the notification method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned notification unit, When a notification is sent, the notification details are adjusted based on the importance of the request. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned notification unit, When sending notifications, different notification methods will be applied depending on the request category. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned notification unit, It estimates the user's emotions and adjusts the length of notifications based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned notification unit, When notifying, the order of notifications will be determined based on when the request was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned notification unit, When sending notifications, the order of notifications will be adjusted based on the relevance of the requests. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the analysis method based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned analysis unit, During analysis, the analysis details are adjusted based on the importance of the request. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned analysis unit, During analysis, different analysis methods are applied depending on the request category. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned analysis unit, The system estimates the user's emotions and determines the priority of analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned analysis unit, During analysis, the order of analysis is determined based on when the request was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of the requests. The system described in Appendix 1, characterized by the features described herein. (Note 33) The generating unit is It estimates the user's emotions and adjusts the method of generating speech based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The generating unit is During generation, the generation details are adjusted based on the importance of the request. The system described in Appendix 1, characterized by the features described herein. (Note 35) The generating unit is During generation, different generation methods are applied depending on the request category. The system described in Appendix 1, characterized by the features described herein. (Note 36) The generating unit is It estimates the user's emotions and determines the priority of the voice output to be generated based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 37) The generating unit is During generation, the generation order is determined based on when the request was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 38) The generating unit is During generation, the generation order is adjusted based on the relevance of the requests. The system described in Appendix 1, characterized by the features described herein. (Note 39) The aforementioned supply unit is, It estimates the user's emotions and adjusts the way it delivers audio based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 40) The aforementioned supply unit is, When providing the service, we will adjust the details of the service based on the importance of the request. The system described in Appendix 1, characterized by the features described herein. (Note 41) The aforementioned supply unit is, When providing a service, different delivery methods will be applied depending on the category of the request. The system described in Appendix 1, characterized by the features described herein. (Note 42) The aforementioned supply unit is, It estimates the user's emotions and determines the priority of the audio content to be delivered based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 43) The aforementioned supply unit is, When providing the service, the order of provision will be determined based on when the request was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 44) The aforementioned supply unit is, When providing the service, we adjust the order of delivery based on the relevance of the requests. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0211] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A reception desk where visually impaired people request audio versions of books from the library, A licensing unit grants permission for AI voice transcription based on a request received by the aforementioned reception unit, A notification unit that notifies the app of a visually impaired person of the information authorized by the aforementioned authorization unit, An analysis unit that analyzes the contents of the book based on the information notified by the notification unit, A generation unit that generates sound effects based on the content analyzed by the analysis unit, The system includes a providing unit that provides the sound generated by the generation unit to a visually impaired person. A system characterized by the following features.
2. The aforementioned supply unit is, It features functions to change speaking speed, voice pitch, and gender. The system according to feature 1.
3. The aforementioned supply unit is, Features a language selection function. The system according to feature 1.
4. The aforementioned analysis unit, Generate sound effects according to the scene in the story. The system according to feature 1.
5. The generating unit is Performs realistic reading aloud based on generated audio. The system according to feature 1.
6. The aforementioned supply unit is, It features a function that allows users to share their opinions on books they have finished reading with AI. The system according to feature 1.
7. The aforementioned supply unit is, The AI has the ability to understand the content of a book and offer its opinion. The system according to feature 1.
8. The aforementioned supply unit is, It has a function to summarize and convey reader feedback found on the internet. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A