system
The system facilitates easy recording and review of daily thoughts and actions by converting audio to text, classifying, and using speech synthesis and virtual reality to recreate past events, enhancing personal reflection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing systems face difficulties in easily recording daily thoughts and actions and reviewing past events.
A system comprising a recording unit, text conversion unit, storage unit, classification unit, and playback unit that converts audio data into text, classifies and stores it, and explains past events using speech synthesis and virtual reality to recreate past life.
Enables users to easily record and review their daily thoughts and actions, providing a vivid recall of past events and promoting reflection on personal development.
Smart Images

Figure 2026072452000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0006] , , , , , ,
[0005] , , ,
[0003] , , , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the prior art, there was a problem that it was difficult to easily record daily thoughts and actions and review past events.
[0005] The system according to the embodiment aims to easily record daily thoughts and actions and review past events.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a recording unit, a text conversion unit, a storage unit, a classification unit, an explanation unit, and a playback unit. The recording unit records audio data. The text conversion unit converts the audio data recorded by the recording unit into text. The storage unit stores the data converted into text by the text conversion unit. The classification unit classifies the data stored in the storage unit. The explanation unit explains past events based on the data classified by the classification unit. The playback unit reconstructs past life based on the data classified by the classification unit. [Effects of the Invention]
[0007] The system according to this embodiment allows users to easily record their daily thoughts and actions and reflect on past events. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between a plurality of computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) An AI system according to an embodiment of the present invention is a system that allows users to easily record their daily thoughts and actions and utilize that information. The purpose of this system is to help users vividly recall their past selves and provide them with a moment of peace. For example, a user records daily events and thoughts in voice. This voice data is input into the AI system. Next, the AI system converts the input voice data into text using natural language processing technology. The converted data is understood and classified as the user's statements, emotions, and thought patterns, and stored in a dedicated database. Photos and video data taken on that day are also stored in association with the data. When a user asks, "What was I doing last year?", the AI system explains the events and thoughts from that time. Also, when a user asks, "Let me see myself in 2020," the AI system plays back an automatically generated life in 2020. In this way, the user can reconnect with their past self. Through this system, the user can look back on their past self and gain a moment of peace. Furthermore, by re-examining their past and present selves, the steps they take towards tomorrow may be different than before. This allows AI systems to easily record users' daily thoughts and actions and utilize that information.
[0029] The AI system according to this embodiment comprises a recording unit, a text conversion unit, a storage unit, a classification unit, an explanation unit, and a playback unit. The recording unit records the user's daily events and thoughts in audio. For example, the recording unit records what the user says using a microphone and saves it as audio data. The recording unit can also record audio using a device such as a smartphone or tablet. Furthermore, the recording unit can convert audio into text data in real time using speech recognition technology. For example, the recording unit records what the user says using a high-precision microphone and saves it as audio data. By using a smartphone or tablet, the user can record audio anytime, anywhere. By using speech recognition technology, audio can be converted into text data in real time and saved immediately. The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing technology. For example, the text conversion unit converts audio data into text data using speech recognition technology. Furthermore, the text conversion unit can also analyze the content of audio data using machine learning algorithms and convert it into text data. Furthermore, the text conversion unit can understand the context of the audio data and convert it into appropriate text data. For example, by using speech recognition technology, highly accurate text data can be generated. By using machine learning algorithms, the content of speech data can be analyzed and more accurate text data can be generated. By understanding the context of speech data, text data can be generated as natural-sounding sentences. The storage unit stores the data converted into text by the text conversion unit in a dedicated database. The storage unit can, for example, save the text data to cloud storage. The storage unit can also save it to local storage. Furthermore, the storage unit can ensure data security by regularly backing up the data. For example, by using cloud storage, data can be stored securely and accessed at any time. By using local storage, data can be stored offline and accessed even without an internet connection. By regularly backing up the data, data loss can be prevented and security can be ensured.The classification unit classifies data stored in the storage unit based on the user's statements, emotions, and thought patterns. The classification unit can classify data using, for example, machine learning algorithms. It can also analyze data using natural language processing techniques and classify it into appropriate categories. Furthermore, the classification unit can understand the user's emotions and thought patterns and classify data based on that understanding. For example, machine learning algorithms can efficiently classify large amounts of data. Natural language processing techniques can analyze the content of data and classify it into appropriate categories. Understanding the user's emotions and thought patterns allows for more accurate classification. The explanation unit explains past events based on the data classified by the classification unit. For example, the explanation unit can explain text data using speech synthesis technology. It can also display text data on a display device. Furthermore, the explanation unit can replay past events as video data. For example, speech synthesis technology can explain text data in natural speech. A display device can visually display and provide text data to the user. Playing video data can visually recreate and provide past events to the user. The playback unit reconstructs past life based on data classified by the classification unit. The playback unit can, for example, play audio data. It can also play video data. Furthermore, the playback unit can recreate past life using virtual reality (VR) technology. For example, by playing audio data, past events can be recreated aurally. By playing video data, past events can be recreated visually. By using virtual reality technology, past life can be recreated in an immersive way and provided to the user. As a result, the AI system according to this embodiment can easily record the user's daily thoughts and actions and utilize that information.
[0030] The recording unit records the user's daily events and thoughts in audio format. For example, the recording unit records what the user says using a microphone and saves it as audio data. Specifically, the recording unit uses a high-sensitivity microphone to record the user's voice clearly. This minimizes background noise and improves the quality of the audio data. The recording unit can also record audio using devices such as smartphones and tablets. This allows users to easily record audio even when they are out and about. Furthermore, the recording unit can convert audio to text data in real time using speech recognition technology. For example, the recording unit records what the user says using a high-precision microphone and saves it as audio data. By using a smartphone or tablet, users can record audio anytime, anywhere. Using speech recognition technology, audio can be converted to text data in real time and saved instantly. The speech recognition technology uses a model based on deep learning, which achieves high recognition accuracy. For example, if a user says, "We discussed a new project at today's meeting," the recording unit can instantly convert this audio into text data and save it. This allows the user to easily review the content later.
[0031] The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing techniques. For example, the text conversion unit can use speech recognition technology to convert the audio data into text data. Specifically, speech recognition technology includes a process of analyzing the audio waveform and breaking it down into phonemes and words. This allows for the conversion of audio data into highly accurate text data. The text conversion unit can also use machine learning algorithms to analyze the content of the audio data and convert it into text data. Furthermore, the text conversion unit can understand the context of the audio data and convert it into appropriate text data. For example, using speech recognition technology can generate highly accurate text data. Using machine learning algorithms can analyze the content of the audio data and generate more accurate text data. Understanding the context of the audio data allows for the generation of text data in a natural-sounding format. For example, if a user says, "What are my plans for tomorrow?", the text conversion unit converts this audio into the text data "What are my plans for tomorrow?". Furthermore, the text conversion unit can refer to past conversation history and user utterance patterns to understand the context. This allows for the generation of more natural and accurate text data.
[0032] The storage unit stores the data converted to text by the text conversion unit in a dedicated database. The storage unit can, for example, store text data in cloud storage. Specifically, by using a cloud storage service, data can be securely stored and accessed anytime, anywhere. The storage unit can also store data in local storage. This allows access to data even without an internet connection. Furthermore, the storage unit can ensure data security by regularly backing up the data. For example, using cloud storage allows for secure data storage and access at any time. Using local storage allows for offline data storage and access even without an internet connection. Regularly backing up data prevents data loss and ensures security. The storage unit can also protect the confidentiality of stored data using data encryption technology. For example, it can encrypt data using encryption algorithms such as AES (Advanced Encryption Standard) to protect against unauthorized access. This allows the storage unit to ensure data security and confidentiality, and protect user privacy.
[0033] The classification unit classifies data stored in the storage unit based on the user's statements, emotions, and thought patterns. The classification unit uses, for example, machine learning algorithms to classify data. Specifically, it analyzes the user's statements and emotions and classifies the data into appropriate categories based on that analysis. For example, if a user says, "I had a great time today," the classification unit classifies this statement as a "positive emotion." The classification unit can also analyze data using natural language processing techniques and classify it into appropriate categories. Furthermore, the classification unit can understand the user's emotions and thought patterns and classify data based on that understanding. For example, machine learning algorithms can efficiently classify large amounts of data. Natural language processing techniques can analyze the content of data and classify it into appropriate categories. Understanding the user's emotions and thought patterns allows for more accurate classification. The classification unit can also learn the user's past statements and behavioral patterns and predict future statements and behaviors. This enables the classification unit to provide personalized services tailored to the user's needs and preferences.
[0034] The explanation unit describes past events based on data classified by the classification unit. For example, the explanation unit can explain text data using speech synthesis technology. Specifically, speech synthesis technology uses algorithms to convert text into natural-sounding speech. This allows the user to hear past events in audio. The explanation unit can also display text data on a display device. This allows the user to visually confirm the information. Furthermore, the explanation unit can play back past events as video data. For example, by using speech synthesis technology, text data can be explained in natural-sounding speech. By using a display device, text data can be visually displayed and provided to the user. By playing back video data, past events can be visually recreated and provided to the user. The explanation unit can also provide relevant information based on the user's past statements and actions. For example, if the user asks, "Tell me about what happened during last summer vacation," the explanation unit will provide relevant information in audio and video based on statements and actions from that period. This allows the explanation unit to supplement the user's memory and help them recall past events more vividly.
[0035] The playback unit replays past life events based on data classified by the classification unit. For example, the playback unit can play audio data. Specifically, it can use high-quality speakers to reproduce clear sound. The playback unit can also play video data, allowing users to visually recreate past events. Furthermore, the playback unit can recreate past life events using virtual reality (VR) technology. For example, by playing audio data, past events can be recreated aurally. By playing video data, past events can be recreated visually. Using virtual reality technology, past life can be recreated in an immersive way and provided to the user. The playback unit can also provide relevant information based on the user's past statements and actions. For example, if a user instructs, "Play back what happened on my birthday last year," the playback unit will play audio and video data from that period and provide it to the user. This allows the playback unit to supplement the user's memory and help them recall past events more vividly. The playback unit can also customize the order and content of the data played according to the user's preferences. This allows the playback unit to recreate and present past events in the most optimal way for the user.
[0036] The storage unit can store photos and video data in association with each other. For example, it can store photos and video data taken by the user together with audio data. The storage unit can also store photos and video data in a dedicated database and link it to audio data. Furthermore, the storage unit can store photos and video data in cloud storage and synchronize it with audio data. For example, storing photos and video data taken by the user together with audio data allows for more detailed recording. Storing data in a dedicated database makes data management easier. Storing data in cloud storage ensures secure storage and access at any time. This allows for more detailed recording by storing photos and video data in association with each other.
[0037] The recording unit can analyze the user's past voice recording history and select the optimal recording method. For example, the recording unit can prioritize suggesting recording methods (voice, text, etc.) that the user has preferred in the past. The recording unit can also suggest appropriate recording timing based on the user's past recording frequency. Furthermore, the recording unit can analyze the user's past recording content and suggest recording methods based on specific themes. For example, by prioritizing suggesting recording methods that the user has preferred in the past, it can provide a recording method that suits the user's preferences. By suggesting appropriate recording timing based on the user's past recording frequency, the efficiency of recording can be improved. By analyzing the user's past recording content and suggesting recording methods based on specific themes, more appropriate recordings can be made. Some or all of the above processing in the recording unit may be performed using AI, for example, or not. For example, the recording unit can input the user's past voice recording history into AI and have the AI select the optimal recording method. This allows the recording unit to provide the user with the most suitable recording method by analyzing the past voice recording history.
[0038] The recording unit can filter audio recordings based on the user's current activity and environment. For example, if the user is in a quiet environment, the recording unit will start recording audio. If the user is in a noisy environment, the recording unit can also apply noise cancellation during recording. Furthermore, if the user is moving, the recording unit can wait to record audio until the user has finished moving. For example, if the user is in a quiet environment, starting recording immediately will allow for the recording of clear audio with minimal noise. If the user is in a noisy environment, applying noise cancellation will remove background noise and improve audio quality. If the user is moving, waiting to record audio until the user has finished moving will avoid noise and vibrations during movement, resulting in stable audio recording. Some or all of the above processing in the recording unit may be performed using AI, or not. For example, the recording unit can input the user's current activity and environment data into the AI and have the AI perform the filtering. This allows for more appropriate recording by filtering audio recordings according to the user's activity and environment.
[0039] The recording unit can prioritize recording audio that is highly relevant, taking into account the user's geographical location information during audio recording. For example, if the user is in a specific location (such as a travel destination), the recording unit will prioritize recording audio related to that location. Furthermore, if the user is at home, the recording unit can also record everyday events. Additionally, if the user is attending a specific event (such as a meeting or party), the recording unit can prioritize recording audio related to that event. For example, if the user is in a specific location, prioritizing the recording of audio related to that location allows for a detailed record of events at the travel destination or specific location. If the user is at home, recording everyday events enriches the record of daily life. If the user is attending a specific event, prioritizing the recording of audio related to that event allows for a detailed record of the event. Some or all of the above processing in the recording unit may be performed using AI, or not. For example, the recording unit can input the user's geographical location information into the AI and have the AI prioritize highly relevant audio. This allows for the recording of highly relevant audio by taking geographical location information into consideration.
[0040] The recording unit can analyze the user's social media activity during audio recording and record relevant audio. For example, if the user posts about a specific topic on social media, the recording unit can record audio related to that topic. It can also record audio related to emotional posts by the user. Furthermore, if the user posts about a specific event on social media, the recording unit can record audio related to that event. For example, if the user posts about a specific topic on social media, recording audio related to that topic allows for a record linked to social media activity. If the user makes an emotional post, recording audio related to that emotion allows for a detailed record of emotional changes. If the user posts about a specific event, recording audio related to that event enriches the event record. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI. For example, the recording unit can input the user's social media activity data into AI and have the AI record relevant audio. This allows for the recording of relevant audio by analyzing social media activity.
[0041] The transcription unit can adjust the level of detail in the transcription based on the importance of the audio. For example, the transcription unit will transcribe audio about important events in detail. It can also transcribe audio about everyday events in a simplified manner. Furthermore, it can transcribe audio about emotional events in a detailed manner that reflects the emotions. For example, transcribing audio about important events in detail ensures that important information is not missed. Transcribing audio about everyday events in a simplified manner allows for efficient recording. Transcribing audio about emotional events in a detailed manner that reflects the emotions allows for a detailed recording of emotional changes. Some or all of the above processing in the transcription unit may be performed using AI, for example, or without AI. For example, the transcription unit can input audio data into AI and have the AI perform the transcription detail adjustment based on importance. This allows important information to be transcribed in detail by adjusting the level of detail in the transcription based on the importance of the audio.
[0042] The text conversion unit can apply different text conversion algorithms depending on the audio category during the text conversion process. For example, the text conversion unit can apply a dialogue-style text conversion algorithm to conversational audio. It can also apply a summary-style text conversion algorithm to lecture audio. Furthermore, it can apply an emotion-reflecting text conversion algorithm to emotional audio. For example, applying a dialogue-style text conversion algorithm to conversational audio can generate natural-sounding dialogue. Applying a summary-style text conversion algorithm to lecture audio can generate a summary that extracts the key points. Applying an emotion-reflecting text conversion algorithm to emotional audio can record emotional changes in detail. Some or all of the above processing in the text conversion unit may be performed using AI, for example, or without AI. For example, the text conversion unit can input audio data into AI and have the AI apply a text conversion algorithm according to the category. This allows for more appropriate text conversion by applying different text conversion algorithms depending on the audio category.
[0043] The text conversion unit can determine the priority of transcription based on the recording date of the audio. For example, the text conversion unit may prioritize the transcription of recently recorded audio. It can also prioritize the transcription of audio related to specific events (such as birthdays or anniversaries). Furthermore, the text conversion unit can prioritize the transcription of audio within a period specified by the user. For example, prioritizing the transcription of recently recorded audio allows for the rapid recording of the latest information. Prioritizing the transcription of audio related to specific events enriches the recording of important events. Prioritizing the transcription of audio within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the text conversion unit may be performed using AI, for example, or without AI. For example, the text conversion unit can input audio data into AI and have the AI prioritize transcription based on the recording date. This allows for the prioritization of important audio by determining the transcription priority based on the recording date of the audio.
[0044] The transcription unit can adjust the order of transcription based on the relevance of the audio. For example, the transcription unit can prioritize the transcription of highly relevant audio. It can also postpone the transcription of less relevant audio. Furthermore, the transcription unit can prioritize the transcription of audio related to a specific theme. For example, prioritizing the transcription of highly relevant audio allows for efficient recording of relevant information. Postponing the transcription of less relevant audio allows for priority recording of important information. Prioritizing the transcription of audio related to a specific theme allows for recording in line with the theme. Some or all of the above processing in the transcription unit may be performed using AI, for example, or without AI. For example, the transcription unit can input audio data into AI and have the AI execute the transcription order based on relevance. This allows for prioritizing the transcription of highly relevant audio by adjusting the transcription order based on the relevance of the audio.
[0045] The storage unit can optimize the storage algorithm by referring to previously saved data during the saving process. For example, the storage unit can refer to the format of previously saved data and save it in the same format. It can also refer to the importance of previously saved data and prioritize saving important data. Furthermore, the storage unit can refer to the category of previously saved data and save it within the same category. For example, referring to the format of previously saved data allows for consistent data storage. Referring to the importance of previously saved data allows for prioritizing the saving of important data. Referring to the category of previously saved data makes data organization easier. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input previously saved data into AI and have the AI optimize the storage algorithm. This allows for optimization of the storage algorithm by referring to previously saved data.
[0046] The storage unit can adjust the level of detail in saving data based on its importance. For example, it can save important data in detail. It can also save everyday data in a simplified manner. Furthermore, it can save emotional data in a detailed manner that reflects the emotions involved. For example, saving important data in detail ensures that important information is not missed. Saving everyday data in a simplified manner allows for efficient recording. Saving emotional data in a detailed manner that reflects emotions allows for a detailed record of emotional changes. Some or all of the above processing in the storage unit may be performed using AI, for example, or not. For example, the storage unit can input data into AI and have the AI perform the level of detail in saving based on importance. This allows important data to be saved in detail by adjusting the level of detail in saving based on the importance of the data.
[0047] The storage unit can weight the data to be saved based on when the data was recorded. For example, the storage unit can prioritize saving recently recorded data. It can also prioritize saving data related to specific events (birthdays, anniversaries, etc.). Furthermore, the storage unit can prioritize saving data within a period specified by the user. For example, prioritizing the saving of recently recorded data allows for the rapid recording of the latest information. Prioritizing the saving of data related to specific events enriches the recording of important events. Prioritizing the saving of data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the storage unit may be performed using AI, for example, or not. For example, the storage unit can input data into AI and have the AI perform weighting of the saved data based on the recording date. This allows for prioritizing the saving of important data by weighting the saved data based on the recording date.
[0048] The storage unit can adjust the order of saving data based on its relevance. For example, the storage unit can prioritize saving highly relevant data. It can also postpone saving less relevant data. Furthermore, the storage unit can prioritize saving data related to a specific theme. For example, prioritizing the saving of highly relevant data allows for efficient recording of relevant information. Postponing the saving of less relevant data allows for prioritizing the saving of important information. Prioritizing the saving of data related to a specific theme allows for theme-based recording. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input data into AI and have the AI execute a saving order based on relevance. This allows for prioritizing the saving of highly relevant data by adjusting the saving order based on data relevance.
[0049] The classification unit can improve the accuracy of classification by considering the interrelationships between data during the classification process. For example, the classification unit can group and classify highly relevant data. It can also analyze the interrelationships between data and classify related data together. Furthermore, the classification unit can efficiently classify data by eliminating duplicate data while considering the interrelationships. For example, grouping and classifying highly relevant data allows for efficient recording of related information. Analyzing the interrelationships between data and classifying related data together makes data organization easier. Eliminating duplicate data while considering the interrelationships allows for efficient classification. Some or all of the above processes in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input data into AI and have the AI perform classification accuracy improvements based on interrelationships. This improves classification accuracy by considering the interrelationships between data.
[0050] The classification unit can perform classification while considering the attribute information of the data submitter. For example, the classification unit can classify data based on the user's age and gender. It can also classify data based on the user's occupation and hobbies. Furthermore, it can classify data based on the user's place of residence and cultural background. For example, classifying data based on the user's age and gender allows for attribute-based recording. Classifying data based on the user's occupation and hobbies allows for recording based on interests and concerns. Classifying data based on the user's place of residence and cultural background allows for recording based on region and culture. Some or all of the above processing in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input user attribute information into AI and have the AI perform attribute-based classification. This allows for more appropriate classification by considering the attribute information of the data submitter.
[0051] The classification unit can perform classification while considering the geographical distribution of the data. For example, the classification unit can classify data based on the user's place of residence. It can also classify data based on places the user has visited. Furthermore, it can classify data based on the user's travel destinations. For example, classifying data based on the user's place of residence allows for regionally appropriate records. Classifying data based on places the user has visited allows for more detailed records of visited locations. Classifying data based on the user's travel destinations allows for more detailed travel records. Some or all of the above processing in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input geographical distribution data into AI and have the AI perform classification based on geographical distribution. This allows for more appropriate classification by considering the geographical distribution of the data.
[0052] The classification unit can improve the accuracy of classification by referring to relevant literature during the classification process. For example, the classification unit can adjust the classification criteria for the data by referring to relevant literature. The classification unit can also analyze the interrelationships of the data by referring to relevant literature. Furthermore, the classification unit can eliminate data duplication by referring to relevant literature. For example, by adjusting the classification criteria for the data by referring to relevant literature, the accuracy of classification can be improved. By analyzing the interrelationships of the data by referring to relevant literature, relevant information can be efficiently recorded. By eliminating data duplication by referring to relevant literature, classification can be performed efficiently. Some or all of the above processes in the classification unit may be performed using AI, for example, or not using AI. For example, the classification unit can input relevant literature data into AI and have the AI perform the classification accuracy improvement. This improves the accuracy of classification by referring to relevant literature for the data.
[0053] The explanation unit can adjust the level of detail in its explanations based on the importance of the data. For example, it can provide detailed explanations for important data. It can also provide simplified explanations for everyday data. Furthermore, it can provide detailed explanations that reflect emotions for emotional data. For example, providing detailed explanations for important data ensures that important information is not missed. Simplified explanations for everyday data allow for efficient recording. Detailed explanations that reflect emotions for emotional data allow for detailed recording of emotional changes. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI perform level of detail in the explanations based on importance. This allows for detailed explanations of important data by adjusting the level of detail in the explanations based on the importance of the data.
[0054] The explanation unit can apply different explanation algorithms depending on the data category during explanation. For example, the explanation unit can apply a dialogue-style explanation algorithm to conversational data. It can also apply a summary-style explanation algorithm to lecture data. Furthermore, it can apply an emotion-reflecting explanation algorithm to emotional data. For example, applying a dialogue-style explanation algorithm to conversational data allows for a natural, dialogue-style explanation. Applying a summary-style explanation algorithm to lecture data allows for a summary that extracts the important points. Applying an emotion-reflecting explanation algorithm to emotional data allows for a detailed recording of emotional changes. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI apply a category-appropriate explanation algorithm. This allows for more appropriate explanations by applying different explanation algorithms depending on the data category.
[0055] The explanation unit can determine the priority of explanations based on the recording date of the data during explanation. For example, the explanation unit may prioritize explaining recently recorded data. It can also prioritize explaining data related to specific events (birthdays, anniversaries, etc.). Furthermore, the explanation unit can prioritize explaining data within a period specified by the user. For example, prioritizing recently recorded data allows for the rapid recording of the latest information. Prioritizing data related to specific events enriches the recording of important events. Prioritizing data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI prioritize explanations based on the recording date. This allows for prioritizing explanations based on the recording date of the data, thereby prioritizing the explanation of important data.
[0056] The explanation unit can adjust the order of explanations based on the relevance of the data during the explanation process. For example, the explanation unit can prioritize explaining highly relevant data. It can also postpone explaining less relevant data. Furthermore, the explanation unit can prioritize explaining data related to a specific theme. For example, prioritizing highly relevant data allows for efficient recording of relevant information. Postponing the explanation of less relevant data allows for priority recording of important information. Prioritizing the explanation of data related to a specific theme allows for theme-aligned recording. Some or all of the above processing in the explanation unit may be performed using AI, for example, or without AI. For example, the explanation unit can input data into AI and have the AI execute the explanation order based on relevance. This allows for prioritizing the explanation of highly relevant data by adjusting the order of explanations based on the relevance of the data.
[0057] The playback unit can adjust the level of detail during playback based on the importance of the data. For example, the playback unit will play important data in detail. It can also play everyday data in a simplified manner. Furthermore, the playback unit can play emotional data in a detailed manner that reflects the emotions. For example, by playing important data in detail, important information can be recorded without being missed. By playing everyday data in a simplified manner, recording can be done efficiently. By playing emotional data in a detailed manner that reflects the emotions, changes in emotions can be recorded in detail. Some or all of the above processing in the playback unit may be performed using AI, for example, or not using AI. For example, the playback unit can input data into AI and have the AI perform playback detail based on importance. This allows important data to be played back in detail by adjusting the level of detail of playback based on the importance of the data.
[0058] The playback unit can apply different playback algorithms depending on the data category during playback. For example, the playback unit can apply a dialogue-style playback algorithm to conversational data. It can also apply a summary-style playback algorithm to lecture data. Furthermore, the playback unit can apply an emotion-reflecting playback algorithm to emotional data. For example, applying a dialogue-style playback algorithm to conversational data allows for natural-sounding dialogue playback. Applying a summary-style playback algorithm to lecture data allows for summarization that extracts important points. Applying an emotion-reflecting playback algorithm to emotional data allows for detailed recording of emotional changes. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI apply a playback algorithm according to the category. This allows for more appropriate playback by applying different playback algorithms depending on the data category.
[0059] The playback unit can determine playback priority based on the recording date of the data during playback. For example, the playback unit may prioritize playback of recently recorded data. It can also prioritize playback of data related to specific events (such as birthdays or anniversaries). Furthermore, the playback unit can prioritize playback of data within a period specified by the user. For example, prioritizing playback of recently recorded data allows for the rapid recording of the latest information. Prioritizing playback of data related to specific events enriches the recording of important events. Prioritizing playback of data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI prioritize playback based on the recording date. This allows for the priority of playback of important data by determining playback priority based on the recording date of the data.
[0060] The playback unit can adjust the playback order based on the relevance of the data during playback. For example, the playback unit can prioritize playback of highly relevant data. It can also postpone playback of less relevant data. Furthermore, the playback unit can prioritize playback of data related to a specific theme. For example, prioritizing playback of highly relevant data allows for efficient recording of relevant information. Postponing playback of less relevant data allows for priority recording of important information. Prioritizing playback of data related to a specific theme allows for recording in line with the theme. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI execute a playback order based on relevance. This allows for prioritizing playback of highly relevant data by adjusting the playback order based on the relevance of the data.
[0061] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0062] The recording unit can monitor the user's health condition while recording their voice data, and adjust the frequency and timing of recording accordingly. For example, if the user is tired, the recording unit will reduce the frequency of voice recording and encourage recording when the user is relaxed. If the user is stressed, the recording unit can play relaxing music to reduce stress before recording voice. Furthermore, if the user is in good health, the recording unit can record voice at the normal frequency. This ensures that voice recording is performed at the appropriate time according to the user's health condition.
[0063] The text conversion unit can analyze the user's speaking speed and tone when converting user audio data into text, thereby improving the accuracy of the conversion. For example, if the user is speaking quickly, the text conversion unit takes the speaking speed into consideration. Furthermore, if the user is speaking emotionally, the unit can reflect that tone in the text. Additionally, if the user is speaking slowly, the text conversion unit can adjust to that pace. This enables highly accurate text conversion tailored to the user's speaking style.
[0064] The storage unit can prioritize saving user audio data based on its importance. For example, audio data about important events can be saved first, while audio data about everyday events can be saved later. Audio data about emotional events can also be saved first to record emotional changes in detail. Furthermore, audio data related to specific themes specified by the user can be saved first. This allows for the priority saving of important data and efficient data management.
[0065] The classification unit can classify user voice data based on the user's lifestyle and hobbies. For example, if a user enjoys sports, it can prioritize classifying sports-related voice data. Similarly, if a user enjoys traveling, it can prioritize classifying travel-related voice data. Furthermore, if a user has a specific hobby (e.g., cooking or reading), it can prioritize classifying voice data related to that hobby. This allows for appropriate data classification tailored to the user's lifestyle and hobbies.
[0066] The explanation unit can adjust the content of the explanation by referring to the user's past behavior patterns and statements when explaining the user's voice data. For example, it can provide detailed explanations for themes that the user has frequently mentioned in the past. It can also prioritize explaining information related to topics that the user has shown interest in in the past. Furthermore, it can reflect the emotions the user felt when talking about events in the past. This allows for appropriate explanations based on the user's past behavior and statements.
[0067] The following briefly describes the processing flow for example form 1.
[0068] Step 1: The recording unit records the user's daily events and thoughts in audio format. For example, it records what the user says using a microphone and saves it as audio data. Audio can also be recorded using devices such as smartphones and tablets. Furthermore, speech recognition technology can be used to convert audio into text data in real time. Step 2: The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing technology. For example, it uses speech recognition technology or machine learning algorithms to convert the audio data into text data, understands the context of the audio data, and generates appropriate text data. Step 3: The storage unit saves the data converted to text by the text conversion unit to a dedicated database. For example, it can be saved to cloud storage or local storage, and data security is ensured by regularly backing up the data. Step 4: The classification unit classifies the data stored in the storage unit based on the user's statements, emotions, and thought patterns. For example, it analyzes the data using machine learning algorithms and natural language processing techniques to classify it into the appropriate category. Step 5: The explanation unit explains past events based on the data classified by the classification unit. For example, it can explain text data using speech synthesis technology, display text data on a display device, or play it back as video data. Step 6: The playback unit recreates past life based on the data classified by the classification unit. For example, it can play audio or video data, or recreate past life using virtual reality (VR) technology.
[0069] (Example of form 2) An AI system according to an embodiment of the present invention is a system that allows users to easily record their daily thoughts and actions and utilize that information. The purpose of this system is to help users vividly recall their past selves and provide them with a moment of peace. For example, a user records daily events and thoughts in voice. This voice data is input into the AI system. Next, the AI system converts the input voice data into text using natural language processing technology. The converted data is understood and classified as the user's statements, emotions, and thought patterns, and stored in a dedicated database. Photos and video data taken on that day are also stored in association with the data. When a user asks, "What was I doing last year?", the AI system explains the events and thoughts from that time. Also, when a user asks, "Let me see myself in 2020," the AI system plays back an automatically generated life in 2020. In this way, the user can reconnect with their past self. Through this system, the user can look back on their past self and gain a moment of peace. Furthermore, by re-examining their past and present selves, the steps they take towards tomorrow may be different than before. This allows AI systems to easily record users' daily thoughts and actions and utilize that information.
[0070] The AI system according to this embodiment comprises a recording unit, a text conversion unit, a storage unit, a classification unit, an explanation unit, and a playback unit. The recording unit records the user's daily events and thoughts in audio. For example, the recording unit records what the user says using a microphone and saves it as audio data. The recording unit can also record audio using a device such as a smartphone or tablet. Furthermore, the recording unit can convert audio into text data in real time using speech recognition technology. For example, the recording unit records what the user says using a high-precision microphone and saves it as audio data. By using a smartphone or tablet, the user can record audio anytime, anywhere. By using speech recognition technology, audio can be converted into text data in real time and saved immediately. The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing technology. For example, the text conversion unit converts audio data into text data using speech recognition technology. Furthermore, the text conversion unit can also analyze the content of audio data using machine learning algorithms and convert it into text data. Furthermore, the text conversion unit can understand the context of the audio data and convert it into appropriate text data. For example, by using speech recognition technology, highly accurate text data can be generated. By using machine learning algorithms, the content of speech data can be analyzed and more accurate text data can be generated. By understanding the context of speech data, text data can be generated as natural-sounding sentences. The storage unit stores the data converted into text by the text conversion unit in a dedicated database. The storage unit can, for example, save the text data to cloud storage. The storage unit can also save it to local storage. Furthermore, the storage unit can ensure data security by regularly backing up the data. For example, by using cloud storage, data can be stored securely and accessed at any time. By using local storage, data can be stored offline and accessed even without an internet connection. By regularly backing up the data, data loss can be prevented and security can be ensured.The classification unit classifies data stored in the storage unit based on the user's statements, emotions, and thought patterns. The classification unit can classify data using, for example, machine learning algorithms. It can also analyze data using natural language processing techniques and classify it into appropriate categories. Furthermore, the classification unit can understand the user's emotions and thought patterns and classify data based on that understanding. For example, machine learning algorithms can efficiently classify large amounts of data. Natural language processing techniques can analyze the content of data and classify it into appropriate categories. Understanding the user's emotions and thought patterns allows for more accurate classification. The explanation unit explains past events based on the data classified by the classification unit. For example, the explanation unit can explain text data using speech synthesis technology. It can also display text data on a display device. Furthermore, the explanation unit can replay past events as video data. For example, speech synthesis technology can explain text data in natural speech. A display device can visually display and provide text data to the user. Playing video data can visually recreate and provide past events to the user. The playback unit reconstructs past life based on data classified by the classification unit. The playback unit can, for example, play audio data. It can also play video data. Furthermore, the playback unit can recreate past life using virtual reality (VR) technology. For example, by playing audio data, past events can be recreated aurally. By playing video data, past events can be recreated visually. By using virtual reality technology, past life can be recreated in an immersive way and provided to the user. As a result, the AI system according to this embodiment can easily record the user's daily thoughts and actions and utilize that information.
[0071] The recording unit records the user's daily events and thoughts in audio format. For example, the recording unit records what the user says using a microphone and saves it as audio data. Specifically, the recording unit uses a high-sensitivity microphone to record the user's voice clearly. This minimizes background noise and improves the quality of the audio data. The recording unit can also record audio using devices such as smartphones and tablets. This allows users to easily record audio even when they are out and about. Furthermore, the recording unit can convert audio to text data in real time using speech recognition technology. For example, the recording unit records what the user says using a high-precision microphone and saves it as audio data. By using a smartphone or tablet, users can record audio anytime, anywhere. Using speech recognition technology, audio can be converted to text data in real time and saved instantly. The speech recognition technology uses a model based on deep learning, which achieves high recognition accuracy. For example, if a user says, "We discussed a new project at today's meeting," the recording unit can instantly convert this audio into text data and save it. This allows the user to easily review the content later.
[0072] The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing techniques. For example, the text conversion unit can use speech recognition technology to convert the audio data into text data. Specifically, speech recognition technology includes a process of analyzing the audio waveform and breaking it down into phonemes and words. This allows for the conversion of audio data into highly accurate text data. The text conversion unit can also use machine learning algorithms to analyze the content of the audio data and convert it into text data. Furthermore, the text conversion unit can understand the context of the audio data and convert it into appropriate text data. For example, using speech recognition technology can generate highly accurate text data. Using machine learning algorithms can analyze the content of the audio data and generate more accurate text data. Understanding the context of the audio data allows for the generation of text data in a natural-sounding format. For example, if a user says, "What are my plans for tomorrow?", the text conversion unit converts this audio into the text data "What are my plans for tomorrow?". Furthermore, the text conversion unit can refer to past conversation history and user utterance patterns to understand the context. This allows for the generation of more natural and accurate text data.
[0073] The storage unit stores the data converted to text by the text conversion unit in a dedicated database. The storage unit can, for example, store text data in cloud storage. Specifically, by using a cloud storage service, data can be securely stored and accessed anytime, anywhere. The storage unit can also store data in local storage. This allows access to data even without an internet connection. Furthermore, the storage unit can ensure data security by regularly backing up the data. For example, using cloud storage allows for secure data storage and access at any time. Using local storage allows for offline data storage and access even without an internet connection. Regularly backing up data prevents data loss and ensures security. The storage unit can also protect the confidentiality of stored data using data encryption technology. For example, it can encrypt data using encryption algorithms such as AES (Advanced Encryption Standard) to protect against unauthorized access. This allows the storage unit to ensure data security and confidentiality, and protect user privacy.
[0074] The classification unit classifies data stored in the storage unit based on the user's statements, emotions, and thought patterns. The classification unit uses, for example, machine learning algorithms to classify data. Specifically, it analyzes the user's statements and emotions and classifies the data into appropriate categories based on that analysis. For example, if a user says, "I had a great time today," the classification unit classifies this statement as a "positive emotion." The classification unit can also analyze data using natural language processing techniques and classify it into appropriate categories. Furthermore, the classification unit can understand the user's emotions and thought patterns and classify data based on that understanding. For example, machine learning algorithms can efficiently classify large amounts of data. Natural language processing techniques can analyze the content of data and classify it into appropriate categories. Understanding the user's emotions and thought patterns allows for more accurate classification. The classification unit can also learn the user's past statements and behavioral patterns and predict future statements and behaviors. This enables the classification unit to provide personalized services tailored to the user's needs and preferences.
[0075] The explanation unit describes past events based on data classified by the classification unit. For example, the explanation unit can explain text data using speech synthesis technology. Specifically, speech synthesis technology uses algorithms to convert text into natural-sounding speech. This allows the user to hear past events in audio. The explanation unit can also display text data on a display device. This allows the user to visually confirm the information. Furthermore, the explanation unit can play back past events as video data. For example, by using speech synthesis technology, text data can be explained in natural-sounding speech. By using a display device, text data can be visually displayed and provided to the user. By playing back video data, past events can be visually recreated and provided to the user. The explanation unit can also provide relevant information based on the user's past statements and actions. For example, if the user asks, "Tell me about what happened during last summer vacation," the explanation unit will provide relevant information in audio and video based on statements and actions from that period. This allows the explanation unit to supplement the user's memory and help them recall past events more vividly.
[0076] The playback unit replays past life events based on data classified by the classification unit. For example, the playback unit can play audio data. Specifically, it can use high-quality speakers to reproduce clear sound. The playback unit can also play video data, allowing users to visually recreate past events. Furthermore, the playback unit can recreate past life events using virtual reality (VR) technology. For example, by playing audio data, past events can be recreated aurally. By playing video data, past events can be recreated visually. Using virtual reality technology, past life can be recreated in an immersive way and provided to the user. The playback unit can also provide relevant information based on the user's past statements and actions. For example, if a user instructs, "Play back what happened on my birthday last year," the playback unit will play audio and video data from that period and provide it to the user. This allows the playback unit to supplement the user's memory and help them recall past events more vividly. The playback unit can also customize the order and content of the data played according to the user's preferences. This allows the playback unit to recreate and present past events in the most optimal way for the user.
[0077] The storage unit can store photos and video data in association with each other. For example, it can store photos and video data taken by the user together with audio data. The storage unit can also store photos and video data in a dedicated database and link it to audio data. Furthermore, the storage unit can store photos and video data in cloud storage and synchronize it with audio data. For example, storing photos and video data taken by the user together with audio data allows for more detailed recording. Storing data in a dedicated database makes data management easier. Storing data in cloud storage ensures secure storage and access at any time. This allows for more detailed recording by storing photos and video data in association with each other.
[0078] The recording unit can estimate the user's emotions and adjust the timing of voice recording based on the estimated emotions. For example, if the user is feeling stressed, the recording unit can prompt voice recording at a time when the user can relax. The recording unit can also start voice recording immediately if the user is excited. Furthermore, if the user is tired, the recording unit can prompt voice recording after a break. For example, if the user is stressed, prompting voice recording at a time when the user can relax allows for a more natural recording. If the user is excited, starting voice recording immediately at that moment allows for the recording of important events. If the user is tired, prompting voice recording after a break reduces the user's burden. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate timing of voice recording by adjusting the timing according to the user's emotions.
[0079] The recording unit can analyze the user's past voice recording history and select the optimal recording method. For example, the recording unit can prioritize suggesting recording methods (voice, text, etc.) that the user has preferred in the past. The recording unit can also suggest appropriate recording timing based on the user's past recording frequency. Furthermore, the recording unit can analyze the user's past recording content and suggest recording methods based on specific themes. For example, by prioritizing suggesting recording methods that the user has preferred in the past, it can provide a recording method that suits the user's preferences. By suggesting appropriate recording timing based on the user's past recording frequency, the efficiency of recording can be improved. By analyzing the user's past recording content and suggesting recording methods based on specific themes, more appropriate recordings can be made. Some or all of the above processing in the recording unit may be performed using AI, for example, or not. For example, the recording unit can input the user's past voice recording history into AI and have the AI select the optimal recording method. This allows the recording unit to provide the user with the most suitable recording method by analyzing the past voice recording history.
[0080] The recording unit can filter audio recordings based on the user's current activity and environment. For example, if the user is in a quiet environment, the recording unit will start recording audio. If the user is in a noisy environment, the recording unit can also apply noise cancellation during recording. Furthermore, if the user is moving, the recording unit can wait to record audio until the user has finished moving. For example, if the user is in a quiet environment, starting recording immediately will allow for the recording of clear audio with minimal noise. If the user is in a noisy environment, applying noise cancellation will remove background noise and improve audio quality. If the user is moving, waiting to record audio until the user has finished moving will avoid noise and vibrations during movement, resulting in stable audio recording. Some or all of the above processing in the recording unit may be performed using AI, or not. For example, the recording unit can input the user's current activity and environment data into the AI and have the AI perform the filtering. This allows for more appropriate recording by filtering audio recordings according to the user's activity and environment.
[0081] The recording unit can estimate the user's emotions and determine the priority of audio recording based on the estimated emotions. For example, if the user is talking about an emotionally important event, the recording unit will prioritize recording that audio. Conversely, if the user is talking about an everyday event, the recording unit can postpone recording that. Furthermore, if the user is expressing a specific emotion (joy, sadness, etc.), the recording unit can prioritize recording that audio. For example, by prioritizing recording an emotionally important event, important information can be captured without being missed. By postponing recording everyday events, important events can be prioritized. By prioritizing recording a specific emotion, changes in emotion can be recorded in detail. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows important audio to be prioritized by determining the priority of audio according to the user's emotions.
[0082] The recording unit can prioritize recording audio that is highly relevant, taking into account the user's geographical location information during audio recording. For example, if the user is in a specific location (such as a travel destination), the recording unit will prioritize recording audio related to that location. Furthermore, if the user is at home, the recording unit can also record everyday events. Additionally, if the user is attending a specific event (such as a meeting or party), the recording unit can prioritize recording audio related to that event. For example, if the user is in a specific location, prioritizing the recording of audio related to that location allows for a detailed record of events at the travel destination or specific location. If the user is at home, recording everyday events enriches the record of daily life. If the user is attending a specific event, prioritizing the recording of audio related to that event allows for a detailed record of the event. Some or all of the above processing in the recording unit may be performed using AI, or not. For example, the recording unit can input the user's geographical location information into the AI and have the AI prioritize highly relevant audio. This allows for the recording of highly relevant audio by taking geographical location information into consideration.
[0083] The recording unit can analyze the user's social media activity during audio recording and record relevant audio. For example, if the user posts about a specific topic on social media, the recording unit can record audio related to that topic. It can also record audio related to emotional posts by the user. Furthermore, if the user posts about a specific event on social media, the recording unit can record audio related to that event. For example, if the user posts about a specific topic on social media, recording audio related to that topic allows for a record linked to social media activity. If the user makes an emotional post, recording audio related to that emotion allows for a detailed record of emotional changes. If the user posts about a specific event, recording audio related to that event enriches the event record. Some or all of the above processing in the recording unit may be performed using AI, for example, or without AI. For example, the recording unit can input the user's social media activity data into AI and have the AI record relevant audio. This allows for the recording of relevant audio by analyzing social media activity.
[0084] The text generation unit can estimate the user's emotions and adjust the textual expression based on the estimated emotions. For example, if the user is speaking emotionally, the text generation unit will use an expression that reflects those emotions. The text generation unit can also use a factual expression when the user is speaking calmly. Furthermore, if the user is excited, the text generation unit can use an expression that reflects that excitement. For example, if the user is speaking emotionally, using an expression that reflects those emotions allows for a detailed record of emotional changes. If the user is speaking calmly, using a factual expression allows for the recording of accurate information. If the user is excited, using an expression that reflects that excitement allows for a detailed record of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate text generation by adjusting the textual expression according to the user's emotions.
[0085] The transcription unit can adjust the level of detail in the transcription based on the importance of the audio. For example, the transcription unit will transcribe audio about important events in detail. It can also transcribe audio about everyday events in a simplified manner. Furthermore, it can transcribe audio about emotional events in a detailed manner that reflects the emotions. For example, transcribing audio about important events in detail ensures that important information is not missed. Transcribing audio about everyday events in a simplified manner allows for efficient recording. Transcribing audio about emotional events in a detailed manner that reflects the emotions allows for a detailed recording of emotional changes. Some or all of the above processing in the transcription unit may be performed using AI, for example, or without AI. For example, the transcription unit can input audio data into AI and have the AI perform the transcription detail adjustment based on importance. This allows important information to be transcribed in detail by adjusting the level of detail in the transcription based on the importance of the audio.
[0086] The text conversion unit can apply different text conversion algorithms depending on the audio category during the text conversion process. For example, the text conversion unit can apply a dialogue-style text conversion algorithm to conversational audio. It can also apply a summary-style text conversion algorithm to lecture audio. Furthermore, it can apply an emotion-reflecting text conversion algorithm to emotional audio. For example, applying a dialogue-style text conversion algorithm to conversational audio can generate natural-sounding dialogue. Applying a summary-style text conversion algorithm to lecture audio can generate a summary that extracts the key points. Applying an emotion-reflecting text conversion algorithm to emotional audio can record emotional changes in detail. Some or all of the above processing in the text conversion unit may be performed using AI, for example, or without AI. For example, the text conversion unit can input audio data into AI and have the AI apply a text conversion algorithm according to the category. This allows for more appropriate text conversion by applying different text conversion algorithms depending on the audio category.
[0087] The text generation unit can estimate the user's emotions and adjust the length of the text based on the estimated emotions. For example, if the user is speaking emotionally, the text generation unit will perform detailed text. It can also perform concise text if the user is speaking calmly. Furthermore, if the user is excited, the text generation unit can perform longer text that reflects their excitement. For example, if the user is speaking emotionally, detailed text can record changes in emotion in detail. If the user is speaking calmly, concise text can record efficiently. If the user is excited, longer text that reflects their excitement can record the degree of excitement in detail. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate text generation by adjusting the length of the text according to the user's emotions.
[0088] The text conversion unit can determine the priority of transcription based on the recording date of the audio. For example, the text conversion unit may prioritize the transcription of recently recorded audio. It can also prioritize the transcription of audio related to specific events (such as birthdays or anniversaries). Furthermore, the text conversion unit can prioritize the transcription of audio within a period specified by the user. For example, prioritizing the transcription of recently recorded audio allows for the rapid recording of the latest information. Prioritizing the transcription of audio related to specific events enriches the recording of important events. Prioritizing the transcription of audio within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the text conversion unit may be performed using AI, for example, or without AI. For example, the text conversion unit can input audio data into AI and have the AI prioritize transcription based on the recording date. This allows for the prioritization of important audio by determining the transcription priority based on the recording date of the audio.
[0089] The transcription unit can adjust the order of transcription based on the relevance of the audio. For example, the transcription unit can prioritize the transcription of highly relevant audio. It can also postpone the transcription of less relevant audio. Furthermore, the transcription unit can prioritize the transcription of audio related to a specific theme. For example, prioritizing the transcription of highly relevant audio allows for efficient recording of relevant information. Postponing the transcription of less relevant audio allows for priority recording of important information. Prioritizing the transcription of audio related to a specific theme allows for recording in line with the theme. Some or all of the above processing in the transcription unit may be performed using AI, for example, or without AI. For example, the transcription unit can input audio data into AI and have the AI execute the transcription order based on relevance. This allows for prioritizing the transcription of highly relevant audio by adjusting the transcription order based on the relevance of the audio.
[0090] The storage unit can estimate the user's emotions and select data to save based on the estimated emotions. For example, the storage unit can prioritize saving audio data where the user is speaking emotionally. It can also postpone saving audio data where the user is speaking calmly. Furthermore, the storage unit can prioritize saving audio data where the user is excited. For example, prioritizing the saving of audio data where the user is speaking emotionally allows for a detailed record of emotional changes. Postponing the saving of audio data where the user is speaking calmly allows for the prioritization of important data. Prioritizing the saving of audio data where the user is excited allows for a detailed record of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for the prioritization of important data by selecting data to save according to the user's emotions.
[0091] The storage unit can optimize the storage algorithm by referring to previously saved data during the saving process. For example, the storage unit can refer to the format of previously saved data and save it in the same format. It can also refer to the importance of previously saved data and prioritize saving important data. Furthermore, the storage unit can refer to the category of previously saved data and save it within the same category. For example, referring to the format of previously saved data allows for consistent data storage. Referring to the importance of previously saved data allows for prioritizing the saving of important data. Referring to the category of previously saved data makes data organization easier. Some or all of the above processes in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input previously saved data into AI and have the AI optimize the storage algorithm. This allows for optimization of the storage algorithm by referring to previously saved data.
[0092] The storage unit can adjust the level of detail in saving data based on its importance. For example, it can save important data in detail. It can also save everyday data in a simplified manner. Furthermore, it can save emotional data in a detailed manner that reflects the emotions involved. For example, saving important data in detail ensures that important information is not missed. Saving everyday data in a simplified manner allows for efficient recording. Saving emotional data in a detailed manner that reflects emotions allows for a detailed record of emotional changes. Some or all of the above processing in the storage unit may be performed using AI, for example, or not. For example, the storage unit can input data into AI and have the AI perform the level of detail in saving based on importance. This allows important data to be saved in detail by adjusting the level of detail in saving based on the importance of the data.
[0093] The storage unit can estimate the user's emotions and determine the priority of saved data based on the estimated user emotions. For example, the storage unit can prioritize saving data in which the user is speaking emotionally. It can also postpone saving data in which the user is speaking calmly. Furthermore, the storage unit can prioritize saving data in which the user is agitated. For example, prioritizing the saving of data in which the user is speaking emotionally allows for a detailed record of emotional changes. Postponing the saving of data in which the user is speaking calmly allows for the prioritization of important data. Prioritizing the saving of data in which the user is agitated allows for a detailed record of the degree of agitation. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for the prioritization of important data by determining the priority of saved data according to the user's emotions.
[0094] The storage unit can weight the data to be saved based on when the data was recorded. For example, the storage unit can prioritize saving recently recorded data. It can also prioritize saving data related to specific events (birthdays, anniversaries, etc.). Furthermore, the storage unit can prioritize saving data within a period specified by the user. For example, prioritizing the saving of recently recorded data allows for the rapid recording of the latest information. Prioritizing the saving of data related to specific events enriches the recording of important events. Prioritizing the saving of data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the storage unit may be performed using AI, for example, or not. For example, the storage unit can input data into AI and have the AI perform weighting of the saved data based on the recording date. This allows for prioritizing the saving of important data by weighting the saved data based on the recording date.
[0095] The storage unit can adjust the order of saving data based on its relevance. For example, the storage unit can prioritize saving highly relevant data. It can also postpone saving less relevant data. Furthermore, the storage unit can prioritize saving data related to a specific theme. For example, prioritizing the saving of highly relevant data allows for efficient recording of relevant information. Postponing the saving of less relevant data allows for prioritizing the saving of important information. Prioritizing the saving of data related to a specific theme allows for theme-based recording. Some or all of the above processing in the storage unit may be performed using AI, for example, or without AI. For example, the storage unit can input data into AI and have the AI execute a saving order based on relevance. This allows for prioritizing the saving of highly relevant data by adjusting the saving order based on data relevance.
[0096] The classification unit can estimate the user's emotions and adjust the classification criteria based on the estimated emotions. For example, the classification unit can prioritize classifying data in which the user is speaking emotionally. It can also prioritize classifying data in which the user is speaking calmly. Furthermore, it can prioritize classifying data in which the user is agitated. For example, prioritizing the classification of data in which the user is speaking emotionally allows for detailed recording of emotional changes. Prioritizing the classification of data in which the user is speaking calmly allows for prioritizing the classification of important data. Prioritizing the classification of data in which the user is agitated allows for detailed recording of the degree of agitation. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate classification by adjusting the classification criteria according to the user's emotions.
[0097] The classification unit can improve the accuracy of classification by considering the interrelationships between data during the classification process. For example, the classification unit can group and classify highly relevant data. It can also analyze the interrelationships between data and classify related data together. Furthermore, the classification unit can efficiently classify data by eliminating duplicate data while considering the interrelationships. For example, grouping and classifying highly relevant data allows for efficient recording of related information. Analyzing the interrelationships between data and classifying related data together makes data organization easier. Eliminating duplicate data while considering the interrelationships allows for efficient classification. Some or all of the above processes in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input data into AI and have the AI perform classification accuracy improvements based on interrelationships. This improves classification accuracy by considering the interrelationships between data.
[0098] The classification unit can perform classification while considering the attribute information of the data submitter. For example, the classification unit can classify data based on the user's age and gender. It can also classify data based on the user's occupation and hobbies. Furthermore, it can classify data based on the user's place of residence and cultural background. For example, classifying data based on the user's age and gender allows for attribute-based recording. Classifying data based on the user's occupation and hobbies allows for recording based on interests and concerns. Classifying data based on the user's place of residence and cultural background allows for recording based on region and culture. Some or all of the above processing in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input user attribute information into AI and have the AI perform attribute-based classification. This allows for more appropriate classification by considering the attribute information of the data submitter.
[0099] The classification unit can estimate the user's emotions and adjust the order in which the classification results are displayed based on the estimated emotions. For example, the classification unit can prioritize displaying data in which the user is speaking emotionally. It can also postpone displaying data in which the user is speaking calmly. Furthermore, it can prioritize displaying data in which the user is agitated. For example, prioritizing the display of data in which the user is speaking emotionally allows for detailed recording of emotional changes. Postponing the display of data in which the user is speaking calmly allows for prioritizing the display of important data. Prioritizing the display of data in which the user is agitated allows for detailed recording of the degree of agitation. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate display by adjusting the order in which the classification results are displayed according to the user's emotions.
[0100] The classification unit can perform classification while considering the geographical distribution of the data. For example, the classification unit can classify data based on the user's place of residence. It can also classify data based on places the user has visited. Furthermore, it can classify data based on the user's travel destinations. For example, classifying data based on the user's place of residence allows for regionally appropriate records. Classifying data based on places the user has visited allows for more detailed records of visited locations. Classifying data based on the user's travel destinations allows for more detailed travel records. Some or all of the above processing in the classification unit may be performed using AI, for example, or without AI. For example, the classification unit can input geographical distribution data into AI and have the AI perform classification based on geographical distribution. This allows for more appropriate classification by considering the geographical distribution of the data.
[0101] The classification unit can improve the accuracy of classification by referring to relevant literature during the classification process. For example, the classification unit can adjust the classification criteria for the data by referring to relevant literature. The classification unit can also analyze the interrelationships of the data by referring to relevant literature. Furthermore, the classification unit can eliminate data duplication by referring to relevant literature. For example, by adjusting the classification criteria for the data by referring to relevant literature, the accuracy of classification can be improved. By analyzing the interrelationships of the data by referring to relevant literature, relevant information can be efficiently recorded. By eliminating data duplication by referring to relevant literature, classification can be performed efficiently. Some or all of the above processes in the classification unit may be performed using AI, for example, or not using AI. For example, the classification unit can input relevant literature data into AI and have the AI perform the classification accuracy improvement. This improves the accuracy of classification by referring to relevant literature for the data.
[0102] The explanation unit can estimate the user's emotions and adjust the way it expresses the explanation based on the estimated emotions. For example, if the user is speaking emotionally, the explanation unit will use an expression that reflects those emotions. The explanation unit can also use a factually-based expression if the user is speaking calmly. Furthermore, if the user is excited, the explanation unit can use an expression that reflects that excitement. For example, if the user is speaking emotionally, using an expression that reflects those emotions allows for a detailed record of emotional changes. If the user is speaking calmly, using a factually-based expression allows for the recording of accurate information. If the user is excited, using an expression that reflects that excitement allows for a detailed record of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate explanations by adjusting the way the explanation is expressed according to the user's emotions.
[0103] The explanation unit can adjust the level of detail in its explanations based on the importance of the data. For example, it can provide detailed explanations for important data. It can also provide simplified explanations for everyday data. Furthermore, it can provide detailed explanations that reflect emotions for emotional data. For example, providing detailed explanations for important data ensures that important information is not missed. Simplified explanations for everyday data allow for efficient recording. Detailed explanations that reflect emotions for emotional data allow for detailed recording of emotional changes. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI perform level of detail in the explanations based on importance. This allows for detailed explanations of important data by adjusting the level of detail in the explanations based on the importance of the data.
[0104] The explanation unit can apply different explanation algorithms depending on the data category during explanation. For example, the explanation unit can apply a dialogue-style explanation algorithm to conversational data. It can also apply a summary-style explanation algorithm to lecture data. Furthermore, it can apply an emotion-reflecting explanation algorithm to emotional data. For example, applying a dialogue-style explanation algorithm to conversational data allows for a natural, dialogue-style explanation. Applying a summary-style explanation algorithm to lecture data allows for a summary that extracts the important points. Applying an emotion-reflecting explanation algorithm to emotional data allows for a detailed recording of emotional changes. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI apply a category-appropriate explanation algorithm. This allows for more appropriate explanations by applying different explanation algorithms depending on the data category.
[0105] The descriptive unit can estimate the user's emotions and adjust the length of the explanation based on the estimated emotions. For example, if the user is speaking emotionally, the descriptive unit will provide a detailed explanation. If the user is speaking calmly, the descriptive unit can provide a concise explanation. Furthermore, if the user is excited, the descriptive unit can provide a longer explanation that reflects that excitement. For example, if the user is speaking emotionally, providing a detailed explanation allows for a detailed recording of changes in emotions. If the user is speaking calmly, providing a concise explanation allows for efficient recording. If the user is excited, providing a longer explanation that reflects that excitement allows for a detailed recording of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate explanations by adjusting the length of the explanation according to the user's emotions.
[0106] The explanation unit can determine the priority of explanations based on the recording date of the data during explanation. For example, the explanation unit may prioritize explaining recently recorded data. It can also prioritize explaining data related to specific events (birthdays, anniversaries, etc.). Furthermore, the explanation unit can prioritize explaining data within a period specified by the user. For example, prioritizing recently recorded data allows for the rapid recording of the latest information. Prioritizing data related to specific events enriches the recording of important events. Prioritizing data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the explanation unit may be performed using AI, for example, or not. For example, the explanation unit can input data into AI and have the AI prioritize explanations based on the recording date. This allows for prioritizing explanations based on the recording date of the data, thereby prioritizing the explanation of important data.
[0107] The explanation unit can adjust the order of explanations based on the relevance of the data during the explanation process. For example, the explanation unit can prioritize explaining highly relevant data. It can also postpone explaining less relevant data. Furthermore, the explanation unit can prioritize explaining data related to a specific theme. For example, prioritizing highly relevant data allows for efficient recording of relevant information. Postponing the explanation of less relevant data allows for priority recording of important information. Prioritizing the explanation of data related to a specific theme allows for theme-aligned recording. Some or all of the above processing in the explanation unit may be performed using AI, for example, or without AI. For example, the explanation unit can input data into AI and have the AI execute the explanation order based on relevance. This allows for prioritizing the explanation of highly relevant data by adjusting the order of explanations based on the relevance of the data.
[0108] The playback unit can estimate the user's emotions and adjust the playback method based on the estimated emotions. For example, if the user is speaking emotionally, the playback unit will play back in a way that reflects those emotions. The playback unit can also play back in a factually-based way if the user is speaking calmly. Furthermore, if the user is excited, the playback unit can play back in a way that reflects that excitement. For example, if the user is speaking emotionally, playing back in a way that reflects those emotions allows for a detailed recording of emotional changes. If the user is speaking calmly, playing back in a factually-based way allows for the recording of accurate information. If the user is excited, playing back in a way that reflects that excitement allows for a detailed recording of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate playback by adjusting the playback method according to the user's emotions.
[0109] The playback unit can adjust the level of detail during playback based on the importance of the data. For example, the playback unit will play important data in detail. It can also play everyday data in a simplified manner. Furthermore, the playback unit can play emotional data in a detailed manner that reflects the emotions. For example, by playing important data in detail, important information can be recorded without being missed. By playing everyday data in a simplified manner, recording can be done efficiently. By playing emotional data in a detailed manner that reflects the emotions, changes in emotions can be recorded in detail. Some or all of the above processing in the playback unit may be performed using AI, for example, or not using AI. For example, the playback unit can input data into AI and have the AI perform playback detail based on importance. This allows important data to be played back in detail by adjusting the level of detail of playback based on the importance of the data.
[0110] The playback unit can apply different playback algorithms depending on the data category during playback. For example, the playback unit can apply a dialogue-style playback algorithm to conversational data. It can also apply a summary-style playback algorithm to lecture data. Furthermore, the playback unit can apply an emotion-reflecting playback algorithm to emotional data. For example, applying a dialogue-style playback algorithm to conversational data allows for natural-sounding dialogue playback. Applying a summary-style playback algorithm to lecture data allows for summarization that extracts important points. Applying an emotion-reflecting playback algorithm to emotional data allows for detailed recording of emotional changes. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI apply a playback algorithm according to the category. This allows for more appropriate playback by applying different playback algorithms depending on the data category.
[0111] The playback unit can estimate the user's emotions and adjust the playback order based on the estimated emotions. For example, the playback unit can prioritize playback of data in which the user is speaking emotionally. It can also postpone playback of data in which the user is speaking calmly. Furthermore, it can prioritize playback of data in which the user is excited. For example, prioritizing playback of data in which the user is speaking emotionally allows for detailed recording of emotional changes. Postponing playback of data in which the user is speaking calmly allows for priority playback of important data. Prioritizing playback of data in which the user is excited allows for detailed recording of the degree of excitement. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI includes, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. This allows for more appropriate playback by adjusting the playback order according to the user's emotions.
[0112] The playback unit can determine playback priority based on the recording date of the data during playback. For example, the playback unit may prioritize playback of recently recorded data. It can also prioritize playback of data related to specific events (such as birthdays or anniversaries). Furthermore, the playback unit can prioritize playback of data within a period specified by the user. For example, prioritizing playback of recently recorded data allows for the rapid recording of the latest information. Prioritizing playback of data related to specific events enriches the recording of important events. Prioritizing playback of data within a period specified by the user allows for recording tailored to the user's needs. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI prioritize playback based on the recording date. This allows for the priority of playback of important data by determining playback priority based on the recording date of the data.
[0113] The playback unit can adjust the playback order based on the relevance of the data during playback. For example, the playback unit can prioritize playback of highly relevant data. It can also postpone playback of less relevant data. Furthermore, the playback unit can prioritize playback of data related to a specific theme. For example, prioritizing playback of highly relevant data allows for efficient recording of relevant information. Postponing playback of less relevant data allows for priority recording of important information. Prioritizing playback of data related to a specific theme allows for recording in line with the theme. Some or all of the above processing in the playback unit may be performed using AI, for example, or without AI. For example, the playback unit can input data into AI and have the AI execute a playback order based on relevance. This allows for prioritizing playback of highly relevant data by adjusting the playback order based on the relevance of the data.
[0114] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0115] The recording unit can monitor the user's health condition while recording their voice data, and adjust the frequency and timing of recording accordingly. For example, if the user is tired, the recording unit will reduce the frequency of voice recording and encourage recording when the user is relaxed. If the user is stressed, the recording unit can play relaxing music to reduce stress before recording voice. Furthermore, if the user is in good health, the recording unit can record voice at the normal frequency. This ensures that voice recording is performed at the appropriate time according to the user's health condition.
[0116] The text conversion unit can analyze the user's speaking speed and tone when converting user audio data into text, thereby improving the accuracy of the conversion. For example, if the user is speaking quickly, the text conversion unit takes the speaking speed into consideration. Furthermore, if the user is speaking emotionally, the unit can reflect that tone in the text. Additionally, if the user is speaking slowly, the text conversion unit can adjust to that pace. This enables highly accurate text conversion tailored to the user's speaking style.
[0117] The storage unit can prioritize saving user audio data based on its importance. For example, audio data about important events can be saved first, while audio data about everyday events can be saved later. Audio data about emotional events can also be saved first to record emotional changes in detail. Furthermore, audio data related to specific themes specified by the user can be saved first. This allows for the priority saving of important data and efficient data management.
[0118] The classification unit can classify user voice data based on the user's lifestyle and hobbies. For example, if a user enjoys sports, it can prioritize classifying sports-related voice data. Similarly, if a user enjoys traveling, it can prioritize classifying travel-related voice data. Furthermore, if a user has a specific hobby (e.g., cooking or reading), it can prioritize classifying voice data related to that hobby. This allows for appropriate data classification tailored to the user's lifestyle and hobbies.
[0119] The explanation unit can adjust the content of the explanation by referring to the user's past behavior patterns and statements when explaining the user's voice data. For example, it can provide detailed explanations for themes that the user has frequently mentioned in the past. It can also prioritize explaining information related to topics that the user has shown interest in in the past. Furthermore, it can reflect the emotions the user felt when talking about events in the past. This allows for appropriate explanations based on the user's past behavior and statements.
[0120] The recording unit can estimate the user's emotions and filter the content of the audio recording based on those emotions. For example, if the user is sad, it can prioritize recording positive content to alleviate those emotions. If the user is happy, it can record content that emphasizes that happiness. Furthermore, if the user is angry, it can record calming content to soothe that anger. This allows for audio recordings with content appropriate to the user's emotions.
[0121] The text conversion unit can estimate the user's emotions and adjust the text conversion style based on the estimated emotions. For example, if the user is speaking emotionally, the text will be converted using expressions that reflect those emotions. If the user is speaking calmly, the text can be converted using expressions that are factual. Furthermore, if the user is excited, the text can be converted using expressions that reflect that excitement. This allows for appropriate text conversion that is tailored to the user's emotions.
[0122] The storage unit can estimate the user's emotions and determine the priority of saved data based on the estimated emotions. For example, it can prioritize saving data where the user is speaking emotionally. It can also postpone saving data where the user is speaking calmly. Furthermore, it can prioritize saving data where the user is excited. This allows for the priority saving of important data in accordance with the user's emotions.
[0123] The classification unit can estimate the user's emotions and adjust the classification criteria based on those emotions. For example, it can prioritize classifying data where the user is speaking emotionally, and postpone classifying data where the user is speaking calmly. Furthermore, it can prioritize classifying data where the user is excited. This allows for appropriate classification based on the user's emotions.
[0124] The playback unit can estimate the user's emotions and adjust the playback method based on the estimated emotions. For example, if the user is speaking emotionally, the playback will reflect those emotions. If the user is speaking calmly, the playback can be based on facts. Furthermore, if the user is excited, the playback can reflect that excitement. This allows for appropriate playback according to the user's emotions.
[0125] The following briefly describes the processing flow for example form 2.
[0126] Step 1: The recording unit records the user's daily events and thoughts in audio format. For example, it records what the user says using a microphone and saves it as audio data. Audio can also be recorded using devices such as smartphones and tablets. Furthermore, speech recognition technology can be used to convert audio into text data in real time. Step 2: The text conversion unit converts the audio data recorded by the recording unit into text using natural language processing technology. For example, it uses speech recognition technology or machine learning algorithms to convert the audio data into text data, understands the context of the audio data, and generates appropriate text data. Step 3: The storage unit saves the data converted to text by the text conversion unit to a dedicated database. For example, it can be saved to cloud storage or local storage, and data security is ensured by regularly backing up the data. Step 4: The classification unit classifies the data stored in the storage unit based on the user's statements, emotions, and thought patterns. For example, it analyzes the data using machine learning algorithms and natural language processing techniques to classify it into the appropriate category. Step 5: The explanation unit explains past events based on the data classified by the classification unit. For example, it can explain text data using speech synthesis technology, display text data on a display device, or play it back as video data. Step 6: The playback unit recreates past life based on the data classified by the classification unit. For example, it can play audio or video data, or recreate past life using virtual reality (VR) technology.
[0127] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0128] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0129] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0130] Each of the multiple elements described above, including the recording unit, text conversion unit, storage unit, classification unit, explanation unit, and playback unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 38B of the smart device 14 and stores it as voice data using the control unit 46A. The text conversion unit converts the voice data into text data using the identification processing unit 290 of the data processing unit 12. The storage unit stores the text data in the database 24. The classification unit classifies the data using the identification processing unit 290 of the data processing unit 12. The explanation unit explains past events using the identification processing unit 290 of the data processing unit 12, and the playback unit plays back the voice and video data using the output device 40 of the smart device 14. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0131] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0132] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0133] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0134] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0135] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0136] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0137] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0138] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0139] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0140] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0141] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0142] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0143] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0144] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0145] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0146] Each of the multiple elements described above, including the recording unit, text conversion unit, storage unit, classification unit, explanation unit, and playback unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the smart glasses 214 and stores it as voice data using the control unit 46A. The text conversion unit converts the voice data into text data using the identification processing unit 290 of the data processing unit 12. The storage unit stores the text data in the database 24. The classification unit classifies the data using the identification processing unit 290 of the data processing unit 12. The explanation unit explains past events using the identification processing unit 290 of the data processing unit 12, and the playback unit plays back voice and video data using the speaker 240 of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0147] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0148] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0149] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0150] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0151] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0152] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0153] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0154] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0155] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0156] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0157] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0158] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0159] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0160] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0161] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0162] Each of the multiple elements described above, including the recording unit, text conversion unit, storage unit, classification unit, explanation unit, and playback unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the headset terminal 314 and stores it as voice data using the control unit 46A. The text conversion unit converts the voice data into text data using the identification processing unit 290 of the data processing unit 12. The storage unit stores the text data in the database 24. The classification unit classifies the data using the identification processing unit 290 of the data processing unit 12. The explanation unit explains past events using the identification processing unit 290 of the data processing unit 12, and the playback unit plays back voice and video data using the speaker 240 of the headset terminal 314. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0163] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0164] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0165] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0166] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0167] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0168] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0169] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0170] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0171] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0172] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0173] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0174] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0175] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0176] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0177] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0178] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0179] Each of the multiple elements described above, including the recording unit, text conversion unit, storage unit, classification unit, explanation unit, and playback unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the recording unit records the user's voice using the microphone 238 of the robot 414 and stores it as voice data using the control unit 46A. The text conversion unit converts the voice data into text data using the identification processing unit 290 of the data processing unit 12. The storage unit stores the text data in the database 24. The classification unit classifies the data using the identification processing unit 290 of the data processing unit 12. The explanation unit explains past events using the identification processing unit 290 of the data processing unit 12, and the playback unit plays back voice and video data using the speaker 240 of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0180] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0181] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0182] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0183] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0184] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0185] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0186] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0187] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0188] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0189] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0190] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0191] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0192] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0193] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0194] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0195] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0196] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0197] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0198] (Note 1) A recording unit for recording audio data, A text conversion unit converts the audio data recorded by the recording unit into text, A storage unit for storing the data converted into text by the text conversion unit, A classification unit that classifies the data stored in the aforementioned storage unit, An explanatory unit that explains past events based on the data classified by the aforementioned classification unit, The system includes a playback unit that recreates past life based on data classified by the aforementioned classification unit. A system characterized by the following features. (Note 2) The aforementioned storage unit is Save photos and video data together. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned recording unit is It estimates the user's emotions and adjusts the timing of voice recordings based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned recording unit is The system analyzes the user's past voice recording history and selects the optimal recording method. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned recording unit is When recording audio, filtering is performed based on the user's current activity and environment. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned recording unit is It estimates the user's emotions and determines the priority of audio recordings based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned recording unit is During voice recording, the system prioritizes recording relevant audio by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned recording unit is During audio recording, the system analyzes the user's social media activity and records relevant audio. The system described in Appendix 1, characterized by the features described herein. (Note 9) The text conversion unit, It estimates the user's emotions and adjusts the textual expression based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The text conversion unit, When transcribing, adjust the level of detail in the text based on the importance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 11) The text conversion unit, When converting to text, different text conversion algorithms are applied depending on the audio category. The system described in Appendix 1, characterized by the features described herein. (Note 12) The text conversion unit, It estimates the user's emotions and adjusts the length of the text based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The text conversion unit, When transcribing audio, the transcription priority is determined based on when the audio was recorded. The system described in Appendix 1, characterized by the features described herein. (Note 14) The text conversion unit, During the transcription process, the order of transcription is adjusted based on the relevance of the audio. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned storage unit is The system estimates the user's emotions and selects data to store based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned storage unit is When saving, the saving algorithm is optimized by referring to previously saved data. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned storage unit is When saving, adjust the level of detail based on the importance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned storage unit is It estimates the user's emotions and determines the priority of stored data based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned storage unit is During saving, the saved data is weighted based on the recording date. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned storage unit is When saving, adjust the saving order based on the relevance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned classification unit is It estimates the user's emotions and adjusts the classification criteria based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned classification unit is When classifying data, consider the interrelationships between data to improve classification accuracy. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned classification unit is When classifying data, the attribute information of the data submitter is taken into consideration. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned classification unit is It estimates the user's emotions and adjusts the order in which the classification results are displayed based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned classification unit is When classifying data, consider its geographical distribution. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned classification unit is When classifying data, we refer to relevant literature to improve the accuracy of the classification. The system described in Appendix 1, characterized by the features described herein. (Note 27) The above explanatory section is, It estimates the user's emotions and adjusts the way explanations are presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The above explanatory section is, When explaining, adjust the level of detail in the explanation based on the importance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 29) The above explanatory section is, When describing the data, different descriptive algorithms are applied depending on the data category. The system described in Appendix 1, characterized by the features described herein. (Note 30) The above explanatory section is, It estimates the user's emotions and adjusts the length of the explanation based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 31) The above explanatory section is, During the explanation, prioritize the explanation based on when the data was recorded. The system described in Appendix 1, characterized by the features described herein. (Note 32) The above explanatory section is, When explaining, adjust the order of explanations based on the relevance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned regeneration unit is It estimates the user's emotions and adjusts the playback method based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned regeneration unit is During playback, adjust the level of detail based on the importance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned regeneration unit is During playback, different playback algorithms are applied depending on the data category. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned regeneration unit is It estimates the user's emotions and adjusts the playback order based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 37) The aforementioned regeneration unit is During playback, the playback priority is determined based on the recording date of the data. The system described in Appendix 1, characterized by the features described herein. (Note 38) The aforementioned regeneration unit is During playback, the playback order is adjusted based on the relevance of the data. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]
[0199] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. A recording unit for recording audio data, A text conversion unit converts the audio data recorded by the recording unit into text, A storage unit for storing the data converted into text by the text conversion unit, A classification unit that classifies the data stored in the aforementioned storage unit, An explanatory unit that explains past events based on the data classified by the aforementioned classification unit, The system includes a playback unit that recreates past life based on data classified by the aforementioned classification unit. A system characterized by the following features.
2. The aforementioned storage unit is Save photos and video data together. The system according to feature 1.
3. The aforementioned recording unit is It estimates the user's emotions and adjusts the timing of voice recordings based on the estimated emotions. The system according to feature 1.
4. The aforementioned recording unit is The system analyzes the user's past voice recording history and selects the optimal recording method. The system according to feature 1.
5. The aforementioned recording unit is When recording audio, filtering is performed based on the user's current activity and environment. The system according to feature 1.
6. The aforementioned recording unit is It estimates the user's emotions and determines the priority of audio recordings based on the estimated emotions. The system according to feature 1.
7. The aforementioned recording unit is During voice recording, the system prioritizes recording relevant audio by considering the user's geographical location. The system according to feature 1.
8. The aforementioned recording unit is During audio recording, the system analyzes the user's social media activity and records relevant audio. The system according to feature 1.
9. The text conversion unit, It estimates the user's emotions and adjusts the textual expression based on the estimated emotions. The system according to feature 1.
10. The text conversion unit, When transcribing, adjust the level of detail in the text based on the importance of the audio. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A