system
The system addresses communication challenges in meetings by collecting and analyzing voice and mouth movements to supplement misunderstood words, enhancing meeting efficiency and clarity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Communication obstacles during meetings due to insufficient words or mistakes lead to difficulty in smoothly conducting the meeting.
A system comprising a collection unit, analysis unit, and supplementation unit that collects voice and mouth movements, analyzes dialects and mispronunciations, and supplements words not understood during meetings.
The system compensates for omissions and slips of the tongue, ensuring smoother communication and preventing misunderstandings by providing real-time corrections and supplements.
Smart Images

Figure 2026073012000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the prior art, there are communication obstacles due to insufficient words or mistakes during a meeting, and there is a problem that it is difficult to smoothly conduct the meeting.
[0005] The system according to the embodiment aims to complement insufficient words or mistakes that occur during a meeting and make the meeting proceed more smoothly.
Means for Solving the Problems
[0006] The system according to this embodiment comprises a collection unit, an analysis unit, and a supplementation unit. The collection unit collects the voice and mouth movements of the subject. The analysis unit analyzes the data collected by the collection unit and identifies points of dialect and mispronunciation. The supplementation unit supplements words that the other party does not understand during the meeting based on the points identified by the analysis unit. [Effects of the Invention]
[0007] The system according to this embodiment can compensate for any omissions or slips of the tongue that occur during a meeting, thereby enabling the meeting to proceed more smoothly. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The meeting completion system according to an embodiment of the present invention is a system for making meetings run more smoothly by compensating for insufficient explanations and slips of the tongue during meetings. This meeting completion system first learns the voice and mouth movements of the target participants in advance. This allows it to learn dialects and points of slip of the tongue. Next, it fills in words that the participants do not understand during the meeting via chat or text. For example, it can fill in business terms and specific phrases (e.g., "all-out team," "Apples to Apples," etc.). This system facilitates smooth communication during meetings and prevents misunderstandings and confusion. First, it learns the voice and mouth movements of the target participants in advance. At this time, it learns in detail the dialects and points of slip of the tongue that the participants speak. For example, it learns specific dialects and pronunciation habits and identifies points that should be supplemented based on them. This allows the system to automatically compensate even if the participants make a mistake during the meeting. Next, it fills in words that the participants do not understand during the meeting via chat or text. For example, if business terms or specific phrases come up, it supplements them in real time. For example, if the phrase "all-out team" comes up, it supplements it with "working together as a team." Furthermore, if the phrase "Apples to Apples" comes up, it is corrected to mean "comparison under the same conditions." This makes it easier for all meeting participants to understand and enables smoother communication. In addition, the system corrects what participants say in real time during the meeting. For example, if a participant says "Good morning," it is corrected to "Good morning." Similarly, if a participant says "Thank you for your cooperation today," it is corrected to "Thank you for your cooperation today." This ensures that what is said during the meeting is accurately conveyed and prevents misunderstandings and confusion. This system facilitates smooth communication during meetings and prevents misunderstandings and confusion. For example, if business terms or specific phrases come up, correcting them in real time makes it easier for all meeting participants to understand. Also, correcting what participants say in real time enables accurate communication. This improves the efficiency of meetings and allows for smoother progress.This allows the meeting completion system to fill in any gaps or slips of the tongue during meetings, making them run more smoothly.
[0029] The conference enhancement system according to this embodiment comprises a collection unit, an analysis unit, and an enhancement unit. The collection unit collects the voice and mouth movements of the subject. The collection unit can, for example, collect audio data of the subject speaking and video data of the subject's mouth movements. The collection unit can, for example, collect audio data and video data using a microphone or camera. The collection unit can also collect facial expression data of the subject when they speak. For example, the collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. The analysis unit analyzes the data collected by the collection unit to identify points of dialect and mispronunciation. The analysis unit can, for example, analyze audio data using speech recognition technology to identify patterns of dialect and mispronunciation. The analysis unit can, for example, convert audio data into text data and extract specific patterns of dialect and mispronunciation. The analysis unit can also analyze video data and identify points of dialect and mispronunciation based on changes in mouth movements and facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. The completion unit completes words that participants do not understand during a meeting, based on points identified by the analysis unit. The completion unit can, for example, complete business terms or specific phrases in real time. For example, the completion unit can automatically recognize business terms spoken during a meeting and complete their meaning in chat or text. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This facilitates smoother communication during meetings and prevents misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can perform completion using an AI model that takes business terms or specific phrases as input and outputs their meanings. As a result, the meeting completion system according to this embodiment can complete omissions and slips of the tongue during meetings, making meetings run more smoothly.
[0030] The data collection unit collects the subject's voice and mouth movements. For example, the data collection unit can collect audio data of the subject speaking and video data of their mouth movements. Specifically, the data collection unit uses a high-sensitivity microphone to clearly capture the subject's speech. This improves the quality of the audio data and makes processing easier in the analysis unit. The data collection unit also uses a high-resolution camera to record the subject's mouth movements and facial expressions in detail. This allows for the capture of even subtle changes in mouth movements and facial expressions, improving the accuracy in the analysis unit. Furthermore, the data collection unit can also collect facial expression data while the subject is speaking. For example, the data collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. This allows for a more accurate understanding of the subject's emotions and intentions. The data collection unit collects this data in real time and transmits it to a central database. This allows the collected data to be processed immediately in the analysis unit, enabling rapid supplementation during meetings. In addition, the data collection unit can use multiple microphones and cameras to collect data from multiple subjects simultaneously. This allows for efficient data collection and smooth processing in the analysis unit, even in large meetings or situations with multiple participants. Furthermore, the data acquisition unit can use noise cancellation technology to remove ambient noise and improve the quality of the audio data. This improves the accuracy of speech recognition in the analysis unit, and allows for more accurate interpolation in the interpolation unit.
[0031] The analysis unit analyzes the data collected by the collection unit to identify points of dialect and mispronunciation. For example, the analysis unit can analyze audio data using speech recognition technology to identify patterns of dialect and mispronunciation. Specifically, the analysis unit converts audio data into text data and extracts specific patterns of dialect and mispronunciation. Deep learning models are used for speech recognition technology, enabling highly accurate speech recognition. For example, the analysis unit divides the audio data into frames and extracts features from each frame. This allows it to capture subtle changes in the audio data and identify points of dialect and mispronunciation. The analysis unit can also analyze video data and identify points of dialect and mispronunciation based on mouth movements and changes in facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. This allows it to confirm the consistency between the subject's speech and mouth movements and identify points of dialect and mispronunciation. Furthermore, the analysis unit can analyze facial expression data to understand the subject's emotions and intentions. This allows the analysis unit to understand the underlying intent behind dialects and slips of the tongue, enabling more appropriate completion by the completion unit. The analysis unit provides these analysis results to the completion unit in real time, ensuring rapid completion during meetings. Furthermore, the analysis unit can learn patterns of dialects and slips of the tongue by utilizing past data and statistical information, thereby improving analysis accuracy. As a result, the analysis unit always performs highly accurate analysis based on the latest information, leading to more effective completion by the completion unit.
[0032] The completion unit completes the meaning of words that participants may not understand during a meeting, based on points identified by the analysis unit. For example, the completion unit can complete business terms and specific phrases in real time. Specifically, the completion unit automatically recognizes business terms spoken during a meeting and completes their meaning in chat or text. This allows participants to immediately understand the meaning of terms, making the meeting proceed more smoothly. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This allows participants to accurately understand the intent of what is being said, preventing misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or not. For example, the completion unit can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings. Specifically, it can use generative AI to generate the meaning of the input business terms and phrases in natural language and display them in chat or text. This allows the completion unit to complete any omissions or slips of the tongue during a meeting, making the meeting run more smoothly. Furthermore, the supplementary component can collect participant feedback and continuously improve the accuracy and effectiveness of its supplementary content. For example, it can collect participants' evaluations of the supplemented content and use them as training data for the AI model. This allows the supplementary component to always provide highly accurate supplementary information based on the latest data, thereby improving the quality of meetings.
[0033] The completion function can complete business terms and specific phrases in real time. For example, it can automatically recognize business terms spoken during a meeting and complete their meaning in chat or text. For instance, if the phrase "ROI" is used, the completion function will complete it with "Return on Investment." Similarly, if the phrase "KPI" is used, it can complete it with "Key Performance Indicator." Furthermore, if the phrase "Apples to Apples" is used, it can complete it with "Comparison under the same conditions." This real-time completion of business terms and specific phrases facilitates smoother communication during meetings. Some or all of the above processing in the completion function may be performed using AI, for example, or without AI. For example, the completion function can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings.
[0034] The completion function can convert specific phrases into more general expressions. For example, it can complete the phrase "team baseball" to "everyone working together." It can also complete the phrase "Apples to Apples" to "comparison under the same conditions." Furthermore, it can complete the phrase "ROI" to "return on investment." By converting specific phrases into more general expressions, it makes it easier for all meeting participants to understand. Some or all of the processing described above in the completion function may be performed using AI, for example, or not. For example, the completion function can perform completion using an AI model that takes a specific phrase as input and outputs a general expression.
[0035] The meeting completion system includes a correction unit that corrects the content of speech during a meeting in real time. For example, if a participant says "Good morning," the correction unit can correct it to "Good morning." For example, if a participant says "Thank you for your cooperation today," the correction unit can correct it to "Thank you for your cooperation today." The correction unit can also correct if a participant says "Excuse me, please wait a moment," to "Excuse me, please wait a moment." This allows for real-time correction of speech during a meeting, preventing misunderstandings and confusion. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can use an AI model that takes the participant's speech as input and outputs an accurate expression to perform the correction.
[0036] The correction unit can correct the content of the target person's speech to an accurate expression. For example, if the target person says "Ohayou gozaimashu," the correction unit can correct it to "Ohayou gozaimasu." For example, if the target person says "Kyou mo yoroshiu onegaishimasu," the correction unit can correct it to "Kyou mo yoroshiku onegaishimasu." Also, if the target person says "Sumimasen, chotto matte kudasai," the correction unit can correct it to "Sumimasen, chotto matte kudasai." By correcting the content of the target person's speech to an accurate expression, accurate communication becomes possible. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can perform corrections using an AI model that takes the content of the target person's speech as input and outputs an accurate expression.
[0037] The meeting completion system includes a presentation unit that provides the meaning of specific phrases when they appear. For example, if the phrase "team baseball" appears, the presentation unit might provide the meaning "working together as a team." If the phrase "Apples to Apples" appears, the presentation unit might provide the meaning "comparison under the same conditions." Furthermore, if the phrase "ROI" appears, the presentation unit might provide the meaning "return on investment." This makes it easier for all meeting participants to understand by providing the meaning of specific phrases when they appear. Some or all of the processing described above in the presentation unit may be performed using AI, for example, or not. For example, the presentation unit could use an AI model that takes a specific phrase as input and outputs its meaning to provide the presentation.
[0038] The data collection unit can analyze the subject's past statement history and select the optimal collection method. For example, the data collection unit can collect data at specific time periods based on past statement history. For example, the data collection unit can prioritize collecting statements related to specific topics based on past statement history. The data collection unit can also analyze past statement history and adjust the collection interval according to the frequency of statements. This allows for the selection of the optimal collection method by analyzing past statement history. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can select a collection method using an AI model that takes past statement history data as input and outputs the optimal collection method.
[0039] The data collection unit can filter data based on the subject's current situation and areas of interest during collection. For example, if the subject is in a meeting, the collection unit will only collect statements related to the meeting. If the subject is interested in a particular project, the collection unit can prioritize collecting statements related to that project. The collection unit can also filter and collect statements related to a specific topic if the subject is discussing that topic. This allows for the collection of highly relevant data by filtering based on the subject's current situation and areas of interest. Some or all of the processing described above in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can perform filtering using an AI model that takes data on the subject's current situation and areas of interest as input and outputs filtered results.
[0040] The data collection unit can prioritize collecting highly relevant data based on the subject's geographical location information during data collection. For example, if the subject is in a specific region, the data collection unit will prioritize collecting statements related to that region. For example, if the subject is on the move, the data collection unit can prioritize collecting statements related to their destination. Furthermore, if the subject is in a specific location, the data collection unit can prioritize collecting statements related to that location. This allows for the collection of more useful data by collecting highly relevant data based on the subject's geographical location information. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can collect data using an AI model that takes the subject's geographical location information as input and outputs highly relevant data.
[0041] The data collection unit can analyze the subject's social media activity and collect relevant data during the collection process. For example, the data collection unit can collect words that the subject frequently uses on social media. For example, the data collection unit can collect statements related to topics that the subject is interested in on social media. The data collection unit can also collect statements related to accounts that the subject follows on social media. In this way, relevant data can be collected by analyzing the subject's social media activity. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can collect data using an AI model that takes the subject's social media activity data as input and outputs relevant data.
[0042] The analysis unit can adjust the level of detail of the analysis based on the importance of the data during the analysis. For example, the analysis unit can perform a detailed analysis on important data. For example, it can perform a simplified analysis on less important data. The analysis unit can also determine the priority of the analysis according to the importance of the data. This allows for detailed analysis of important data by adjusting the level of detail based on the importance of the data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the level of detail using an AI model that takes data importance as input and outputs the level of detail of the analysis.
[0043] The analysis unit can apply different analysis algorithms depending on the data category during analysis. For example, the analysis unit can apply a specific analysis algorithm to business terms. For example, it can apply a different analysis algorithm to dialects. It can also apply a standard analysis algorithm to common phrases. By applying different analysis algorithms depending on the data category, more accurate analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can apply an algorithm using an AI model that takes data categories as input and outputs an analysis algorithm.
[0044] The analysis unit can determine the priority of analysis based on the data collection timing during analysis. For example, the analysis unit may prioritize the analysis of the most recent data. For example, the analysis unit may perform analysis while referring to past data. The analysis unit can also adjust the priority of analysis according to the data collection timing. This allows for prioritization of the analysis of the most recent data by determining the priority of analysis based on the data collection timing. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can determine the priority using an AI model that takes the data collection timing as input and outputs the priority of analysis.
[0045] The analysis unit can adjust the order of analysis based on the relevance of the data during analysis. For example, the analysis unit can prioritize the analysis of highly relevant data. For example, the analysis unit can postpone the analysis of less relevant data. The analysis unit can also adjust the order of analysis according to the relevance of the data. This allows for prioritizing the analysis of highly relevant data by adjusting the order of analysis based on the relevance of the data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the order using an AI model that takes the relevance of the data as input and outputs the order of analysis.
[0046] The completion unit can adjust the level of detail of completion based on the importance of the words. For example, the completion unit can perform detailed completion for important words, and simplified completion for less important words. The completion unit can also determine the priority of completion based on the importance of the words. This allows for detailed completion of important words by adjusting the level of detail based on the importance of the words. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the level of detail using an AI model that takes the importance of words as input and outputs the level of detail of completion.
[0047] The completion unit can apply different completion algorithms depending on the word category during completion. For example, the completion unit can apply a specific completion algorithm to business terms. For example, it can apply a different completion algorithm to dialects. It can also apply a standard completion algorithm to common phrases. This allows for more accurate completion by applying different completion algorithms depending on the word category. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can apply an algorithm using an AI model that takes a word category as input and outputs a completion algorithm.
[0048] The completion unit can determine the priority of completion based on the timing of word usage during completion. For example, the completion unit may prioritize the most recent words. For example, the completion unit can perform completion while referring to past words. The completion unit can also adjust the priority of completion according to the timing of word usage. This allows for prioritizing the completion of the most recent words by determining the priority of completion based on the timing of word usage. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can determine the priority using an AI model that takes the timing of word usage as input and outputs the priority of completion.
[0049] The completion unit can adjust the order of completion based on the relevance of the words during completion. For example, the completion unit may prioritize completing highly relevant words. For example, the completion unit may postpone completing less relevant words. The completion unit can also adjust the order of completion according to the relevance of the words. This allows for prioritizing the completion of highly relevant words by adjusting the order of completion based on the relevance of the words. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the order using an AI model that takes the relevance of words as input and outputs the order of completion.
[0050] The editing unit can adjust the level of detail of the edits based on the importance of the statements during the editing process. For example, the editing unit can perform detailed edits on important statements, and simplified edits on less important statements. The editing unit can also determine the priority of edits based on the importance of the statements. This allows for detailed edits on important statements by adjusting the level of detail based on the importance of the statements. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the level of detail using an AI model that takes the importance of a statement as input and outputs the level of detail of the edit.
[0051] The editing unit can apply different editing algorithms depending on the category of the utterance during editing. For example, the editing unit can apply a specific editing algorithm to business terms. For example, it can apply a different editing algorithm to dialects. It can also apply a standard editing algorithm to common phrases. By applying different editing algorithms depending on the category of the utterance, more accurate editing becomes possible. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can apply an algorithm using an AI model that takes the category of the utterance as input and outputs an editing algorithm.
[0052] The editing unit can determine the priority of revisions based on when the statements were used. For example, the editing unit will prioritize revising the most recent statements. For example, the editing unit can perform revisions while referring to past statements. The editing unit can also adjust the priority of revisions according to when the statements were used. This allows the editing unit to prioritize revising the most recent statements by determining the priority of revisions based on when the statements were used. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can determine the priority using an AI model that takes the timing of statement use as input and outputs the priority of revisions.
[0053] The editing unit can adjust the order of editing based on the relevance of the statements during editing. For example, the editing unit can prioritize editing highly relevant statements. For example, the editing unit can postpone editing less relevant statements. The editing unit can also adjust the order of editing according to the relevance of the statements. This allows for prioritizing the editing of highly relevant statements by adjusting the order of editing based on the relevance of the statements. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the order using an AI model that takes the relevance of statements as input and outputs the order of editing.
[0054] The presentation unit can adjust the level of detail of the presentation based on the importance of the phrases. For example, the presentation unit can provide detailed presentations for important phrases, and simplified presentations for less important phrases. The presentation unit can also determine the priority of presentations based on the importance of the phrases. This allows for detailed presentations for important phrases by adjusting the level of detail based on the importance of the phrases. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can adjust the level of detail using an AI model that takes the importance of the phrases as input and outputs the level of detail of the presentation.
[0055] The presentation unit can determine the presentation priority based on when the phrases were used. For example, the presentation unit can prioritize the presentation of the most recent phrases. For example, the presentation unit can make presentations while referring to past phrases. The presentation unit can also adjust the presentation priority according to when the phrases were used. This allows the presentation unit to prioritize the presentation of the most recent phrases by determining the presentation priority based on when the phrases were used. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can determine the priority using an AI model that takes the phrase usage time as input and outputs the presentation priority.
[0056] The presentation unit can adjust the presentation order based on the relevance of the phrases during presentation. For example, the presentation unit can prioritize the presentation of highly relevant phrases. For example, the presentation unit can postpone the presentation of less relevant phrases. The presentation unit can also adjust the presentation order according to the relevance of the phrases. This allows for the priority presentation of highly relevant phrases by adjusting the presentation order based on the relevance of the phrases. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can adjust the order using an AI model that takes the relevance of phrases as input and outputs the presentation order.
[0057] The presentation unit can provide the meaning of specific phrases when they appear during the presentation. For example, if the phrase "team baseball" appears, the presentation unit can provide the meaning "working together as a team." If the phrase "Apples to Apples" appears, the presentation unit can provide the meaning "comparison under the same conditions." Furthermore, if the phrase "ROI" appears, the presentation unit can provide the meaning "return on investment." By providing the meaning of specific phrases when they appear, it becomes easier for all meeting participants to understand. Some or all of the above processing in the presentation unit may be performed using AI, for example, or not. For example, the presentation unit can provide meaning using an AI model that takes a specific phrase as input and outputs its meaning.
[0058] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0059] The meeting enhancement system includes a translation unit that translates the content of a meeting in real time. For example, if a participant speaks in Japanese, the translation unit will translate it into English. For example, if a participant speaks in English, the translation unit can translate it into Japanese. Furthermore, if a participant speaks in French, the translation unit can translate it into Spanish. This facilitates communication between participants who speak different languages by translating the content of a meeting in real time. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can perform translation using an AI model that takes the content of the speech as input and outputs the translation result.
[0060] The meeting completion system includes a summarization unit that summarizes the content of discussions during a meeting. The summarization unit can, for example, summarize and present the key points of the meeting at the end of the meeting. The summarization unit can, for example, summarize the content of discussions during the meeting in real time and provide it to participants. The summarization unit can also provide summaries on specific topics during the meeting. This makes it easier for participants to grasp the key points by summarizing the content of discussions during the meeting. Some or all of the above processing in the summarization unit may be performed using, for example, AI, or not using AI. For example, the summarization unit can perform summarization using an AI model that takes the content of discussion as input and outputs a summary result.
[0061] The meeting completion system includes a recording unit that records the content of speech during a meeting. The recording unit can, for example, record the content of speech during the meeting as audio data. The recording unit can, for example, record the content of speech during the meeting as text data. The recording unit can also record the content of speech during the meeting as video data. This allows for later review of the content of speech during the meeting. Some or all of the above processing in the recording unit may be performed using, for example, AI, or without AI. For example, the recording unit can perform recording using an AI model that takes the content of speech as input and outputs recorded data.
[0062] The meeting enhancement system includes a visualization unit that visualizes the content of discussions during a meeting. The visualization unit can, for example, visualize the content of discussions during a meeting as graphs or charts. The visualization unit can also, for example, visualize the content of discussions during a meeting as a mind map. Furthermore, the visualization unit can visualize the content of discussions during a meeting as a timeline. By visualizing the content of discussions during a meeting, it becomes easier for participants to understand the content. Some or all of the above processing in the visualization unit may be performed using, for example, AI, or without AI. For example, the visualization unit can perform visualization using an AI model that takes the content of discussions as input and outputs visualized data.
[0063] The meeting enhancement system includes an evaluation unit that evaluates the content of statements made during a meeting in real time. The evaluation unit can, for example, evaluate the content of statements made during a meeting and determine the quality of those statements. The evaluation unit can, for example, score the content of statements made during a meeting and evaluate the importance of those statements. The evaluation unit can also rank the content of statements made during a meeting and evaluate the impact of those statements. This makes it easier to grasp the quality and importance of statements by evaluating the content of statements made during a meeting in real time. Some or all of the above processing in the evaluation unit may be performed using, for example, AI, or not using AI. For example, the evaluation unit can perform evaluations using an AI model that takes the content of statements as input and outputs evaluation results.
[0064] The meeting completion system includes a comparison unit that compares the content of statements made during a meeting in real time. The comparison unit can, for example, compare the content of statements made during the meeting with past statements. The comparison unit can also, for example, compare the content of statements made during the meeting with statements made by other participants. Furthermore, the comparison unit can also compare the content of statements made during the meeting with industry standards. This makes it easier to evaluate the consistency and validity of statements by comparing the content of statements made during the meeting in real time. Some or all of the above processing in the comparison unit may be performed using, for example, AI, or not using AI. For example, the comparison unit can perform the comparison using an AI model that takes the content of statements as input and outputs the comparison results.
[0065] The meeting enhancement system includes a prediction unit that predicts the content of discussions during a meeting in real time. The prediction unit predicts, for example, what will be said next based on what has been said during the meeting. The prediction unit can also predict the next topic to be discussed based on what has been said during the meeting. Furthermore, the prediction unit can predict the next problem that will arise based on what has been said during the meeting. This allows for smoother meeting progress by predicting the content of discussions in real time. Some or all of the above processing in the prediction unit may be performed using, for example, AI, or not using AI. For example, the prediction unit can perform predictions using an AI model that takes the content of the discussion as input and outputs prediction results.
[0066] The meeting completion system includes a correction unit that corrects the content of speech during a meeting in real time. For example, if a participant says "Good morning," the correction unit can correct it to "Good morning." For example, if a participant says "Thank you for your cooperation today," the correction unit can correct it to "Thank you for your cooperation today." The correction unit can also correct if a participant says "Excuse me, please wait a moment," to "Excuse me, please wait a moment." This allows for real-time correction of speech during a meeting, preventing misunderstandings and confusion. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can use an AI model that takes the participant's speech as input and outputs an accurate expression to perform the correction.
[0067] The following briefly describes the processing flow for example form 1.
[0068] Step 1: The data collection unit collects the subject's voice and mouth movements. For example, the data collection unit can collect audio data of the subject speaking and video data of their mouth movements. The data collection unit uses microphones and cameras to collect audio and video data. The data collection unit can also collect facial expression data of the subject while they are speaking. For example, the data collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. Step 2: The analysis unit analyzes the data collected by the collection unit to identify dialects and points of mispronunciation. For example, the analysis unit can analyze audio data using speech recognition technology to identify patterns of dialects and mispronunciation. The analysis unit converts the audio data into text data and extracts specific patterns of dialects and mispronunciation. The analysis unit can also analyze video data and identify points of dialects and mispronunciation based on mouth movements and changes in facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. Step 3: The completion unit completes words that the other party does not understand during the meeting, based on the points identified by the analysis unit. For example, the completion unit can complete business terms and specific phrases in real time. The completion unit automatically recognizes business terms spoken during the meeting and completes their meaning in chat or text. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This facilitates smooth communication during the meeting and prevents misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or not. For example, the completion unit can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings.
[0069] (Example of form 2) The meeting completion system according to an embodiment of the present invention is a system for making meetings run more smoothly by compensating for insufficient explanations and slips of the tongue during meetings. This meeting completion system first learns the voice and mouth movements of the target participants in advance. This allows it to learn dialects and points of slip of the tongue. Next, it fills in words that the participants do not understand during the meeting via chat or text. For example, it can fill in business terms and specific phrases (e.g., "all-out team," "Apples to Apples," etc.). This system facilitates smooth communication during meetings and prevents misunderstandings and confusion. First, it learns the voice and mouth movements of the target participants in advance. At this time, it learns in detail the dialects and points of slip of the tongue that the participants speak. For example, it learns specific dialects and pronunciation habits and identifies points that should be supplemented based on them. This allows the system to automatically compensate even if the participants make a mistake during the meeting. Next, it fills in words that the participants do not understand during the meeting via chat or text. For example, if business terms or specific phrases come up, it supplements them in real time. For example, if the phrase "all-out team" comes up, it supplements it with "working together as a team." Furthermore, if the phrase "Apples to Apples" comes up, it is corrected to mean "comparison under the same conditions." This makes it easier for all meeting participants to understand and enables smoother communication. In addition, the system corrects what participants say in real time during the meeting. For example, if a participant says "Good morning," it is corrected to "Good morning." Similarly, if a participant says "Thank you for your cooperation today," it is corrected to "Thank you for your cooperation today." This ensures that what is said during the meeting is accurately conveyed and prevents misunderstandings and confusion. This system facilitates smooth communication during meetings and prevents misunderstandings and confusion. For example, if business terms or specific phrases come up, correcting them in real time makes it easier for all meeting participants to understand. Also, correcting what participants say in real time enables accurate communication. This improves the efficiency of meetings and allows for smoother progress.This allows the meeting completion system to fill in any gaps or slips of the tongue during meetings, making them run more smoothly.
[0070] The conference enhancement system according to this embodiment comprises a collection unit, an analysis unit, and an enhancement unit. The collection unit collects the voice and mouth movements of the subject. The collection unit can, for example, collect audio data of the subject speaking and video data of the subject's mouth movements. The collection unit can, for example, collect audio data and video data using a microphone or camera. The collection unit can also collect facial expression data of the subject when they speak. For example, the collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. The analysis unit analyzes the data collected by the collection unit to identify points of dialect and mispronunciation. The analysis unit can, for example, analyze audio data using speech recognition technology to identify patterns of dialect and mispronunciation. The analysis unit can, for example, convert audio data into text data and extract specific patterns of dialect and mispronunciation. The analysis unit can also analyze video data and identify points of dialect and mispronunciation based on changes in mouth movements and facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. The completion unit completes words that participants do not understand during a meeting, based on points identified by the analysis unit. The completion unit can, for example, complete business terms or specific phrases in real time. For example, the completion unit can automatically recognize business terms spoken during a meeting and complete their meaning in chat or text. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This facilitates smoother communication during meetings and prevents misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can perform completion using an AI model that takes business terms or specific phrases as input and outputs their meanings. As a result, the meeting completion system according to this embodiment can complete omissions and slips of the tongue during meetings, making meetings run more smoothly.
[0071] The data collection unit collects the subject's voice and mouth movements. For example, the data collection unit can collect audio data of the subject speaking and video data of their mouth movements. Specifically, the data collection unit uses a high-sensitivity microphone to clearly capture the subject's speech. This improves the quality of the audio data and makes processing easier in the analysis unit. The data collection unit also uses a high-resolution camera to record the subject's mouth movements and facial expressions in detail. This allows for the capture of even subtle changes in mouth movements and facial expressions, improving the accuracy in the analysis unit. Furthermore, the data collection unit can also collect facial expression data while the subject is speaking. For example, the data collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. This allows for a more accurate understanding of the subject's emotions and intentions. The data collection unit collects this data in real time and transmits it to a central database. This allows the collected data to be processed immediately in the analysis unit, enabling rapid supplementation during meetings. In addition, the data collection unit can use multiple microphones and cameras to collect data from multiple subjects simultaneously. This allows for efficient data collection and smooth processing in the analysis unit, even in large meetings or situations with multiple participants. Furthermore, the data acquisition unit can use noise cancellation technology to remove ambient noise and improve the quality of the audio data. This improves the accuracy of speech recognition in the analysis unit, and allows for more accurate interpolation in the interpolation unit.
[0072] The analysis unit analyzes the data collected by the collection unit to identify points of dialect and mispronunciation. For example, the analysis unit can analyze audio data using speech recognition technology to identify patterns of dialect and mispronunciation. Specifically, the analysis unit converts audio data into text data and extracts specific patterns of dialect and mispronunciation. Deep learning models are used for speech recognition technology, enabling highly accurate speech recognition. For example, the analysis unit divides the audio data into frames and extracts features from each frame. This allows it to capture subtle changes in the audio data and identify points of dialect and mispronunciation. The analysis unit can also analyze video data and identify points of dialect and mispronunciation based on mouth movements and changes in facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. This allows it to confirm the consistency between the subject's speech and mouth movements and identify points of dialect and mispronunciation. Furthermore, the analysis unit can analyze facial expression data to understand the subject's emotions and intentions. This allows the analysis unit to understand the underlying intent behind dialects and slips of the tongue, enabling more appropriate completion by the completion unit. The analysis unit provides these analysis results to the completion unit in real time, ensuring rapid completion during meetings. Furthermore, the analysis unit can learn patterns of dialects and slips of the tongue by utilizing past data and statistical information, thereby improving analysis accuracy. As a result, the analysis unit always performs highly accurate analysis based on the latest information, leading to more effective completion by the completion unit.
[0073] The completion unit completes the meaning of words that participants may not understand during a meeting, based on points identified by the analysis unit. For example, the completion unit can complete business terms and specific phrases in real time. Specifically, the completion unit automatically recognizes business terms spoken during a meeting and completes their meaning in chat or text. This allows participants to immediately understand the meaning of terms, making the meeting proceed more smoothly. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This allows participants to accurately understand the intent of what is being said, preventing misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or not. For example, the completion unit can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings. Specifically, it can use generative AI to generate the meaning of the input business terms and phrases in natural language and display them in chat or text. This allows the completion unit to complete any omissions or slips of the tongue during a meeting, making the meeting run more smoothly. Furthermore, the supplementary component can collect participant feedback and continuously improve the accuracy and effectiveness of its supplementary content. For example, it can collect participants' evaluations of the supplemented content and use them as training data for the AI model. This allows the supplementary component to always provide highly accurate supplementary information based on the latest data, thereby improving the quality of meetings.
[0074] The completion function can complete business terms and specific phrases in real time. For example, it can automatically recognize business terms spoken during a meeting and complete their meaning in chat or text. For instance, if the phrase "ROI" is used, the completion function will complete it with "Return on Investment." Similarly, if the phrase "KPI" is used, it can complete it with "Key Performance Indicator." Furthermore, if the phrase "Apples to Apples" is used, it can complete it with "Comparison under the same conditions." This real-time completion of business terms and specific phrases facilitates smoother communication during meetings. Some or all of the above processing in the completion function may be performed using AI, for example, or without AI. For example, the completion function can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings.
[0075] The completion function can convert specific phrases into more general expressions. For example, it can complete the phrase "team baseball" to "everyone working together." It can also complete the phrase "Apples to Apples" to "comparison under the same conditions." Furthermore, it can complete the phrase "ROI" to "return on investment." By converting specific phrases into more general expressions, it makes it easier for all meeting participants to understand. Some or all of the processing described above in the completion function may be performed using AI, for example, or not. For example, the completion function can perform completion using an AI model that takes a specific phrase as input and outputs a general expression.
[0076] The meeting completion system includes a correction unit that corrects the content of speech during a meeting in real time. For example, if a participant says "Good morning," the correction unit can correct it to "Good morning." For example, if a participant says "Thank you for your cooperation today," the correction unit can correct it to "Thank you for your cooperation today." The correction unit can also correct if a participant says "Excuse me, please wait a moment," to "Excuse me, please wait a moment." This allows for real-time correction of speech during a meeting, preventing misunderstandings and confusion. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can use an AI model that takes the participant's speech as input and outputs an accurate expression to perform the correction.
[0077] The correction unit can correct the content of the target person's speech to an accurate expression. For example, if the target person says "Ohayou gozaimashu," the correction unit can correct it to "Ohayou gozaimasu." For example, if the target person says "Kyou mo yoroshiu onegaishimasu," the correction unit can correct it to "Kyou mo yoroshiku onegaishimasu." Also, if the target person says "Sumimasen, chotto matte kudasai," the correction unit can correct it to "Sumimasen, chotto matte kudasai." By correcting the content of the target person's speech to an accurate expression, accurate communication becomes possible. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can perform corrections using an AI model that takes the content of the target person's speech as input and outputs an accurate expression.
[0078] The meeting completion system includes a presentation unit that provides the meaning of specific phrases when they appear. For example, if the phrase "team baseball" appears, the presentation unit might provide the meaning "working together as a team." If the phrase "Apples to Apples" appears, the presentation unit might provide the meaning "comparison under the same conditions." Furthermore, if the phrase "ROI" appears, the presentation unit might provide the meaning "return on investment." This makes it easier for all meeting participants to understand by providing the meaning of specific phrases when they appear. Some or all of the processing described above in the presentation unit may be performed using AI, for example, or not. For example, the presentation unit could use an AI model that takes a specific phrase as input and outputs its meaning to provide the presentation.
[0079] The data collection unit can estimate the subject's emotions and adjust the timing of voice and mouth movement data collection based on the estimated emotions. For example, if the subject is tense, the data collection unit can delay collection until the subject relaxes. For example, if the subject is relaxed, the data collection unit can start collecting data with tension. The data collection unit can also pause collection if the subject is excited and wait until they calm down. By adjusting the collection timing based on the subject's emotions, more accurate data collection becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can adjust the collection timing using an AI model that takes the subject's emotion data as input and outputs the collection timing.
[0080] The data collection unit can analyze the subject's past statement history and select the optimal collection method. For example, the data collection unit can collect data at specific time periods based on past statement history. For example, the data collection unit can prioritize collecting statements related to specific topics based on past statement history. The data collection unit can also analyze past statement history and adjust the collection interval according to the frequency of statements. This allows for the selection of the optimal collection method by analyzing past statement history. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can select a collection method using an AI model that takes past statement history data as input and outputs the optimal collection method.
[0081] The data collection unit can filter data based on the subject's current situation and areas of interest during collection. For example, if the subject is in a meeting, the collection unit will only collect statements related to the meeting. If the subject is interested in a particular project, the collection unit can prioritize collecting statements related to that project. The collection unit can also filter and collect statements related to a specific topic if the subject is discussing that topic. This allows for the collection of highly relevant data by filtering based on the subject's current situation and areas of interest. Some or all of the processing described above in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can perform filtering using an AI model that takes data on the subject's current situation and areas of interest as input and outputs filtered results.
[0082] The data collection unit can estimate the subject's emotions and determine the priority of data to collect based on the estimated emotions. For example, if the subject is stressed, the data collection unit will prioritize collecting important statements. If the subject is relaxed, the data collection unit can collect all statements equally. Also, if the subject is excited, the data collection unit can prioritize collecting emotional statements. In this way, important data can be collected preferentially by determining the data priority based on the subject's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can determine the priority using an AI model that takes the subject's emotion data as input and outputs the data priority.
[0083] The data collection unit can prioritize collecting highly relevant data based on the subject's geographical location information during data collection. For example, if the subject is in a specific region, the data collection unit will prioritize collecting statements related to that region. For example, if the subject is on the move, the data collection unit can prioritize collecting statements related to their destination. Furthermore, if the subject is in a specific location, the data collection unit can prioritize collecting statements related to that location. This allows for the collection of more useful data by collecting highly relevant data based on the subject's geographical location information. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can collect data using an AI model that takes the subject's geographical location information as input and outputs highly relevant data.
[0084] The data collection unit can analyze the subject's social media activity and collect relevant data during the collection process. For example, the data collection unit can collect words that the subject frequently uses on social media. For example, the data collection unit can collect statements related to topics that the subject is interested in on social media. The data collection unit can also collect statements related to accounts that the subject follows on social media. In this way, relevant data can be collected by analyzing the subject's social media activity. Some or all of the above processing in the data collection unit may be performed using AI, for example, or without AI. For example, the data collection unit can collect data using an AI model that takes the subject's social media activity data as input and outputs relevant data.
[0085] The analysis unit can estimate the subject's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the subject is tense, the analysis unit can provide a simple and highly visual analysis result. For example, if the subject is relaxed, the analysis unit can provide a detailed analysis result. Furthermore, if the subject is excited, the analysis unit can provide a visually stimulating analysis result. In this way, by adjusting the presentation of the analysis based on the subject's emotions, it is possible to provide analysis results that are easier to understand. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the presentation using an AI model that takes the subject's emotion data as input and outputs the presentation of the analysis.
[0086] The analysis unit can adjust the level of detail of the analysis based on the importance of the data during the analysis. For example, the analysis unit can perform a detailed analysis on important data. For example, it can perform a simplified analysis on less important data. The analysis unit can also determine the priority of the analysis according to the importance of the data. This allows for detailed analysis of important data by adjusting the level of detail based on the importance of the data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the level of detail using an AI model that takes data importance as input and outputs the level of detail of the analysis.
[0087] The analysis unit can apply different analysis algorithms depending on the data category during analysis. For example, the analysis unit can apply a specific analysis algorithm to business terms. For example, it can apply a different analysis algorithm to dialects. It can also apply a standard analysis algorithm to common phrases. By applying different analysis algorithms depending on the data category, more accurate analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can apply an algorithm using an AI model that takes data categories as input and outputs an analysis algorithm.
[0088] The analysis unit can estimate the subject's emotions and adjust the length of the analysis based on the estimated emotions. For example, if the subject is in a hurry, the analysis unit can provide a short, concise analysis result. For example, if the subject is relaxed, the analysis unit can provide a detailed analysis result. Furthermore, if the subject is excited, the analysis unit can provide a visually stimulating analysis result. By adjusting the length of the analysis based on the subject's emotions, more appropriate analysis results can be provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the length using an AI model that takes the subject's emotion data as input and outputs the length of the analysis.
[0089] The analysis unit can determine the priority of analysis based on the data collection timing during analysis. For example, the analysis unit may prioritize the analysis of the most recent data. For example, the analysis unit may perform analysis while referring to past data. The analysis unit can also adjust the priority of analysis according to the data collection timing. This allows for prioritization of the analysis of the most recent data by determining the priority of analysis based on the data collection timing. Some or all of the above-described processes in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can determine the priority using an AI model that takes the data collection timing as input and outputs the priority of analysis.
[0090] The analysis unit can adjust the order of analysis based on the relevance of the data during analysis. For example, the analysis unit can prioritize the analysis of highly relevant data. For example, the analysis unit can postpone the analysis of less relevant data. The analysis unit can also adjust the order of analysis according to the relevance of the data. This allows for prioritizing the analysis of highly relevant data by adjusting the order of analysis based on the relevance of the data. Some or all of the above processing in the analysis unit may be performed using AI, for example, or without AI. For example, the analysis unit can adjust the order using an AI model that takes the relevance of the data as input and outputs the order of analysis.
[0091] The completion unit can estimate the subject's emotions and adjust the way it expresses the completion based on the estimated emotions. For example, if the subject is tense, the completion unit will provide a simple and highly visual completion. If the subject is relaxed, the completion unit can provide a detailed completion. Furthermore, if the subject is excited, the completion unit can provide a visually stimulating completion. This allows for more easily understandable completion by adjusting the way it expresses the completion based on the subject's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the expression using an AI model that takes the subject's emotion data as input and outputs a way to express the completion.
[0092] The completion unit can adjust the level of detail of completion based on the importance of the words. For example, the completion unit can perform detailed completion for important words, and simplified completion for less important words. The completion unit can also determine the priority of completion based on the importance of the words. This allows for detailed completion of important words by adjusting the level of detail based on the importance of the words. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the level of detail using an AI model that takes the importance of words as input and outputs the level of detail of completion.
[0093] The completion unit can apply different completion algorithms depending on the word category during completion. For example, the completion unit can apply a specific completion algorithm to business terms. For example, it can apply a different completion algorithm to dialects. It can also apply a standard completion algorithm to common phrases. This allows for more accurate completion by applying different completion algorithms depending on the word category. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can apply an algorithm using an AI model that takes a word category as input and outputs a completion algorithm.
[0094] The completion unit can estimate the subject's emotions and adjust the length of the completion based on the estimated emotions. For example, if the subject is in a hurry, the completion unit will provide a short, concise completion. If the subject is relaxed, the completion unit can provide a detailed completion. Furthermore, if the subject is excited, the completion unit can provide a visually stimulating completion. By adjusting the length of the completion based on the subject's emotions, more appropriate completion becomes possible. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the length using an AI model that takes the subject's emotion data as input and outputs the length of the completion.
[0095] The completion unit can determine the priority of completion based on the timing of word usage during completion. For example, the completion unit may prioritize the most recent words. For example, the completion unit can perform completion while referring to past words. The completion unit can also adjust the priority of completion according to the timing of word usage. This allows for prioritizing the completion of the most recent words by determining the priority of completion based on the timing of word usage. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can determine the priority using an AI model that takes the timing of word usage as input and outputs the priority of completion.
[0096] The completion unit can adjust the order of completion based on the relevance of the words during completion. For example, the completion unit may prioritize completing highly relevant words. For example, the completion unit may postpone completing less relevant words. The completion unit can also adjust the order of completion according to the relevance of the words. This allows for prioritizing the completion of highly relevant words by adjusting the order of completion based on the relevance of the words. Some or all of the above processing in the completion unit may be performed using AI, for example, or without AI. For example, the completion unit can adjust the order using an AI model that takes the relevance of words as input and outputs the order of completion.
[0097] The editing unit can estimate the subject's emotions and adjust the expression of the correction based on the estimated emotions. For example, if the subject is tense, the editing unit can make a simple and highly visible correction. For example, if the subject is relaxed, the editing unit can make a detailed correction. Furthermore, if the subject is excited, the editing unit can make a visually stimulating correction. This allows for more easily understood corrections by adjusting the expression of the correction based on the subject's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the expression using an AI model that takes the subject's emotion data as input and outputs the expression of the correction.
[0098] The editing unit can adjust the level of detail of the edits based on the importance of the statements during the editing process. For example, the editing unit can perform detailed edits on important statements, and simplified edits on less important statements. The editing unit can also determine the priority of edits based on the importance of the statements. This allows for detailed edits on important statements by adjusting the level of detail based on the importance of the statements. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the level of detail using an AI model that takes the importance of a statement as input and outputs the level of detail of the edit.
[0099] The editing unit can apply different editing algorithms depending on the category of the utterance during editing. For example, the editing unit can apply a specific editing algorithm to business terms. For example, it can apply a different editing algorithm to dialects. It can also apply a standard editing algorithm to common phrases. By applying different editing algorithms depending on the category of the utterance, more accurate editing becomes possible. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can apply an algorithm using an AI model that takes the category of the utterance as input and outputs an editing algorithm.
[0100] The editing unit can estimate the subject's emotions and adjust the length of the edit based on the estimated emotions. For example, if the subject is in a hurry, the editing unit can make a short, concise edit. If the subject is relaxed, the editing unit can make a detailed edit. Furthermore, if the subject is excited, the editing unit can make a visually stimulating edit. By adjusting the length of the edit based on the subject's emotions, more appropriate edits can be made. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the length using an AI model that takes the subject's emotion data as input and outputs the length of the edit.
[0101] The editing unit can determine the priority of revisions based on when the statements were used. For example, the editing unit will prioritize revising the most recent statements. For example, the editing unit can perform revisions while referring to past statements. The editing unit can also adjust the priority of revisions according to when the statements were used. This allows the editing unit to prioritize revising the most recent statements by determining the priority of revisions based on when the statements were used. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can determine the priority using an AI model that takes the timing of statement use as input and outputs the priority of revisions.
[0102] The editing unit can adjust the order of editing based on the relevance of the statements during editing. For example, the editing unit can prioritize editing highly relevant statements. For example, the editing unit can postpone editing less relevant statements. The editing unit can also adjust the order of editing according to the relevance of the statements. This allows for prioritizing the editing of highly relevant statements by adjusting the order of editing based on the relevance of the statements. Some or all of the above processing in the editing unit may be performed using AI, for example, or without AI. For example, the editing unit can adjust the order using an AI model that takes the relevance of statements as input and outputs the order of editing.
[0103] The presentation unit can estimate the target's emotions and adjust the presentation's presentation style based on the estimated emotions. For example, if the target is tense, the presentation unit can provide a simple and highly visual presentation. If the target is relaxed, the presentation unit can provide a detailed presentation. Furthermore, if the target is excited, the presentation unit can provide a visually stimulating presentation. By adjusting the presentation style based on the target's emotions, a more easily understandable presentation becomes possible. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the presentation unit may be performed using AI, or not. For example, the presentation unit can adjust its presentation style using an AI model that takes the target's emotion data as input and outputs a presentation style.
[0104] The presentation unit can adjust the level of detail of the presentation based on the importance of the phrases. For example, the presentation unit can provide detailed presentations for important phrases, and simplified presentations for less important phrases. The presentation unit can also determine the priority of presentations based on the importance of the phrases. This allows for detailed presentations for important phrases by adjusting the level of detail based on the importance of the phrases. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can adjust the level of detail using an AI model that takes the importance of the phrases as input and outputs the level of detail of the presentation.
[0105] The presentation unit can estimate the target's emotions and adjust the length of the presentation based on the estimated emotions. For example, if the target is in a hurry, the presentation unit can provide a short, concise presentation. If the target is relaxed, the presentation unit can provide a detailed presentation. Furthermore, if the target is excited, the presentation unit can provide a visually stimulating presentation. By adjusting the length of the presentation based on the target's emotions, a more appropriate presentation becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can adjust the length using an AI model that takes the target's emotion data as input and outputs the length of the presentation.
[0106] The presentation unit can determine the presentation priority based on when the phrases were used. For example, the presentation unit can prioritize the presentation of the most recent phrases. For example, the presentation unit can make presentations while referring to past phrases. The presentation unit can also adjust the presentation priority according to when the phrases were used. This allows the presentation unit to prioritize the presentation of the most recent phrases by determining the presentation priority based on when the phrases were used. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can determine the priority using an AI model that takes the phrase usage time as input and outputs the presentation priority.
[0107] The presentation unit can adjust the presentation order based on the relevance of the phrases during presentation. For example, the presentation unit can prioritize the presentation of highly relevant phrases. For example, the presentation unit can postpone the presentation of less relevant phrases. The presentation unit can also adjust the presentation order according to the relevance of the phrases. This allows for the priority presentation of highly relevant phrases by adjusting the presentation order based on the relevance of the phrases. Some or all of the above processing in the presentation unit may be performed using AI, for example, or without AI. For example, the presentation unit can adjust the order using an AI model that takes the relevance of phrases as input and outputs the presentation order.
[0108] The presentation unit can provide the meaning of specific phrases when they appear during the presentation. For example, if the phrase "team baseball" appears, the presentation unit can provide the meaning "working together as a team." If the phrase "Apples to Apples" appears, the presentation unit can provide the meaning "comparison under the same conditions." Furthermore, if the phrase "ROI" appears, the presentation unit can provide the meaning "return on investment." By providing the meaning of specific phrases when they appear, it becomes easier for all meeting participants to understand. Some or all of the above processing in the presentation unit may be performed using AI, for example, or not. For example, the presentation unit can provide meaning using an AI model that takes a specific phrase as input and outputs its meaning.
[0109] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0110] The meeting enhancement system includes a translation unit that translates the content of a meeting in real time. For example, if a participant speaks in Japanese, the translation unit will translate it into English. For example, if a participant speaks in English, the translation unit can translate it into Japanese. Furthermore, if a participant speaks in French, the translation unit can translate it into Spanish. This facilitates communication between participants who speak different languages by translating the content of a meeting in real time. Some or all of the above processing in the translation unit may be performed using AI, for example, or without AI. For example, the translation unit can perform translation using an AI model that takes the content of the speech as input and outputs the translation result.
[0111] The meeting completion system includes a summarization unit that summarizes the content of discussions during a meeting. The summarization unit can, for example, summarize and present the key points of the meeting at the end of the meeting. The summarization unit can, for example, summarize the content of discussions during the meeting in real time and provide it to participants. The summarization unit can also provide summaries on specific topics during the meeting. This makes it easier for participants to grasp the key points by summarizing the content of discussions during the meeting. Some or all of the above processing in the summarization unit may be performed using, for example, AI, or not using AI. For example, the summarization unit can perform summarization using an AI model that takes the content of discussion as input and outputs a summary result.
[0112] The meeting completion system includes a recording unit that records the content of speech during a meeting. The recording unit can, for example, record the content of speech during the meeting as audio data. The recording unit can, for example, record the content of speech during the meeting as text data. The recording unit can also record the content of speech during the meeting as video data. This allows for later review of the content of speech during the meeting. Some or all of the above processing in the recording unit may be performed using, for example, AI, or without AI. For example, the recording unit can perform recording using an AI model that takes the content of speech as input and outputs recorded data.
[0113] The meeting enhancement system includes an analysis unit that analyzes the content of speech during meetings. The analysis unit can, for example, perform sentiment analysis on the content of speech during meetings to understand the emotional state of the participants. It can also, for example, perform topic analysis on the content of speech during meetings to identify key agenda items. Furthermore, the analysis unit can perform keyword analysis on the content of speech during meetings to extract important keywords. This makes it easier to understand the emotional state of participants and key agenda items by analyzing the content of speech during meetings. Some or all of the above processing in the analysis unit may be performed using, for example, AI, or without AI. For example, the analysis unit can perform analysis using an AI model that takes the content of speech as input and outputs analysis results.
[0114] The meeting enhancement system includes a visualization unit that visualizes the content of discussions during a meeting. The visualization unit can, for example, visualize the content of discussions during a meeting as graphs or charts. The visualization unit can also, for example, visualize the content of discussions during a meeting as a mind map. Furthermore, the visualization unit can visualize the content of discussions during a meeting as a timeline. By visualizing the content of discussions during a meeting, it becomes easier for participants to understand the content. Some or all of the above processing in the visualization unit may be performed using, for example, AI, or without AI. For example, the visualization unit can perform visualization using an AI model that takes the content of discussions as input and outputs visualized data.
[0115] The meeting enhancement system includes an evaluation unit that evaluates the content of statements made during a meeting in real time. The evaluation unit can, for example, evaluate the content of statements made during a meeting and determine the quality of those statements. The evaluation unit can, for example, score the content of statements made during a meeting and evaluate the importance of those statements. The evaluation unit can also rank the content of statements made during a meeting and evaluate the impact of those statements. This makes it easier to grasp the quality and importance of statements by evaluating the content of statements made during a meeting in real time. Some or all of the above processing in the evaluation unit may be performed using, for example, AI, or not using AI. For example, the evaluation unit can perform evaluations using an AI model that takes the content of statements as input and outputs evaluation results.
[0116] The meeting enhancement system includes a feedback unit that provides real-time feedback on what is said during a meeting. The feedback unit can, for example, provide immediate feedback on what is said during the meeting. The feedback unit can, for example, point out areas for improvement in what is said during the meeting. The feedback unit can also provide positive feedback on what is said during the meeting. This allows participants to improve the quality of their contributions by providing real-time feedback on what is said during the meeting. Some or all of the above processing in the feedback unit may be performed using, for example, AI, or not using AI. For example, the feedback unit can provide feedback using an AI model that takes the content of the comments as input and outputs feedback results.
[0117] The meeting completion system includes a comparison unit that compares the content of statements made during a meeting in real time. The comparison unit can, for example, compare the content of statements made during the meeting with past statements. The comparison unit can also, for example, compare the content of statements made during the meeting with statements made by other participants. Furthermore, the comparison unit can also compare the content of statements made during the meeting with industry standards. This makes it easier to evaluate the consistency and validity of statements by comparing the content of statements made during the meeting in real time. Some or all of the above processing in the comparison unit may be performed using, for example, AI, or not using AI. For example, the comparison unit can perform the comparison using an AI model that takes the content of statements as input and outputs the comparison results.
[0118] The meeting enhancement system includes a prediction unit that predicts the content of discussions during a meeting in real time. The prediction unit predicts, for example, what will be said next based on what has been said during the meeting. The prediction unit can also predict the next topic to be discussed based on what has been said during the meeting. Furthermore, the prediction unit can predict the next problem that will arise based on what has been said during the meeting. This allows for smoother meeting progress by predicting the content of discussions in real time. Some or all of the above processing in the prediction unit may be performed using, for example, AI, or not using AI. For example, the prediction unit can perform predictions using an AI model that takes the content of the discussion as input and outputs prediction results.
[0119] The meeting completion system includes a correction unit that corrects the content of speech during a meeting in real time. For example, if a participant says "Good morning," the correction unit can correct it to "Good morning." For example, if a participant says "Thank you for your cooperation today," the correction unit can correct it to "Thank you for your cooperation today." The correction unit can also correct if a participant says "Excuse me, please wait a moment," to "Excuse me, please wait a moment." This allows for real-time correction of speech during a meeting, preventing misunderstandings and confusion. Some or all of the above processing in the correction unit may be performed using AI, for example, or without AI. For example, the correction unit can use an AI model that takes the participant's speech as input and outputs an accurate expression to perform the correction.
[0120] The following briefly describes the processing flow for example form 2.
[0121] Step 1: The data collection unit collects the subject's voice and mouth movements. For example, the data collection unit can collect audio data of the subject speaking and video data of their mouth movements. The data collection unit uses microphones and cameras to collect audio and video data. The data collection unit can also collect facial expression data of the subject while they are speaking. For example, the data collection unit can capture the subject's facial movements with a high-resolution camera and collect facial expression data. Step 2: The analysis unit analyzes the data collected by the collection unit to identify dialects and points of mispronunciation. For example, the analysis unit can analyze audio data using speech recognition technology to identify patterns of dialects and mispronunciation. The analysis unit converts the audio data into text data and extracts specific patterns of dialects and mispronunciation. The analysis unit can also analyze video data and identify points of dialects and mispronunciation based on mouth movements and changes in facial expressions. For example, the analysis unit analyzes video data frame by frame and extracts features of mouth movements. Step 3: The completion unit completes words that the other party does not understand during the meeting, based on the points identified by the analysis unit. For example, the completion unit can complete business terms and specific phrases in real time. The completion unit automatically recognizes business terms spoken during the meeting and completes their meaning in chat or text. The completion unit can also convert specific phrases into more general expressions. For example, the completion unit completes the phrase "team baseball" to "everyone working together." This facilitates smooth communication during the meeting and prevents misunderstandings and confusion. Some or all of the above processing in the completion unit may be performed using AI, for example, or not. For example, the completion unit can perform completion using an AI model that takes business terms and specific phrases as input and outputs their meanings.
[0122] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0123] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0124] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0125] Each of the multiple elements described above, including the collection unit, analysis unit, supplementation unit, modification unit, and presentation unit, is implemented, for example, in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects the subject's voice and mouth movements using the microphone 38B and camera 42 of the smart device 14. The analysis unit analyzes the collected data using the identification processing unit 290 of the data processing unit 12 to identify dialects and points of mispronunciation. The supplementation unit uses the identification processing unit 290 of the data processing unit 12 to supplement words that the other party does not understand during a meeting. The modification unit uses the identification processing unit 290 of the data processing unit 12 to correct the subject's statements in real time. The presentation unit uses the display 40A of the smart device 14 to present the meaning of a specific phrase. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.
[0126] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0127] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0128] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0129] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0130] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0131] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0132] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0133] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0134] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0135] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0136] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0137] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0138] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0139] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0140] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0141] Each of the multiple elements described above, including the collection unit, analysis unit, supplementation unit, correction unit, and presentation unit, is implemented, for example, in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects the subject's voice and mouth movements using the microphone 238 and camera 42 of the smart glasses 214. The analysis unit analyzes the collected data using the identification processing unit 290 of the data processing unit 12 to identify dialects and points of mispronunciation. The supplementation unit uses the identification processing unit 290 of the data processing unit 12 to supplement words that the other party does not understand during a meeting. The correction unit corrects the subject's statements in real time using the identification processing unit 290 of the data processing unit 12. The presentation unit uses the display of the smart glasses 214 to present the meaning of a specific phrase. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.
[0142] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0143] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0144] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0145] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0146] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0147] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0148] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0149] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0150] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0151] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0152] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0153] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0154] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0155] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0156] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0157] Each of the multiple elements described above, including the collection unit, analysis unit, supplementation unit, modification unit, and presentation unit, is implemented, for example, by at least one of the headset terminal 314 and the data processing unit 12. For example, the collection unit collects the subject's voice and mouth movements using the microphone 238 and camera 42 of the headset terminal 314. The analysis unit analyzes the collected data by the identification processing unit 290 of the data processing unit 12 to identify dialects and points of mispronunciation. The supplementation unit uses the identification processing unit 290 of the data processing unit 12 to supplement words that the other party does not understand during the meeting. The modification unit uses the identification processing unit 290 of the data processing unit 12 to correct the subject's statements in real time. The presentation unit uses the display 343 of the headset terminal 314 to present the meaning of a specific phrase. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.
[0158] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0159] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0160] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0161] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0162] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0163] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0164] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0165] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0166] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0167] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0168] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0169] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0170] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0171] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0172] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0173] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0174] Each of the multiple elements described above, including the collection unit, analysis unit, supplementation unit, modification unit, and presentation unit, is implemented, for example, by at least one of the robot 414 and the data processing unit 12. For example, the collection unit collects the subject's voice and mouth movements using the robot 414's microphone 238 and camera 42. The analysis unit analyzes the collected data using the identification processing unit 290 of the data processing unit 12 to identify dialects and points of mispronunciation. The supplementation unit uses the identification processing unit 290 of the data processing unit 12 to supplement words that the other party does not understand during a meeting. The modification unit uses the identification processing unit 290 of the data processing unit 12 to correct the subject's statements in real time. The presentation unit uses the robot 414's display to present the meaning of a specific phrase. The correspondence between each unit and the device or control unit is not limited to the example described above, and various changes are possible.
[0175] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0176] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0177] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0178] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0179] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0180] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0181] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0182] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0183] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0184] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0185] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0186] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0187] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0188] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0189] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0190] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0191] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0192] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0193] (Note 1) The collection unit collects the voices and mouth movements of the subjects, An analysis unit analyzes the data collected by the aforementioned collection unit to identify points of dialect and mispronunciation, The system includes a complementation unit that complements words that the other party does not understand during a meeting, based on points identified by the analysis unit. A system characterized by the following features. (Note 2) The aforementioned supplementary unit is, Provides real-time supplementation for business terms and specific phrases. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned supplementary unit is, Convert a specific phrase into a general expression. The system described in Appendix 1, characterized by the features described herein. (Note 4) It includes a correction function that allows for real-time editing of comments made during meetings. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned modification section is, Correct the subject's statements to ensure accurate wording. The system described in Appendix 4, characterized by the features described herein. (Note 6) It includes a section that provides the meaning of a specific phrase when it appears. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is The system estimates the subject's emotions and adjusts the timing of voice and mouth movement data collection based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is Analyze the subject's past statements and select the most suitable data collection method. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is During data collection, filtering is performed based on the subjects' current situation and areas of interest. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is The system estimates the emotions of the subjects and determines the priority of data to collect based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned collection unit is During data collection, the system prioritizes collecting highly relevant data based on the geographical location information of the subjects. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned collection unit is During data collection, the social media activity of the target individuals is analyzed, and relevant data is collected. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned analysis unit, The system estimates the emotions of the subjects and adjusts the representation of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, During analysis, adjust the level of detail based on the importance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the data category. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, The system estimates the emotions of the subjects and adjusts the length of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when the data was collected. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, adjust the order of analysis based on the relevance of the data. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned supplementary unit is, The system estimates the subject's emotions and adjusts the method of expression for supplementation based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned supplementary unit is, During completion, adjust the level of detail based on the importance of the words. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned supplementary unit is, During completion, different completion algorithms are applied depending on the category of the word. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned supplementary unit is, The system estimates the subject's emotions and adjusts the length of the interpolation based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned supplementary unit is, During completion, the priority of completion is determined based on the time period in which the words were used. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned supplementary unit is, During completion, the order of completion is adjusted based on the relevance of the words. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned modification section is, The system estimates the emotions of the target individual and adjusts the expression of the correction based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned modification section is, When making revisions, adjust the level of detail based on the importance of the statement. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned modification section is, When making corrections, different correction algorithms are applied depending on the category of the statement. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned modification section is, The system estimates the subject's emotions and adjusts the length of the correction based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned modification section is, When making revisions, prioritize revisions based on when the statements were used. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned modification section is, When making revisions, adjust the order of revisions based on the relevance of the statements. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned display unit is, The system estimates the emotions of the target audience and adjusts the presentation method based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The aforementioned display unit is, When presenting, adjust the level of detail based on the importance of the phrase. The system described in Appendix 1, characterized by the features described herein. (Note 33) The aforementioned display unit is, The system estimates the target audience's emotions and adjusts the length of the presentation based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 34) The aforementioned display unit is, When presenting phrases, prioritize their presentation based on when they are used. The system described in Appendix 1, characterized by the features described herein. (Note 35) The aforementioned display unit is, When presenting, adjust the order of presentation based on the relevance of the phrases. The system described in Appendix 1, characterized by the features described herein. (Note 36) The aforementioned display unit is, When presenting information, if a specific phrase appears, its meaning will be explained. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0194] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. The collection unit collects the voices and mouth movements of the subjects, An analysis unit analyzes the data collected by the aforementioned collection unit to identify points of dialect and mispronunciation, The system includes a complementation unit that complements words that the other party does not understand during a meeting, based on points identified by the analysis unit. A system characterized by the following features.
2. The aforementioned supplementary unit is, Provides real-time supplementation for business terms and specific phrases. The system according to feature 1.
3. The aforementioned supplementary unit is, Convert a specific phrase into a general expression. The system according to feature 1.
4. It includes a correction function that allows for real-time editing of comments made during meetings. The system according to feature 1.
5. The aforementioned modification section is, Correct the subject's statements to ensure accurate wording. The system according to feature 4.
6. It includes a section that provides the meaning of a specific phrase when it appears. The system according to feature 1.
7. The aforementioned collection unit is The system estimates the subject's emotions and adjusts the timing of voice and mouth movement data collection based on the estimated emotions. The system according to feature 1.
8. The aforementioned collection unit is Analyze the subject's past statements and select the most suitable data collection method. The system according to feature 1.
9. The aforementioned collection unit is During data collection, filtering is performed based on the subjects' current situation and areas of interest. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A