System
The system addresses the lack of facial expression and pointer position integration in meeting minutes by using audio and video analysis units to create detailed records, improving information sharing and decision-making efficiency.
Patent Information
- Application Number
- JP2024132342
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
Conventional technology fails to incorporate participants' facial expressions, emotions, or pointer position information when generating detailed minutes from video and audio recordings of meetings and lectures, reducing the efficiency of information sharing and decision-making.
A system comprising an audio data analysis unit, a facial expression analysis unit, and a pointer analysis unit that integrates these analyses to generate detailed minutes, including participants' facial expressions, emotions, and pointer positions from video and audio data.
The system effectively generates detailed meeting minutes that incorporate facial expressions, emotions, and pointer positions, enhancing information sharing and decision-making by providing clear records of participants' thoughts and feelings.
Smart Images

Figure 2026029493000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional technology cannot incorporate participants' facial expressions, emotions, or pointer position information when generating detailed minutes from video and audio recordings of meetings and lectures, which can reduce the efficiency of information sharing and decision-making.
[0005] The system of the embodiment aims to generate detailed minutes of meetings and lectures that incorporate participants' facial expressions, emotions, and pointer position information from video and audio data of meetings and lectures. [Means for solving the problem]
[0006] The system according to the embodiment includes an audio data analysis unit, a facial expression analysis unit, a pointer analysis unit, and a minutes generation unit. The audio data analysis unit analyzes audio data and transcribes it. The facial expression analysis unit analyzes video data to identify the facial expressions and emotions of participants. The pointer analysis unit analyzes pointer position data to identify important points in the materials. The minutes generation unit integrates the analysis results of the audio data analysis unit, facial expression analysis unit, and pointer analysis unit to generate detailed minutes. [Effects of the Invention]
[0007] The system according to the embodiment can generate detailed minutes of meetings and lectures that incorporate participants' facial expressions, emotions, and pointer position information from video and audio recordings of the meetings and lectures. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION
[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0010] First, the terms used in the following description will be explained.
[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).
[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.
[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example 1) The emotext system according to an embodiment of the present invention automatically transcribes video and audio data and materials from meetings and lectures, generating detailed minutes that incorporate the facial expressions and emotions of the participants. This allows the emotext system to record the details of meetings and lectures, and share the thoughts and feelings of the participants.
[0029] The emotext system according to the embodiment includes an audio data analysis unit, a facial expression analysis unit, a pointer analysis unit, and a minutes generation unit. The audio data analysis unit analyzes audio data and transcribes it. For example, the audio data analysis unit analyzes recorded data of a meeting or lecture and converts what is being said into text. The audio data analysis unit analyzes audio data in MP3 or WAV format and transcribes what is being said. The audio data analysis unit can also analyze real-time audio streams and instantly convert what is being said into text. The facial expression analysis unit analyzes video data to identify the facial expressions and emotions of participants. For example, the facial expression analysis unit analyzes video data of a meeting or lecture and recognizes the facial expressions of participants. The facial expression analysis unit analyzes video data in MP4 or AVI format and identifies the facial expressions of participants. The facial expression analysis unit can also analyze real-time video streams and instantly identify the facial expressions of participants. The pointer analysis unit analyzes pointer position data to identify important points in the materials. For example, the pointer analysis unit analyzes the movement of a speaker's pointer to identify important points in the materials. The pointer analysis unit, for example, analyzes mouse coordinate data and laser pointer position data to identify important points. The pointer analysis unit can also analyze the frequency and speed of pointer movement to identify emphasized parts. The minutes generation unit generates detailed minutes by integrating the analysis results of the audio data analysis unit, facial expression analysis unit, and pointer analysis unit. For example, the minutes generation unit generates detailed minutes by integrating the results of audio data transcription, facial expression analysis, and pointer analysis. The minutes generation unit generates minutes that include, for example, the content of remarks, the facial expressions of participants, and important points of materials. Furthermore, the minutes generation unit supports information sharing and decision-making among participants based on the generated minutes. As a result, the emotext system according to the embodiment can automatically transcribe video and audio data and materials of meetings and lectures and generate detailed minutes that incorporate the facial expressions and emotions of participants. For example, the emotext system can clearly record the content and key points of a meeting, allowing participants to share their thoughts and the sense of presence of the meeting.In addition, the emotext system can support information sharing and decision-making among participants based on the generated minutes.
[0030] The audio data analysis unit automatically removes background sounds and noise from audio data, making speech clearer. The audio data analysis unit, for example, analyzes audio data and uses technology to automatically filter background sounds and noise. For example, it removes the sound of a conference room air conditioner and other external noise to make speech clearer. The audio data analysis unit also applies a noise removal algorithm to the audio data to extract only the speaker's voice. For example, even when multiple speakers are speaking simultaneously, each speaker's voice can be extracted individually. The audio data analysis unit also uses audio data noise removal technology to achieve high-quality transcription regardless of the recording environment. For example, it accurately records speech even in outdoor meetings or noisy environments. By removing background sounds and noise from the audio data, speech can be recorded more clearly.
[0031] The audio data analysis unit transcribes audio data in real time, allowing minutes to be generated instantly during a meeting. The audio data analysis unit, for example, builds a system that analyzes audio data in real time and instantly converts spoken content into text. For example, transcription occurs as soon as a speech is made during a meeting. The audio data analysis unit also uses real-time transcription technology to instantly share the minutes generated during a meeting with participants. For example, the minutes can be checked immediately after the meeting ends. The audio data analysis unit also introduces a real-time transcription system that instantly records what is said during a meeting, significantly reducing the time it takes to create minutes. For example, this eliminates the need to create minutes after the meeting ends. As a result, minutes can be generated in real time, allowing them to be checked immediately after the meeting ends.
[0032] The audio data analysis unit can simultaneously analyze audio data in different languages and generate minutes that support multiple languages. The audio data analysis unit, for example, builds a system that simultaneously analyzes audio data in different languages and converts the content of speech in each language into text. For example, it transcribes speech in English and Japanese simultaneously. The audio data analysis unit also uses multilingual transcription technology to instantly analyze multiple languages spoken during a meeting and reflect them in the minutes. For example, it translates and records English speech into Japanese. The audio data analysis unit also introduces a system that analyzes audio data in different languages and generates minutes that support multiple languages. For example, it transcribes multiple languages used in international conferences simultaneously. This makes it possible to simultaneously analyze audio data in different languages and generate minutes that support multiple languages.
[0033] The voice data analysis unit can display the results of transcription of voice data as visual infographics to promote understanding. The voice data analysis unit, for example, builds a system that displays the results of transcription of voice data as visual infographics. For example, the content of speech is shown in a graph or chart. The voice data analysis unit also converts the transcription results into infographics to make the content of the meeting easier to understand visually. For example, the speaking time of each speaker is displayed in a pie chart. The voice data analysis unit also uses infographics to visually display the transcription results, allowing participants to intuitively grasp the content of the meeting. For example, important points are highlighted. In this way, by visually displaying the transcription results, participants can intuitively grasp the content of the meeting.
[0034] The facial expression analysis unit integrates video data from different camera angles to perform more accurate facial expression and emotion analysis. For example, the facial expression analysis unit builds a system that integrates video data from different camera angles to improve the accuracy of facial expression and emotion analysis. For example, it analyzes video from multiple cameras simultaneously. The facial expression analysis unit also integrates video data from different camera angles to more accurately analyze participants' facial expressions and emotions. For example, it combines frontal and side video for analysis. The facial expression analysis unit also introduces a system that integrates video from multiple cameras to improve the accuracy of facial expression and emotion analysis. For example, it analyzes video from different angles in real time. By integrating video data from different camera angles, more accurate facial expression and emotion analysis becomes possible.
[0035] The facial expression analysis unit can provide participants with the results of the facial expression and emotional analysis as post-meeting feedback, encouraging self-improvement. For example, the facial expression analysis unit could build a system that provides participants with post-meeting feedback based on the results of the facial expression and emotional analysis. For example, it could provide a report of changes in emotions during the meeting. The facial expression analysis unit could also provide feedback based on the analysis results to help participants improve themselves. For example, it could offer advice such as, "There were many tense moments during the meeting. Try some relaxation techniques." The facial expression analysis unit could also provide feedback based on the results of the facial expression and emotional analysis, introducing a system that helps participants improve themselves. For example, it could display changes in emotions in a graph to clarify areas for improvement. By providing participants with analysis results as feedback, participants can improve themselves.
[0036] The pointer analysis unit can analyze the speed and frequency of pointer movement to identify the speaker's emphasis points. The pointer analysis unit, for example, analyzes the speed of pointer movement to build a system that identifies the points the speaker wants to emphasize. For example, it records points where the pointer moves quickly as emphasis points. The pointer analysis unit also analyzes the frequency of pointer movement to identify points that the speaker has pointed to multiple times. For example, if the same point is pointed to multiple times, it records that point as an important point. The pointer analysis unit also introduces a system that analyzes the speed and frequency of pointer movement to comprehensively evaluate the speaker's emphasis points. For example, it records points with high speed and frequency as emphasis points. In this way, the speaker's emphasis points can be identified by analyzing the speed and frequency of pointer movement.
[0037] The pointer analysis unit can synchronize pointer movement with remarks to automatically highlight important points. The pointer analysis unit, for example, synchronizes pointer movement with remarks to build a system that automatically highlights important points. For example, when a speaker points to a specific part while explaining, that part is highlighted. The pointer analysis unit also analyzes remarks and pointer movement in real time to automatically highlight important points. For example, it records the information as "Speaker C: Point to page 3, figure 2 in the materials." The pointer analysis unit also introduces a system that synchronizes pointer movement with remarks to automatically highlight important points. For example, when a specific part of a slide is pointed to, that part is highlighted. This makes it possible to automatically highlight important points by synchronizing pointer movement with remarks.
[0038] The pointer analysis unit can reproduce pointer movement as a 3D model, making it easier to understand visually. The pointer analysis unit, for example, builds a system that reproduces pointer movement as a 3D model. For example, the movement of a speaker pointing to a specific part of a document is displayed as a 3D model. The pointer analysis unit also reproduces pointer movement as a 3D model, making it easier to understand visually. For example, the movement of a speaker pointing to a specific part of a slide is displayed as a 3D animation. The pointer analysis unit also introduces a system that uses 3D models to visually reproduce pointer movement. For example, the speaker's hand movement is displayed as a 3D model, making the pointing part clear. In this way, reproducing pointer movement as a 3D model makes it easier to understand visually.
[0039] The pointer analysis unit can integrate and analyze pointer movements of different presentation tools and provide common analysis results. The pointer analysis unit, for example, builds a system that integrates and analyzes pointer movements of different presentation tools. For example, it simultaneously analyzes pointer movements of PowerPoint and Keynote. The pointer analysis unit also integrates pointer movements of different presentation tools and provides common analysis results. For example, it analyzes pointer movements used in multiple tools as a single piece of data. The pointer analysis unit also introduces a system that integrates and analyzes pointer movements of presentation tools and provides common analysis results. For example, it displays pointer movements of different tools in a single view. This makes it possible to integrate and analyze pointer movements of different presentation tools and provide common analysis results.
[0040] The pointer analysis unit can feed back the results of the analysis of pointer movement to participants in real time, thereby promoting understanding. The pointer analysis unit, for example, builds a system that feeds back the results of the analysis of pointer movement to participants in real time. For example, the part indicated by the speaker is instantly displayed to the participants. The pointer analysis unit also feeds back the results of the analysis of pointer movement in real time, promoting understanding among participants. For example, when indicating a specific part of a slide, that part is highlighted. The pointer analysis unit also introduces a system that feeds back the results of the analysis of pointer movement in real time, allowing participants to instantly understand the speaker's intention. For example, the part indicated is highlighted. In this way, by feeding back the results of the analysis of pointer movement in real time, understanding among participants can be promoted.
[0041] The minutes generation unit can automatically evaluate the importance of the content of statements and highlight important statements when generating minutes. The minutes generation unit, for example, builds a system that automatically evaluates the importance of the content of statements and highlights important statements. For example, the importance is evaluated based on the frequency of keywords and phrases. The minutes generation unit also analyzes the importance of the content of statements and highlights important statements. For example, it records "Speaker A: Today's agenda item is XX (important)." The minutes generation unit also introduces a system that automatically evaluates the importance of the content of statements and reflects this in the minutes. For example, important statements are highlighted in bold or color. In this way, by automatically evaluating the importance of the content of statements and highlighting important statements, it becomes easier for readers of the minutes to grasp important information.
[0042] The minutes generation unit can provide minutes in a multimedia format synchronized with audio and video. The minutes generation unit, for example, builds a system that provides minutes in a multimedia format synchronized with audio and video. For example, it adds audio and video corresponding to the content of remarks as links. The minutes generation unit also provides minutes synchronized with audio and video so that participants can visually check the content of remarks. For example, it records "Speaker A: Today's agenda is ____ (audio link)." The minutes generation unit also introduces a system that provides minutes in a multimedia format so that participants can check the minutes while referring to the audio and video. For example, it embeds video corresponding to the content of remarks. By providing minutes in a multimedia format synchronized with audio and video, participants can visually check the content of remarks.
[0043] The minutes generation unit can automatically generate minutes in different formats (PDF, Word, HTML, etc.) to meet user needs. The minutes generation unit, for example, builds a system that automatically generates minutes in different formats. For example, minutes are generated in PDF, Word, and HTML formats. The minutes generation unit also provides minutes in different formats according to user needs. For example, minutes in PDF format are sent by email. The minutes generation unit also introduces a system that automatically generates minutes in different formats and allows users to select. For example, minutes in HTML format are published on a website. This makes it possible to meet user needs by automatically generating minutes in different formats.
[0044] The minutes generation unit can add a function to read out the generated minutes aloud through an AI assistant. For example, the minutes generation unit builds a system that reads out the generated minutes aloud through an AI assistant. For example, the generated minutes are read out using speech synthesis technology. The minutes generation unit also adds a function to read out the contents of the minutes aloud using an AI assistant. For example, it plays out loud "Speaker A: Today's agenda is ____." The minutes generation unit also introduces a function to read out the generated minutes aloud so that participants can listen to them. For example, it plays out the minutes aloud after the meeting ends. This allows participants to listen to the generated minutes by reading them out loud.
[0045] The minutes generation unit can automatically generate the agenda for the next meeting based on the generated minutes. The minutes generation unit, for example, builds a system that automatically generates the agenda for the next meeting based on the generated minutes. For example, it analyzes the contents of the minutes and automatically sets the next agenda. The minutes generation unit also automatically generates the agenda for the next meeting based on the contents of the minutes and shares it with the participants. For example, it records "Next agenda: Progress report on XX." The minutes generation unit also introduces a system that automatically generates the agenda for the next meeting based on the generated minutes. For example, it extracts important points from the minutes and sets them as the next agenda. In this way, meeting preparations are made more efficient by automatically generating the agenda for the next meeting based on the generated minutes.
[0046] The minutes generation unit can analyze the contents of the minutes and automatically extract unresolved issues and action items. For example, the minutes generation unit builds a system that analyzes the contents of the minutes and automatically extracts unresolved issues and action items. For example, it extracts items marked "unresolved" from the minutes. The minutes generation unit also automatically extracts unresolved issues and action items based on the contents of the minutes and notifies participants. For example, it records "Action item: Investigate XX." The minutes generation unit also introduces a system that analyzes the contents of the minutes and automatically extracts unresolved issues and action items. For example, it extracts important tasks from the minutes and makes a list. This makes meeting follow-up more efficient by analyzing the contents of the minutes and automatically extracting unresolved issues and action items.
[0047] The minutes generation unit can share minutes on the cloud and enable real-time collaborative editing. The minutes generation unit, for example, builds a system that shares minutes on the cloud and enables real-time collaborative editing. For example, it provides a collaborative editing function like Google Docs. The minutes generation unit also shares minutes on the cloud and enables participants to edit them in real time. For example, minutes can be updated during the meeting so that everyone can see the latest information. The minutes generation unit also introduces a cloud-based minutes sharing system that enables real-time collaborative editing. For example, participants can simultaneously add comments and corrections to the minutes. This allows minutes to be shared on the cloud and enables real-time collaborative editing so that participants can simultaneously add comments and corrections to the minutes.
[0048] The minutes generation department can work with different project management tools to automatically register the contents of the minutes as tasks. For example, the minutes generation department could work with different project management tools to build a system that automatically registers the contents of the minutes as tasks. For example, it could automatically register action items from the minutes in Trello or Asana. The minutes generation department could also automatically reflect the contents of the minutes in the project management tool and register them as tasks. For example, it could add "Action item: Investigate XX" to the project management tool. The minutes generation department could also work with project management tools to introduce a system that automatically registers the contents of the minutes as tasks. For example, it could automatically register important tasks from the minutes in JIRA. This would allow it to work with different project management tools and automatically register the contents of the minutes as tasks, making meeting follow-up more efficient.
[0049] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0050] The emotext system can further include a summarization unit that summarizes what participants have said. For example, the summarization unit can shorten long speeches and extract important points. For example, it can summarize what a speaker has said at length into a few lines. The summarization unit can also summarize what has been said in real time and provide a summary immediately during the meeting. For example, a summary can be displayed as soon as a speech is finished. The summarization unit can also reflect the summary of what has been said in the minutes, allowing readers to immediately grasp the important points. For example, it can record "Speaker A: Progress report on new project (summary)." In this way, summarization of what has been said improves the readability of the minutes.
[0051] The emotext system can further include a speech frequency analysis unit that analyzes the frequency of speech made by each speaker. The speech frequency analysis unit, for example, counts the number of times each participant speaks and reflects this in the minutes. For example, it may record "Speaker B: 10 speeches." The speech frequency analysis unit can also evaluate the balance of speech made during a meeting based on speech frequency. For example, if a particular participant speaks a lot, it will record that information. The speech frequency analysis unit can also visualize the results of the speech frequency analysis in graphs and charts, making it easier to understand the progress of the meeting. For example, it could display the number of speeches as a bar graph. This makes it easier to understand the progress of the meeting by analyzing speech frequency.
[0052] The emotext system can further include a classification unit that classifies the content of participants' comments. For example, the classification unit classifies the content of comments by theme and reflects this in the minutes. For example, it may record "Speaker C: About a new project (theme: project)." The classification unit can also classify the content of comments in real time and provide the classification results immediately during the meeting. For example, the theme may be displayed as soon as a comment is finished. The classification unit can also visualize the classification results of the content of comments in graphs or charts, making it easier to understand the content of the meeting. For example, the number of comments for each theme may be displayed in a pie chart. In this way, classifying the content of comments makes it easier to understand the content of the minutes.
[0053] The emotext system can further include a translation unit that translates the content of participants' statements. The translation unit, for example, automatically translates the content of statements made in different languages and reflects it in the minutes. For example, it may record "Speaker D: I'm looking forward to the new project (English)" and translate it into Japanese. The translation unit can also translate the content of statements in real time and provide the translation results immediately during the meeting. For example, the translation may be displayed as soon as the statement is finished. The translation unit also reflects the translation results in the minutes, making it easier for participants speaking different languages to understand the content of the meeting. For example, it may record "Speaker E: There are problems with this proposal (Japanese)" and translate it into English. By translating the content of statements, it makes it easier for participants speaking different languages to understand the content of the meeting.
[0054] The emotext system can further include an evaluation unit that evaluates the content of participants' comments. The evaluation unit, for example, evaluates the importance and usefulness of the content of the comments and reflects this in the minutes. For example, it may record "Speaker F: New project proposal (important)." The evaluation unit can also evaluate the content of comments in real time and provide the evaluation results immediately during the meeting. For example, the evaluation may be displayed as soon as the comment is finished. The evaluation unit may also visualize the evaluation results in graphs or charts, making the content of the meeting easier to understand. For example, it may highlight important comments. In this way, evaluating the content of the comments makes it easier to understand the content of the minutes.
[0055] The processing flow of the first embodiment will be briefly explained below.
[0056] Step 1: The audio data analysis unit analyzes audio data and transcribes it. For example, it analyzes recordings of meetings and lectures and converts what is being said into text. The audio data analysis unit analyzes audio data in MP3 or WAV format and transcribes what is being said. It can also analyze real-time audio streams and instantly convert what is being said into text. Step 2: The facial expression analyzer analyzes the video data to identify the participants' facial expressions and emotions. For example, it analyzes video data of a meeting or lecture to recognize the participants' facial expressions. The facial expression analyzer analyzes video data in MP4 or AVI format to identify the participants' facial expressions. It can also analyze real-time video streams to instantly identify the participants' facial expressions. Step 3: The pointer analysis unit analyzes the pointer position data to identify important points in the document. For example, it analyzes the movement of the speaker's pointer to identify important points in the document. The pointer analysis unit analyzes mouse coordinate data and laser pointer position data to identify important points. It can also analyze the frequency and speed of pointer movement to identify emphasized points. Step 4: The minutes generator combines the results of the voice data analysis, facial expression analysis, and pointer analysis to generate detailed minutes. For example, it combines the results of voice data transcription, facial expression analysis, and pointer analysis to generate detailed minutes that include the content of remarks, participants' facial expressions, and key points of the materials. The generated minutes also support information sharing and decision-making among participants.
[0057] (Example 2) The emotext system according to an embodiment of the present invention automatically transcribes video and audio data and materials from meetings and lectures, generating detailed minutes that incorporate the facial expressions and emotions of the participants. This allows the emotext system to record the details of meetings and lectures, and share the thoughts and feelings of the participants.
[0058] The emotext system according to the embodiment includes an audio data analysis unit, a facial expression analysis unit, a pointer analysis unit, and a minutes generation unit. The audio data analysis unit analyzes audio data and transcribes it. For example, the audio data analysis unit analyzes recorded data of a meeting or lecture and converts what is being said into text. The audio data analysis unit analyzes audio data in MP3 or WAV format and transcribes what is being said. The audio data analysis unit can also analyze real-time audio streams and instantly convert what is being said into text. The facial expression analysis unit analyzes video data to identify the facial expressions and emotions of participants. For example, the facial expression analysis unit analyzes video data of a meeting or lecture and recognizes the facial expressions of participants. The facial expression analysis unit analyzes video data in MP4 or AVI format and identifies the facial expressions of participants. The facial expression analysis unit can also analyze real-time video streams and instantly identify the facial expressions of participants. The pointer analysis unit analyzes pointer position data to identify important points in the materials. For example, the pointer analysis unit analyzes the movement of a speaker's pointer to identify important points in the materials. The pointer analysis unit, for example, analyzes mouse coordinate data and laser pointer position data to identify important points. The pointer analysis unit can also analyze the frequency and speed of pointer movement to identify emphasized parts. The minutes generation unit generates detailed minutes by integrating the analysis results of the audio data analysis unit, facial expression analysis unit, and pointer analysis unit. For example, the minutes generation unit generates detailed minutes by integrating the results of audio data transcription, facial expression analysis, and pointer analysis. The minutes generation unit generates minutes that include, for example, the content of remarks, the facial expressions of participants, and important points of materials. Furthermore, the minutes generation unit supports information sharing and decision-making among participants based on the generated minutes. As a result, the emotext system according to the embodiment can automatically transcribe video and audio data and materials of meetings and lectures and generate detailed minutes that incorporate the facial expressions and emotions of participants. For example, the emotext system can clearly record the content and key points of a meeting, allowing participants to share their thoughts and the sense of presence of the meeting.In addition, the emotext system can support information sharing and decision-making among participants based on the generated minutes.
[0059] The audio data analysis unit automatically removes background sounds and noise from audio data, making speech clearer. The audio data analysis unit, for example, analyzes audio data and uses technology to automatically filter background sounds and noise. For example, it removes the sound of a conference room air conditioner and other external noise to make speech clearer. The audio data analysis unit also applies a noise removal algorithm to the audio data to extract only the speaker's voice. For example, even when multiple speakers are speaking simultaneously, each speaker's voice can be extracted individually. The audio data analysis unit also uses audio data noise removal technology to achieve high-quality transcription regardless of the recording environment. For example, it accurately records speech even in outdoor meetings or noisy environments. By removing background sounds and noise from the audio data, speech can be recorded more clearly.
[0060] The voice data analysis unit can analyze the tone and speed of a speaker's voice to estimate their emotions and level of tension, which can be reflected in the recording. For example, the voice data analysis unit analyzes the speaker's tone of voice to estimate changes in their emotions. For example, it records information such as a higher voice indicating excitement and a lower voice indicating calm. The voice data analysis unit also analyzes the speaker's speaking speed to estimate their level of tension. For example, it records information such as a faster speaking speed indicating tension, and a slower speaking speed indicating relaxed. The voice data analysis unit also combines the results of the voice tone and speed analysis to comprehensively evaluate the speaker's emotions and level of tension, which can be reflected in the minutes. For example, it records "Speaker A: Today's agenda item is XX (slightly excited)." In this way, the speaker's emotions and level of tension can be estimated and reflected in the minutes, providing more detailed information.
[0061] The voice data analysis unit can use an emotion estimation function to analyze the speaker's emotion and assign emotion tags to the speech content. The voice data analysis unit, for example, analyzes voice data and uses an algorithm to estimate the speaker's emotion. For example, it identifies emotions such as joy, anger, and sadness and assigns tags to the speech content. The voice data analysis unit also uses the emotion estimation function to analyze the speaker's emotion in real time and reflects it in the minutes. For example, it records, "Speaker B: I'm looking forward to the new project (joy)." The voice data analysis unit also assigns emotion tags to the speech content, making it easier for readers of the minutes to understand the speaker's emotion. For example, it records, "Speaker C: There are problems with this proposal (anger)." By assigning emotion tags to the speech content, it makes it easier for readers of the minutes to understand the speaker's emotion.
[0062] The audio data analysis unit transcribes audio data in real time, allowing minutes to be generated instantly during a meeting. The audio data analysis unit, for example, builds a system that analyzes audio data in real time and instantly converts spoken content into text. For example, transcription occurs as soon as a speech is made during a meeting. The audio data analysis unit also uses real-time transcription technology to instantly share the minutes generated during a meeting with participants. For example, the minutes can be checked immediately after the meeting ends. The audio data analysis unit also introduces a real-time transcription system that instantly records what is said during a meeting, significantly reducing the time it takes to create minutes. For example, this eliminates the need to create minutes after the meeting ends. As a result, minutes can be generated in real time, allowing them to be checked immediately after the meeting ends.
[0063] The audio data analysis unit can simultaneously analyze audio data in different languages and generate minutes that support multiple languages. The audio data analysis unit, for example, builds a system that simultaneously analyzes audio data in different languages and converts the content of speech in each language into text. For example, it transcribes speech in English and Japanese simultaneously. The audio data analysis unit also uses multilingual transcription technology to instantly analyze multiple languages spoken during a meeting and reflect them in the minutes. For example, it translates and records English speech into Japanese. The audio data analysis unit also introduces a system that analyzes audio data in different languages and generates minutes that support multiple languages. For example, it transcribes multiple languages used in international conferences simultaneously. This makes it possible to simultaneously analyze audio data in different languages and generate minutes that support multiple languages.
[0064] The voice data analysis unit can display the results of transcription of voice data as visual infographics to promote understanding. The voice data analysis unit, for example, builds a system that displays the results of transcription of voice data as visual infographics. For example, the content of speech is shown in a graph or chart. The voice data analysis unit also converts the transcription results into infographics to make the content of the meeting easier to understand visually. For example, the speaking time of each speaker is displayed in a pie chart. The voice data analysis unit also uses infographics to visually display the transcription results, allowing participants to intuitively grasp the content of the meeting. For example, important points are highlighted. In this way, by visually displaying the transcription results, participants can intuitively grasp the content of the meeting.
[0065] In addition to analyzing facial expressions, the facial expression analysis unit can also analyze body movements and gestures to obtain more detailed emotional information. For example, the facial expression analysis unit constructs a system that analyzes body movements and gestures in addition to analyzing facial expressions. For example, it analyzes hand movements and changes in posture to obtain emotional information. The facial expression analysis unit also analyzes body movements and gestures and integrates them with the results of the facial expression analysis to record detailed emotional information. For example, it records information such as a hand waving motion indicating joy. The facial expression analysis unit also introduces a system that combines facial expression analysis and body movement analysis to comprehensively evaluate the emotions of participants. For example, it emphasizes joy when a smile and hand movements match. In this way, more detailed emotional information can be obtained by analyzing body movements and gestures as well.
[0066] The facial expression analysis unit can compare the past facial expression data of participants and track and record changes in their emotions. The facial expression analysis unit, for example, builds a system that saves past facial expression data of participants and compares it with their current expressions. For example, it compares facial expression data from past meetings with their current expressions to record changes in emotions. The facial expression analysis unit also tracks changes in participants' emotions based on past facial expression data and reflects this in the minutes. For example, it records "Participant B: smiling (expressionless in the previous meeting)." The facial expression analysis unit also introduces a system that analyzes the history of facial expression data and tracks changes in emotions. For example, it compares changes in emotions with past data and displays them in a graph. This makes it possible to track and record changes in emotions by comparing them with past facial expression data.
[0067] The facial expression analysis unit uses the emotion estimation function to analyze the emotions of participants in real time and instantly record emotional changes. The facial expression analysis unit, for example, uses the emotion estimation function to build a system that analyzes the emotions of participants in real time. For example, it analyzes camera footage and instantly records emotional changes. The facial expression analysis unit also analyzes emotions in real time and instantly reflects the emotional changes in the minutes. For example, it records "Participant C: changed from smiling to surprised while speaking." The facial expression analysis unit also uses the emotion estimation function to introduce a system that monitors the emotional changes of participants in real time and reflects them in the minutes. For example, it displays the emotional changes in a graph. This allows the emotion estimation function to record emotional changes in real time.
[0068] The facial expression analysis unit visualizes the results of the analysis of facial expressions and emotions as an emotion map, making it possible to grasp the flow of emotions throughout the entire meeting. For example, the facial expression analysis unit builds a system that visualizes the results of the analysis of facial expressions and emotions as an emotion map. For example, it shows changes in emotions during the meeting using colors or graphs. The facial expression analysis unit also uses the emotion map to visually display the flow of emotions throughout the entire meeting. For example, it adds a graph showing changes in emotions along a time axis. The facial expression analysis unit also introduces a system that updates the emotion map in real time, allowing changes in emotions during the meeting to be immediately grasped. For example, it displays changes in emotions in real time. This makes it possible to visually grasp the flow of emotions throughout the entire meeting using the emotion map.
[0069] The facial expression analysis unit integrates video data from different camera angles to perform more accurate facial expression and emotion analysis. For example, the facial expression analysis unit builds a system that integrates video data from different camera angles to improve the accuracy of facial expression and emotion analysis. For example, it analyzes video from multiple cameras simultaneously. The facial expression analysis unit also integrates video data from different camera angles to more accurately analyze participants' facial expressions and emotions. For example, it combines frontal and side video for analysis. The facial expression analysis unit also introduces a system that integrates video from multiple cameras to improve the accuracy of facial expression and emotion analysis. For example, it analyzes video from different angles in real time. By integrating video data from different camera angles, more accurate facial expression and emotion analysis becomes possible.
[0070] The facial expression analysis unit can provide participants with the results of the facial expression and emotional analysis as post-meeting feedback, encouraging self-improvement. For example, the facial expression analysis unit could build a system that provides participants with post-meeting feedback based on the results of the facial expression and emotional analysis. For example, it could provide a report of changes in emotions during the meeting. The facial expression analysis unit could also provide feedback based on the analysis results to help participants improve themselves. For example, it could offer advice such as, "There were many tense moments during the meeting. Try some relaxation techniques." The facial expression analysis unit could also provide feedback based on the results of the facial expression and emotional analysis, introducing a system that helps participants improve themselves. For example, it could display changes in emotions in a graph to clarify areas for improvement. By providing participants with analysis results as feedback, participants can improve themselves.
[0071] The pointer analysis unit can analyze the speed and frequency of pointer movement to identify the speaker's emphasis points. The pointer analysis unit, for example, analyzes the speed of pointer movement to build a system that identifies the points the speaker wants to emphasize. For example, it records points where the pointer moves quickly as emphasis points. The pointer analysis unit also analyzes the frequency of pointer movement to identify points that the speaker has pointed to multiple times. For example, if the same point is pointed to multiple times, it records that point as an important point. The pointer analysis unit also introduces a system that analyzes the speed and frequency of pointer movement to comprehensively evaluate the speaker's emphasis points. For example, it records points with high speed and frequency as emphasis points. In this way, the speaker's emphasis points can be identified by analyzing the speed and frequency of pointer movement.
[0072] The pointer analysis unit can synchronize pointer movement with remarks to automatically highlight important points. The pointer analysis unit, for example, synchronizes pointer movement with remarks to build a system that automatically highlights important points. For example, when a speaker points to a specific part while explaining, that part is highlighted. The pointer analysis unit also analyzes remarks and pointer movement in real time to automatically highlight important points. For example, it records the information as "Speaker C: Point to page 3, figure 2 in the materials." The pointer analysis unit also introduces a system that synchronizes pointer movement with remarks to automatically highlight important points. For example, when a specific part of a slide is pointed to, that part is highlighted. This makes it possible to automatically highlight important points by synchronizing pointer movement with remarks.
[0073] The pointer analysis unit can use the emotion estimation function to analyze the speaker's emotion associated with the pointer's movement and reflect it in the recording. The pointer analysis unit, for example, builds a system that analyzes the speaker's emotion associated with the pointer's movement. For example, if the pointer moves quickly, it estimates that the speaker is excited and records that information. The pointer analysis unit also uses the emotion estimation function to analyze the speaker's emotion associated with the pointer's movement in real time and reflect it in the minutes. For example, it records "Speaker C: Pointing to page 3, Figure 2 of the materials (slightly excited)." The pointer analysis unit also introduces a system that comprehensively analyzes the pointer's movement and the speaker's emotion and reflects it in the minutes. For example, it records that the speaker is excited if the pointer moves quickly and the tone of the voice is high. In this way, by analyzing the speaker's emotion associated with the pointer's movement and reflecting it in the recording, more detailed minutes can be generated.
[0074] The pointer analysis unit can reproduce pointer movement as a 3D model, making it easier to understand visually. The pointer analysis unit, for example, builds a system that reproduces pointer movement as a 3D model. For example, the movement of a speaker pointing to a specific part of a document is displayed as a 3D model. The pointer analysis unit also reproduces pointer movement as a 3D model, making it easier to understand visually. For example, the movement of a speaker pointing to a specific part of a slide is displayed as a 3D animation. The pointer analysis unit also introduces a system that uses 3D models to visually reproduce pointer movement. For example, the speaker's hand movement is displayed as a 3D model, making the pointing part clear. In this way, reproducing pointer movement as a 3D model makes it easier to understand visually.
[0075] The pointer analysis unit can integrate and analyze pointer movements of different presentation tools and provide common analysis results. The pointer analysis unit, for example, builds a system that integrates and analyzes pointer movements of different presentation tools. For example, it simultaneously analyzes pointer movements of PowerPoint and Keynote. The pointer analysis unit also integrates pointer movements of different presentation tools and provides common analysis results. For example, it analyzes pointer movements used in multiple tools as a single piece of data. The pointer analysis unit also introduces a system that integrates and analyzes pointer movements of presentation tools and provides common analysis results. For example, it displays pointer movements of different tools in a single view. This makes it possible to integrate and analyze pointer movements of different presentation tools and provide common analysis results.
[0076] The pointer analysis unit can feed back the results of the analysis of pointer movement to participants in real time, thereby promoting understanding. The pointer analysis unit, for example, builds a system that feeds back the results of the analysis of pointer movement to participants in real time. For example, the part indicated by the speaker is instantly displayed to the participants. The pointer analysis unit also feeds back the results of the analysis of pointer movement in real time, promoting understanding among participants. For example, when indicating a specific part of a slide, that part is highlighted. The pointer analysis unit also introduces a system that feeds back the results of the analysis of pointer movement in real time, allowing participants to instantly understand the speaker's intention. For example, the part indicated is highlighted. In this way, by feeding back the results of the analysis of pointer movement in real time, understanding among participants can be promoted.
[0077] The minutes generation unit can automatically evaluate the importance of the content of statements and highlight important statements when generating minutes. The minutes generation unit, for example, builds a system that automatically evaluates the importance of the content of statements and highlights important statements. For example, the importance is evaluated based on the frequency of keywords and phrases. The minutes generation unit also analyzes the importance of the content of statements and highlights important statements. For example, it records "Speaker A: Today's agenda item is XX (important)." The minutes generation unit also introduces a system that automatically evaluates the importance of the content of statements and reflects this in the minutes. For example, important statements are highlighted in bold or color. In this way, by automatically evaluating the importance of the content of statements and highlighting important statements, it becomes easier for readers of the minutes to grasp important information.
[0078] The minutes generation unit can use the emotion estimation function to assign emotion tags to the minutes and visualize the flow of emotions. The minutes generation unit, for example, uses the emotion estimation function to build a system that assigns emotion tags to the minutes. For example, emotions regarding the content of remarks are recorded as tags. The minutes generation unit also assigns emotion tags and visualizes the flow of emotions in the minutes. For example, it records "Speaker C: I'm looking forward to the new project (joyful)." The minutes generation unit also introduces a system that uses the emotion estimation function to assign emotion tags to the minutes and visualizes the flow of emotions in graphs and charts. For example, it indicates changes in emotions during a meeting with colors. In this way, by assigning emotion tags and visualizing the flow of emotions, readers of the minutes can more easily understand the flow of emotions in a meeting.
[0079] The minutes generation unit can provide minutes in a multimedia format synchronized with audio and video. The minutes generation unit, for example, builds a system that provides minutes in a multimedia format synchronized with audio and video. For example, it adds audio and video corresponding to the content of remarks as links. The minutes generation unit also provides minutes synchronized with audio and video so that participants can visually check the content of remarks. For example, it records "Speaker A: Today's agenda is ____ (audio link)." The minutes generation unit also introduces a system that provides minutes in a multimedia format so that participants can check the minutes while referring to the audio and video. For example, it embeds video corresponding to the content of remarks. By providing minutes in a multimedia format synchronized with audio and video, participants can visually check the content of remarks.
[0080] The minutes generation unit can automatically generate minutes in different formats (PDF, Word, HTML, etc.) to meet user needs. The minutes generation unit, for example, builds a system that automatically generates minutes in different formats. For example, minutes are generated in PDF, Word, and HTML formats. The minutes generation unit also provides minutes in different formats according to user needs. For example, minutes in PDF format are sent by email. The minutes generation unit also introduces a system that automatically generates minutes in different formats and allows users to select. For example, minutes in HTML format are published on a website. This makes it possible to meet user needs by automatically generating minutes in different formats.
[0081] The minutes generation unit can add a function to read out the generated minutes aloud through an AI assistant. For example, the minutes generation unit builds a system that reads out the generated minutes aloud through an AI assistant. For example, the generated minutes are read out using speech synthesis technology. The minutes generation unit also adds a function to read out the contents of the minutes aloud using an AI assistant. For example, it plays out loud "Speaker A: Today's agenda is ____." The minutes generation unit also introduces a function to read out the generated minutes aloud so that participants can listen to them. For example, it plays out the minutes aloud after the meeting ends. This allows participants to listen to the generated minutes by reading them out loud.
[0082] The minutes generation unit can automatically generate the agenda for the next meeting based on the generated minutes. The minutes generation unit, for example, builds a system that automatically generates the agenda for the next meeting based on the generated minutes. For example, it analyzes the contents of the minutes and automatically sets the next agenda. The minutes generation unit also automatically generates the agenda for the next meeting based on the contents of the minutes and shares it with the participants. For example, it records "Next agenda: Progress report on XX." The minutes generation unit also introduces a system that automatically generates the agenda for the next meeting based on the generated minutes. For example, it extracts important points from the minutes and sets them as the next agenda. In this way, meeting preparations are made more efficient by automatically generating the agenda for the next meeting based on the generated minutes.
[0083] The minutes generation unit can analyze the contents of the minutes and automatically extract unresolved issues and action items. For example, the minutes generation unit builds a system that analyzes the contents of the minutes and automatically extracts unresolved issues and action items. For example, it extracts items marked "unresolved" from the minutes. The minutes generation unit also automatically extracts unresolved issues and action items based on the contents of the minutes and notifies participants. For example, it records "Action item: Investigate XX." The minutes generation unit also introduces a system that analyzes the contents of the minutes and automatically extracts unresolved issues and action items. For example, it extracts important tasks from the minutes and makes a list. This makes meeting follow-up more efficient by analyzing the contents of the minutes and automatically extracting unresolved issues and action items.
[0084] The minutes generation unit can use the emotion estimation function to analyze participants' emotional reactions to the contents of the minutes and use the results as a reference for decision-making. For example, the minutes generation unit uses the emotion estimation function to build a system that analyzes participants' emotional reactions to the contents of the minutes. For example, it calculates an emotion score for each item in the minutes. The minutes generation unit also analyzes the contents of the minutes based on the participants' emotional reactions and uses the results as a reference for decision-making. For example, it records "Participant A: I'm looking forward to the new project (happy)." The minutes generation unit also introduces a system that uses the emotion estimation function to analyze participants' emotional reactions to the contents of the minutes in real time and uses the results as a reference for decision-making. For example, it prioritizes discussion of items with high emotion scores. In this way, analyzing participants' emotional reactions using the emotion estimation function can be used as a reference for decision-making.
[0085] The minutes generation unit can share minutes on the cloud and enable real-time collaborative editing. The minutes generation unit, for example, builds a system that shares minutes on the cloud and enables real-time collaborative editing. For example, it provides a collaborative editing function like Google Docs. The minutes generation unit also shares minutes on the cloud and enables participants to edit them in real time. For example, minutes can be updated during the meeting so that everyone can see the latest information. The minutes generation unit also introduces a cloud-based minutes sharing system that enables real-time collaborative editing. For example, participants can simultaneously add comments and corrections to the minutes. This allows minutes to be shared on the cloud and enables real-time collaborative editing so that participants can simultaneously add comments and corrections to the minutes.
[0086] The minutes generation department can work with different project management tools to automatically register the contents of the minutes as tasks. For example, the minutes generation department could work with different project management tools to build a system that automatically registers the contents of the minutes as tasks. For example, it could automatically register action items from the minutes in Trello or Asana. The minutes generation department could also automatically reflect the contents of the minutes in the project management tool and register them as tasks. For example, it could add "Action item: Investigate XX" to the project management tool. The minutes generation department could also work with project management tools to introduce a system that automatically registers the contents of the minutes as tasks. For example, it could automatically register important tasks from the minutes in JIRA. This would allow it to work with different project management tools and automatically register the contents of the minutes as tasks, making meeting follow-up more efficient.
[0087] The minutes generation unit can use the emotion estimation function to collect participants' emotional feedback on the contents of the minutes and reflect it in the next meeting. For example, the minutes generation unit uses the emotion estimation function to build a system that collects participants' emotional feedback on the contents of the minutes. For example, it collects an emotion score for each item in the minutes. The minutes generation unit also adjusts the content of the next meeting based on the participants' emotional feedback. For example, it records "Participant B: Dissatisfied with this topic (emotion score: low)" and changes the next agenda. The minutes generation unit also uses the emotion estimation function to introduce a system that collects participants' emotional feedback on the contents of the minutes and reflects it in the next meeting. For example, it improves items with low emotion scores. In this way, the quality of the meeting can be improved by using the emotion estimation function to collect participants' emotional feedback on the contents of the minutes and reflecting it in the next meeting.
[0088] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0089] The emotext system can further include a summarization unit that summarizes what participants have said. For example, the summarization unit can shorten long speeches and extract important points. For example, it can summarize what a speaker has said at length into a few lines. The summarization unit can also summarize what has been said in real time and provide a summary immediately during the meeting. For example, a summary can be displayed as soon as a speech is finished. The summarization unit can also reflect the summary of what has been said in the minutes, allowing readers to immediately grasp the important points. For example, it can record "Speaker A: Progress report on new project (summary)." In this way, summarization of what has been said improves the readability of the minutes.
[0090] The emotext system can further include a speech frequency analysis unit that analyzes the frequency of speech made by each speaker. The speech frequency analysis unit, for example, counts the number of times each participant speaks and reflects this in the minutes. For example, it may record "Speaker B: 10 speeches." The speech frequency analysis unit can also evaluate the balance of speech made during a meeting based on speech frequency. For example, if a particular participant speaks a lot, it will record that information. The speech frequency analysis unit can also visualize the results of the speech frequency analysis in graphs and charts, making it easier to understand the progress of the meeting. For example, it could display the number of speeches as a bar graph. This makes it easier to understand the progress of the meeting by analyzing speech frequency.
[0091] The emotext system can further include a classification unit that classifies the content of participants' comments. For example, the classification unit classifies the content of comments by theme and reflects this in the minutes. For example, it may record "Speaker C: About a new project (theme: project)." The classification unit can also classify the content of comments in real time and provide the classification results immediately during the meeting. For example, the theme may be displayed as soon as a comment is finished. The classification unit can also visualize the classification results of the content of comments in graphs or charts, making it easier to understand the content of the meeting. For example, the number of comments for each theme may be displayed in a pie chart. In this way, classifying the content of comments makes it easier to understand the content of the minutes.
[0092] The emotext system can further include a translation unit that translates the content of participants' statements. The translation unit, for example, automatically translates the content of statements made in different languages and reflects it in the minutes. For example, it may record "Speaker D: I'm looking forward to the new project (English)" and translate it into Japanese. The translation unit can also translate the content of statements in real time and provide the translation results immediately during the meeting. For example, the translation may be displayed as soon as the statement is finished. The translation unit also reflects the translation results in the minutes, making it easier for participants speaking different languages to understand the content of the meeting. For example, it may record "Speaker E: There are problems with this proposal (Japanese)" and translate it into English. By translating the content of statements, it makes it easier for participants speaking different languages to understand the content of the meeting.
[0093] The emotext system can further include an evaluation unit that evaluates the content of participants' comments. The evaluation unit, for example, evaluates the importance and usefulness of the content of the comments and reflects this in the minutes. For example, it may record "Speaker F: New project proposal (important)." The evaluation unit can also evaluate the content of comments in real time and provide the evaluation results immediately during the meeting. For example, the evaluation may be displayed as soon as the comment is finished. The evaluation unit may also visualize the evaluation results in graphs or charts, making the content of the meeting easier to understand. For example, it may highlight important comments. In this way, evaluating the content of the comments makes it easier to understand the content of the minutes.
[0094] The emotext system may further include an emotion display unit that estimates the emotions of participants and displays changes in emotion in real time. The emotion display unit, for example, estimates the emotions of participants and displays changes in emotion using colors or graphs. For example, happiness is displayed in yellow and anger in red. The emotion display unit also displays changes in emotion in real time, allowing participants to immediately grasp changes in emotion during a meeting. For example, changes in emotion are displayed as soon as a speech finishes. The emotion display unit also visualizes changes in emotion using graphs or charts, making it easier to grasp the flow of emotions in a meeting. For example, a graph showing changes in emotion along a time axis may be added. This makes it easier to grasp the flow of emotions in a meeting by displaying changes in emotion in real time.
[0095] The emotext system may further include an emotion reflection unit that estimates the emotions of participants and reflects changes in their emotions in the minutes. The emotion reflection unit, for example, estimates the emotions of participants and reflects changes in their emotions in the minutes. For example, it may record, "Speaker G: I'm looking forward to the new project (happy)." The emotion reflection unit may also reflect changes in emotions in the minutes in real time, recording changes in emotions instantly during the meeting. For example, changes in emotions may be recorded as soon as a speech is finished. The emotion reflection unit may also visualize changes in emotions in graphs or charts, making it easier to understand the flow of emotions in the minutes. For example, a graph showing changes in emotions along a time axis may be added. In this way, changes in emotions may be reflected in the minutes, making it easier to understand the flow of emotions in the minutes.
[0096] The emotext system can further include an emotion evaluation unit that estimates the emotions of participants and evaluates the content of their comments based on changes in their emotions. The emotion evaluation unit, for example, estimates the emotions of participants and evaluates the importance and usefulness of their comments based on changes in their emotions. For example, it may record "Speaker H: New project proposal (joy, important)." The emotion evaluation unit can also evaluate changes in emotions in real time and provide evaluation results immediately during the meeting. For example, the evaluation may be displayed as soon as a comment is finished. The emotion evaluation unit may also visualize the evaluation results in graphs or charts, making it easier to understand the content of the meeting. For example, it may highlight important comments. This allows the content of comments to be evaluated based on changes in emotions, making it easier to understand the content of the meeting minutes.
[0097] The emotext system may further include an emotion adjustment unit that estimates the emotions of participants and adjusts the progress of the meeting based on changes in their emotions. The emotion adjustment unit, for example, estimates the emotions of participants and adjusts the progress of the meeting based on changes in their emotions. For example, if a participant is nervous, it may suggest a break to help them relax. The emotion adjustment unit may also grasp changes in emotions in real time and instantly adjust the progress during the meeting. For example, changes in emotions may be displayed as soon as a participant finishes speaking, and the progress may be adjusted. The emotion adjustment unit may also visualize changes in emotions in graphs or charts, making it easier to understand the progress of the meeting. For example, a graph showing changes in emotions over time may be added. This may improve the quality of the meeting by adjusting the progress of the meeting based on changes in emotions.
[0098] The emotext system can further include an emotional agenda unit that estimates the emotions of participants and adjusts the agenda for the next meeting based on changes in their emotions. The emotional agenda unit, for example, estimates the emotions of participants and adjusts the agenda for the next meeting based on changes in their emotions. For example, an agenda item for which a participant expressed dissatisfaction is discussed again at the next meeting. The emotional agenda unit can also grasp changes in emotions in real time and immediately adjust the agenda for the next meeting. For example, changes in emotions are displayed as soon as a meeting ends, and the next agenda is adjusted. The emotional agenda unit can also visualize changes in emotions in graphs or charts, making it easier to understand the agenda for the next meeting. For example, a graph showing changes in emotions over time can be added. This allows the quality of meetings to be improved by adjusting the agenda for the next meeting based on changes in emotions.
[0099] The processing flow of the second embodiment will be briefly explained below.
[0100] Step 1: The audio data analysis unit analyzes audio data and transcribes it. For example, it analyzes recordings of meetings and lectures and converts what is being said into text. The audio data analysis unit analyzes audio data in MP3 or WAV format and transcribes what is being said. It can also analyze real-time audio streams and instantly convert what is being said into text. Step 2: The facial expression analyzer analyzes the video data to identify the participants' facial expressions and emotions. For example, it analyzes video data of a meeting or lecture to recognize the participants' facial expressions. The facial expression analyzer analyzes video data in MP4 or AVI format to identify the participants' facial expressions. It can also analyze real-time video streams to instantly identify the participants' facial expressions. Step 3: The pointer analysis unit analyzes the pointer position data to identify important points in the document. For example, it analyzes the movement of the speaker's pointer to identify important points in the document. The pointer analysis unit analyzes mouse coordinate data and laser pointer position data to identify important points. It can also analyze the frequency and speed of pointer movement to identify emphasized points. Step 4: The minutes generator combines the results of the voice data analysis, facial expression analysis, and pointer analysis to generate detailed minutes. For example, it combines the results of voice data transcription, facial expression analysis, and pointer analysis to generate detailed minutes that include the content of remarks, participants' facial expressions, and key points of the materials. The generated minutes also support information sharing and decision-making among participants.
[0101] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0102] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0103] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0104] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0105] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0106] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0107] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0108] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0109] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0110] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0111] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0112] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0113] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0114] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0115] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0116] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0117] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0118] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0119] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0120] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0121] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0122] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0123] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0124] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0125] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0126] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0127] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0128] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0129] In the headset type terminal 314, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0130] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0131] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0132] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0133] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0134] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0135] 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0136] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0137] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0138] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0140] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0141] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0142] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0143] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0144] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0145] In the robot 414, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The robot 414 also has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0146] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0147] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0148] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0149] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0150] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0151] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[0152] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[0153] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[0154] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.
[0155] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[0156] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[0157] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.
[0158] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[0159] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0160] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[0161] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[0162] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[0163] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[0164] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[0165] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.
[0166] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0167] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]
[0168] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot
Claims
1. an audio data analysis unit that analyzes audio data and transcribes it; An expression analysis unit that analyzes video data to identify participants' expressions and emotions; a pointer analysis unit that analyzes pointer position data to identify important points on the document; a minutes generation unit that generates detailed minutes by integrating the analysis results of the voice data analysis unit, the facial expression analysis unit, and the pointer analysis unit. A system characterized by:
2. The voice data analysis unit Automatically removes background sounds and noise from the audio data to make the spoken content clearer 2. The system of claim 1.
3. The voice data analysis unit Analyze the tone and speed of the speaker's voice to estimate their emotions and tension and reflect them in the recording.
2. The system of claim 1.
4. The voice data analysis unit Analyze the speaker's emotions and assign emotion tags to the content of the speech 2. The system of claim 1.
5. The voice data analysis unit The audio data is transcribed in real time, and the minutes are generated instantly during the meeting.
2. The system of claim 1.
6. The voice data analysis unit The speech data in different languages is analyzed simultaneously to generate the minutes in multiple languages.
2. The system of claim 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A