System

The system addresses the challenge of evaluating meeting progress and individual contributions by using real-time transcription and analysis to enhance meeting management through detailed evaluations and visual tracking.

JP2026029718APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024132572
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Conventional techniques face challenges in efficiently evaluating the progress and content of individual remarks during meetings.

Method used

A system incorporating a transcription unit, progress evaluation unit, and utterance evaluation unit to transcribe, analyze, and evaluate meeting data in real time, providing detailed assessments of meeting progress and individual contributions.

Benefits of technology

The system efficiently evaluates meeting progress and individual remarks, improving meeting management by offering real-time transcription, noise removal, emotional analysis, multilingual support, and visual progress tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026029718000001_ABST
    Figure 2026029718000001_ABST
Patent Text Reader

Abstract

An object of the system according to the embodiment is to efficiently evaluate the progress status of the conference and the statement content of each individual.SOLUTION: A system includes a transcription unit, a progress evaluation unit, and a statement evaluation unit. The transcription unit transcribes the audio data of the conference. The progress evaluation unit evaluates the progress status of the entire conference based on the transcription data generated by the transcription unit. A speech evaluation part evaluates the speech contents of each individual on the basis of the transcription data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques have had the problem of making it difficult to efficiently evaluate the progress of a meeting and the content of each individual's comments.

[0005] The system according to the embodiment aims to efficiently evaluate the progress of a meeting and the content of each individual's remarks. [Means for solving the problem]

[0006] The system according to the embodiment includes a transcription unit, a progress evaluation unit, and a utterance evaluation unit. The transcription unit transcribes audio data of a conference. The progress evaluation unit evaluates the progress of the entire conference based on the transcription data generated by the transcription unit. The utterance evaluation unit evaluates the content of each individual's utterance based on the transcription data. [Effects of the Invention]

[0007] The system according to the embodiment can efficiently evaluate the progress of a meeting and the content of each individual's remarks. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION

[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0010] First, the terms used in the following description will be explained.

[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).

[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.

[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.

[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example 1) A system for improving meeting management according to an embodiment of the present invention is a system that transcribes the contents of a meeting, evaluates the progress of the entire meeting and the attitude of each individual based on the transcription data, and aims to improve the management of the meeting. As a result, the system for improving meeting management can evaluate in detail the progress of the meeting and the content of each individual's remarks, and aim to improve the management of the meeting.

[0029] A system for improving meeting management according to an embodiment includes a transcription unit, a progress evaluation unit, and a comment evaluation unit. The transcription unit transcribes audio data from a meeting. For example, the recording of the meeting is input into a generation AI, which then transcribes the data using speech recognition technology. The transcription unit can also transcribe the audio data from a meeting in real time. For example, as each comment is made during a meeting, the content of that comment is simultaneously displayed as text data. The progress evaluation unit evaluates the overall progress of the meeting based on the transcription data generated by the transcription unit. For example, the generation AI analyzes the transcription data and calculates the ratio of the number of speakers to the number of participants. The system also analyzes the duration of the meeting and evaluates the effectiveness of time utilization. The system also evaluates whether the main points are clearly communicated at each stage: explanation, discussion, and summary. The comment evaluation unit evaluates the content of each individual's comment based on the transcription data. For example, the generation AI analyzes the transcription data and measures the length of each comment. The system also detects and points out redundant expressions, emotional comments, and notable catchphrases. As a result, the system for improving conference management according to the embodiment can evaluate in detail the progress of the conference and the content of each individual's remarks, thereby improving the management of the conference.

[0030] The transcription unit can transcribe audio data in real time and provide text data instantly during a meeting. For example, the transcription unit inputs meeting audio data into a generation AI in real time, which then transcribes it instantly. For example, as soon as a comment is made during a meeting, the content of that comment is displayed as text data. The generation AI also analyzes the audio data in real time and generates text data for each speaker. For example, the content of comments is converted into text sequentially as the meeting progresses. A system can also be built that transcribes meeting audio data in real time, allowing participants to instantly check the text data during the meeting. For example, the content of comments is displayed on a screen as the meeting progresses. This allows text data to be provided instantly during the meeting, allowing participants to check the content in real time.

[0031] The transcription unit can automatically remove background noise from audio data to generate more accurate transcription data. For example, the generation AI analyzes audio data and uses an algorithm to automatically remove background noise. For example, it filters out noise in the conference room and external sounds to accurately transcribe only what is being said. The generation AI also uses noise removal technology from audio data to more accurately convert what is being said into text. For example, it removes air conditioner noise and keyboard typing sounds to extract only the speaker's voice. In addition, a system is built in which the generation AI removes background noise from audio data in real time to improve transcription accuracy. For example, it automatically filters out noise generated during a meeting to accurately transcribe what is being said. This improves transcription accuracy by removing background noise.

[0032] The transcription unit can analyze audio and video data and reflect the speaker's facial expressions and gestures in the text data. In the transcription unit, for example, the generation AI analyzes video data of a meeting and reflects the speaker's facial expressions and gestures in the text data. For example, in addition to the content of what is said, the speaker's facial expressions and movements are converted into text. In addition, video data analysis technology is used to integrate the speaker's non-verbal information into the text data. For example, the speaker's smile and hand gestures are recorded in the text data. In addition, a system is built in which the generation AI simultaneously analyzes audio and video data and reflects the speaker's facial expressions and gestures in the text data. For example, the speaker's gaze and posture are converted into text along with the content of what is said. In this way, the content of the meeting can be understood in more detail by reflecting the speaker's non-verbal information in the text data.

[0033] The transcription unit can simultaneously transcribe meetings in different languages ​​and generate multilingual text data in real time. In the transcription unit, for example, the generation AI simultaneously analyzes audio data in different languages ​​and generates multilingual text data in real time. For example, meetings in English and Japanese can be simultaneously transcribed and text data in both languages ​​can be provided. In addition, using multilingual speech recognition technology, the generation AI transcribes meetings in different languages ​​in real time. For example, text data is generated in multiple languages ​​as comments are made during a meeting. In addition, a system is constructed in which the generation AI analyzes audio data in different languages ​​and generates multilingual text data in real time. For example, comments are converted into text in multiple languages ​​as the meeting progresses. This makes it possible to simultaneously transcribe meetings in different languages ​​and operate multilingual meetings.

[0034] The progress evaluation unit can monitor the progress of the meeting in real time and issue an alert if the progress is behind schedule. For example, the progress evaluation unit will build a system in which the generation AI monitors the progress of the meeting in real time and issues an alert if the progress is behind schedule. For example, an alert will be displayed if the scheduled time is exceeded. The progress evaluation unit will also analyze the progress of the meeting and automatically issue an alert if the progress is behind schedule. For example, it will monitor the progress time for each agenda item and notify if a delay occurs. In addition, a system will be developed in which the generation AI monitors the progress of the meeting in real time and issues an alert if the progress is behind schedule. For example, it will monitor the progress as the meeting progresses and issue an alert if a delay occurs. This will prevent delays by issuing an alert if the meeting is behind schedule.

[0035] The progress evaluation unit can automatically record the progress of a meeting and generate a timeline that can be played back later. For example, the progress evaluation unit will build a system in which a generation AI automatically records the progress of a meeting and generates a timeline that can be played back later. For example, each statement and agenda item in the meeting will be displayed on the timeline. The progress of the meeting will also be automatically recorded and a timeline that can be played back later. For example, the start and end times of the meeting and the progress of each agenda item will be displayed on the timeline. A system will also be developed in which a generation AI will record the progress of a meeting in real time and generate a timeline that can be played back later. For example, each statement and agenda item will be displayed on the timeline as the meeting progresses. This will make it easier to review the content of the meeting by recording the progress of the meeting and generating a timeline that can be played back later.

[0036] The progress evaluation unit can compare the progress evaluation of a meeting with meetings in different industries or fields and provide a benchmark. For example, the progress evaluation unit builds a system in which a generation AI compares the progress evaluation of a meeting with meetings in different industries or fields and provides a benchmark. For example, it compares and evaluates the progress of meetings at companies in the same industry. In addition, data on meetings from different industries or fields is collected, and the generation AI evaluates the progress based on that data. For example, it compares and provides a benchmark by comparing the progress of meetings in different industries. In addition, a system is developed in which a generation AI compares the progress evaluation of a meeting with meetings in different industries or fields and provides a benchmark. For example, it compares and evaluates the progress of meetings in different industries and suggests areas for improvement. This makes it possible to objectively evaluate the progress of a meeting by comparing it with meetings in different industries or fields.

[0037] The progress evaluation unit can visualize the progress of the meeting and display it visually in graphs and charts. For example, the progress evaluation unit will build a system in which the generation AI visualizes the progress of the meeting and displays it visually in graphs and charts. For example, the progress of the meeting will be displayed in a pie chart or bar graph. The generation AI will also analyze the progress of the meeting and visualize it based on that data. For example, the progress of each agenda item will be displayed in a timeline or chart. A system will also be developed in which the generation AI will visualize the progress of the meeting in real time and display it visually in graphs and charts. For example, the progress will be displayed in graphs and charts as the meeting progresses. This will allow the progress of the meeting to be visually displayed, making it possible to grasp the progress at a glance.

[0038] The utterance evaluation unit can develop an algorithm that analyzes the content of each utterance and evaluates its quality and usefulness. For example, the utterance evaluation unit develops an algorithm in which the generation AI analyzes the content of each utterance and evaluates its quality and usefulness. For example, it evaluates based on the specificity and logic of the utterance. The utterance content is also analyzed, and the generation AI evaluates the quality and usefulness of the utterance based on that data. For example, it evaluates based on the originality and novelty of the utterance. The generation AI also analyzes the content of each utterance in real time and develops an algorithm that evaluates its quality and usefulness. For example, it evaluates based on the influence and feasibility of the utterance. In this way, the efficiency of meetings is improved by evaluating the quality and usefulness of utterances.

[0039] The speech evaluation unit can automatically record the frequency and length of each statement and evaluate the balance of statements. For example, the speech evaluation unit will build a system in which the generation AI automatically records the frequency and length of each statement and evaluates the balance of statements. For example, it will tally up the number of times each participant spoke and the length of time they spoke. It will also analyze the frequency and length of statements and have the generation AI evaluate the balance of statements based on that data. For example, it will detect bias and imbalance in statements. It will also develop a system in which the generation AI records the frequency and length of each statement in real time and evaluates the balance of statements. For example, it will identify participants who speak a lot and participants who speak a little. This will improve the fairness of meetings by recording the frequency and length of statements and evaluating the balance.

[0040] The speech evaluation unit can compare each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, the speech evaluation unit constructs a system in which a generation AI compares each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, it compares and evaluates with past meeting data. In addition, speech data from different meetings and projects is collected, and the generation AI evaluates speech based on that data. For example, it compares and evaluates the content of speech made by the same participant. In addition, a system can be developed in which the generation AI compares each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, it compares and evaluates the content of speech made in different projects and suggests areas for improvement. In this way, individual performance can be improved by comparing speech made in different meetings and projects.

[0041] The speech evaluation unit can summarize the content of each speech and extract the important points to evaluate it. For example, the speech evaluation unit will build a system in which a generation AI summarizes the content of each speech and extracts the important points to evaluate it. For example, it will extract the main points and keywords of the speech. The content of the speech will also be analyzed, and the generation AI will summarize based on that data and evaluate the important points. For example, it will extract the core parts and conclusions of the speech. In addition, a system will be developed in which the generation AI will summarize the content of each speech in real time, extract the important points, and evaluate them. For example, it will automatically summarize the main points of the speech and reflect them in the evaluation. In this way, the efficiency of meetings will be improved by extracting and evaluating the main points of speech.

[0042] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0043] The system for improving meeting management can further include a concentration measurement unit that measures the concentration level of participants. The concentration measurement unit, for example, uses a generation AI to analyze the gaze and posture of participants during a meeting to evaluate their concentration level. For example, it measures the concentration level based on the amount of time participants spend looking at the screen and changes in their posture. The concentration measurement unit can also analyze the frequency and content of participants' comments to evaluate their concentration level. For example, it measures the concentration level based on the frequency of comments and the relevance of the content. The concentration measurement unit can also analyze the biometric data of participants during a meeting to evaluate their concentration level. For example, it measures the concentration level based on changes in heart rate and breathing. In this way, the efficiency of meetings can be improved by evaluating the concentration level of participants.

[0044] The system for improving conference management can further include an opinion collection unit that anonymously collects participants' opinions. The opinion collection unit, for example, provides an interface that allows participants to anonymously post opinions during the conference. For example, participants use a smartphone or tablet to input their opinions anonymously, and the system collects them. The opinion collection unit can also include a function that allows participants to provide anonymous feedback after the conference. For example, opinions are collected in the form of a questionnaire after the conference ends. The opinion collection unit can also analyze the collected opinions and suggest areas for improvement in the conference. For example, the progress and content of the conference can be evaluated based on the participants' opinions, and areas for improvement can be suggested. In this way, the quality of the conference can be improved by anonymously collecting participants' opinions.

[0045] The meeting management improvement system can further include an agenda progress visualization unit that visualizes the progress for each agenda item. The agenda progress visualization unit, for example, uses a generation AI to analyze the progress of each agenda item in real time and display it in a graph or chart. For example, it displays the progress time for each agenda item and the number of speakers in a pie chart or bar graph. The agenda progress visualization unit can also display the progress of each agenda item on a timeline. For example, it can display the start and end times of each agenda item on a timeline, allowing the progress to be understood at a glance. The agenda progress visualization unit can also analyze the progress of each agenda item and issue an alert if progress is behind schedule. For example, it can issue an alert if the scheduled time is exceeded. In this way, by visualizing the progress for each agenda item, the progress of the meeting can be managed efficiently.

[0046] The system for improving meeting management can further include a statement classification unit that automatically classifies the content of participants' statements. The statement classification unit, for example, uses a generation AI to analyze the content of each statement and classify it by category. For example, it may classify it into categories such as proposals, questions, and opinions. The statement classification unit can also analyze the content of statements and classify them by related agenda. For example, it may display each statement linked to the related agenda. The statement classification unit can also analyze the content of statements and classify them according to importance. For example, it may display important statements preferentially. This makes it easier to organize the content of the meeting by automatically classifying the content of statements.

[0047] The system for improving meeting management can further include a speech summarization unit that automatically summarizes what participants say. The speech summarization unit, for example, uses a generation AI to analyze the content of each speech and extract and summarize the key points. For example, it may concisely summarize the main points of the speech. The speech summarization unit can also analyze the content of speech and extract and summarize important information. For example, it may summarize the conclusions or proposals of the speech. The speech summarization unit can also summarize the content of speech in real time and display it during the meeting. For example, it may display a summary as soon as a speech is made. This allows the content of the meeting to be grasped efficiently by automatically summarizing the content of speech.

[0048] The system for improving meeting management can further include a comment evaluation unit that automatically evaluates the content of participants' comments. The comment evaluation unit, for example, uses a generation AI to analyze the content of each comment and evaluate it based on evaluation criteria. For example, it may evaluate based on the specificity and logic of the comment. The comment evaluation unit can also analyze the content of comments and evaluate the usefulness of the comment. For example, it may evaluate based on the originality and novelty of the comment. The comment evaluation unit can also evaluate the content of comments in real time and provide feedback during the meeting. For example, it may display the evaluation as soon as a comment is made. This makes it possible to improve the quality of meetings by automatically evaluating the content of comments.

[0049] The processing flow of the first embodiment will be briefly explained below.

[0050] Step 1: The transcription unit transcribes the audio data of the meeting. For example, the recording of the meeting is input into the generation AI, which then transcribes it using speech recognition technology. The transcription unit can also transcribe the audio data of the meeting in real time. For example, as someone speaks during the meeting, the content of that speech is displayed as text data at the same time. Step 2: The progress evaluation section evaluates the progress of the entire meeting based on the transcription data generated by the transcription section. For example, the generation AI analyzes the transcription data and calculates the ratio of the number of speakers to the number of participants. It also analyzes the duration of the meeting and evaluates how effectively the time was used. It also evaluates whether the main points were made clear at each stage: explanation, discussion, and summary. Step 3: The speech evaluation unit evaluates the content of each individual's speech based on the transcription data. For example, the generation AI analyzes the transcription data and measures the length of each speech. It also detects and points out redundant expressions, emotional statements, and notable catchphrases.

[0051] (Example 2) A system for improving meeting management according to an embodiment of the present invention is a system that transcribes the contents of a meeting, evaluates the progress of the entire meeting and the attitude of each individual based on the transcription data, and aims to improve the management of the meeting. As a result, the system for improving meeting management can evaluate in detail the progress of the meeting and the content of each individual's remarks, and aim to improve the management of the meeting.

[0052] A system for improving meeting management according to an embodiment includes a transcription unit, a progress evaluation unit, and a comment evaluation unit. The transcription unit transcribes audio data from a meeting. For example, the recording of the meeting is input into a generation AI, which then transcribes the data using speech recognition technology. The transcription unit can also transcribe the audio data from a meeting in real time. For example, as each comment is made during a meeting, the content of that comment is simultaneously displayed as text data. The progress evaluation unit evaluates the overall progress of the meeting based on the transcription data generated by the transcription unit. For example, the generation AI analyzes the transcription data and calculates the ratio of the number of speakers to the number of participants. The system also analyzes the duration of the meeting and evaluates the effectiveness of time utilization. The system also evaluates whether the main points are clearly communicated at each stage: explanation, discussion, and summary. The comment evaluation unit evaluates the content of each individual's comment based on the transcription data. For example, the generation AI analyzes the transcription data and measures the length of each comment. The system also detects and points out redundant expressions, emotional comments, and notable catchphrases. As a result, the system for improving conference management according to the embodiment can evaluate in detail the progress of the conference and the content of each individual's remarks, thereby improving the management of the conference.

[0053] The transcription unit can transcribe audio data in real time and provide text data instantly during a meeting. For example, the transcription unit inputs meeting audio data into a generation AI in real time, which then transcribes it instantly. For example, as soon as a comment is made during a meeting, the content of that comment is displayed as text data. The generation AI also analyzes the audio data in real time and generates text data for each speaker. For example, the content of comments is converted into text sequentially as the meeting progresses. A system can also be built that transcribes meeting audio data in real time, allowing participants to instantly check the text data during the meeting. For example, the content of comments is displayed on a screen as the meeting progresses. This allows text data to be provided instantly during the meeting, allowing participants to check the content in real time.

[0054] The transcription unit can automatically remove background noise from audio data to generate more accurate transcription data. For example, the generation AI analyzes audio data and uses an algorithm to automatically remove background noise. For example, it filters out noise in the conference room and external sounds to accurately transcribe only what is being said. The generation AI also uses noise removal technology from audio data to more accurately convert what is being said into text. For example, it removes air conditioner noise and keyboard typing sounds to extract only the speaker's voice. In addition, a system is built in which the generation AI removes background noise from audio data in real time to improve transcription accuracy. For example, it automatically filters out noise generated during a meeting to accurately transcribe what is being said. This improves transcription accuracy by removing background noise.

[0055] The transcription unit can use an emotion estimation function to estimate the speaker's emotional state and assign emotion tags to the transcript. For example, the transcription unit uses an algorithm in which the generation AI analyzes audio data and estimates the speaker's emotional state. For example, it estimates emotions based on the tone and speed of speech and assigns emotion tags to the transcript. It also uses the emotion estimation function to analyze the speaker's emotional state in real time and reflects this in the transcript. For example, it tags emotions such as joy, anger, and sadness. It also builds a system in which the generation AI estimates the speaker's emotional state and assigns emotion tags to the transcript. For example, it automatically adds emotion tags based on the content of speech and visualizes changes in emotion. This makes it easier to understand the atmosphere of a meeting by visualizing the speaker's emotional state.

[0056] The transcription unit can analyze audio and video data and reflect the speaker's facial expressions and gestures in the text data. In the transcription unit, for example, the generation AI analyzes video data of a meeting and reflects the speaker's facial expressions and gestures in the text data. For example, in addition to the content of what is said, the speaker's facial expressions and movements are converted into text. In addition, video data analysis technology is used to integrate the speaker's non-verbal information into the text data. For example, the speaker's smile and hand gestures are recorded in the text data. In addition, a system is built in which the generation AI simultaneously analyzes audio and video data and reflects the speaker's facial expressions and gestures in the text data. For example, the speaker's gaze and posture are converted into text along with the content of what is said. In this way, the content of the meeting can be understood in more detail by reflecting the speaker's non-verbal information in the text data.

[0057] The transcription unit can simultaneously transcribe meetings in different languages ​​and generate multilingual text data in real time. In the transcription unit, for example, the generation AI simultaneously analyzes audio data in different languages ​​and generates multilingual text data in real time. For example, meetings in English and Japanese can be simultaneously transcribed and text data in both languages ​​can be provided. In addition, using multilingual speech recognition technology, the generation AI transcribes meetings in different languages ​​in real time. For example, text data is generated in multiple languages ​​as comments are made during a meeting. In addition, a system is constructed in which the generation AI analyzes audio data in different languages ​​and generates multilingual text data in real time. For example, comments are converted into text in multiple languages ​​as the meeting progresses. This makes it possible to simultaneously transcribe meetings in different languages ​​and operate multilingual meetings.

[0058] The transcription unit uses an emotion estimation function to analyze the overall flow of emotions during a meeting and reflect the atmosphere of the meeting in the text data. In the transcription unit, for example, the generation AI analyzes audio data during a meeting and estimates the overall flow of emotions. For example, changes in emotions as the meeting progresses are reflected in the text data. The emotion estimation function also analyzes the emotional state of all participants during the meeting and reflects this in the text data. For example, the atmosphere of the meeting and rising emotions are converted into text. In addition, a system is built in which the generation AI analyzes the flow of emotions during a meeting in real time and reflects this in the text data. For example, changes in emotions are recorded in the text data as the meeting progresses. In this way, by reflecting the atmosphere of the meeting in the text data, the progress of the meeting can be understood in more detail.

[0059] The progress evaluation unit can monitor the progress of the meeting in real time and issue an alert if the progress is behind schedule. For example, the progress evaluation unit will build a system in which the generation AI monitors the progress of the meeting in real time and issues an alert if the progress is behind schedule. For example, an alert will be displayed if the scheduled time is exceeded. The progress evaluation unit will also analyze the progress of the meeting and automatically issue an alert if the progress is behind schedule. For example, it will monitor the progress time for each agenda item and notify if a delay occurs. In addition, a system will be developed in which the generation AI monitors the progress of the meeting in real time and issues an alert if the progress is behind schedule. For example, it will monitor the progress as the meeting progresses and issue an alert if a delay occurs. This will prevent delays by issuing an alert if the meeting is behind schedule.

[0060] The progress evaluation unit can automatically record the progress of a meeting and generate a timeline that can be played back later. For example, the progress evaluation unit will build a system in which a generation AI automatically records the progress of a meeting and generates a timeline that can be played back later. For example, each statement and agenda item in the meeting will be displayed on the timeline. The progress of the meeting will also be automatically recorded and a timeline that can be played back later. For example, the start and end times of the meeting and the progress of each agenda item will be displayed on the timeline. A system will also be developed in which a generation AI will record the progress of a meeting in real time and generate a timeline that can be played back later. For example, each statement and agenda item will be displayed on the timeline as the meeting progresses. This will make it easier to review the content of the meeting by recording the progress of the meeting and generating a timeline that can be played back later.

[0061] The progress evaluation unit can use the emotion estimation function to analyze changes in the emotions of speakers and participants as the meeting progresses and reflect them in the progress evaluation. For example, the progress evaluation unit will build a system in which a generation AI analyzes changes in the emotions of speakers and participants as the meeting progresses and reflects them in the progress evaluation. For example, it will reflect increases and decreases in emotion in the progress evaluation. In addition, the emotion estimation function will be used to analyze changes in the emotions of speakers and participants in real time as the meeting progresses and reflect them in the progress evaluation. For example, it will evaluate the progress based on changes in emotion. In addition, a system will be developed in which a generation AI analyzes changes in the emotions of speakers and participants as the meeting progresses and reflects them in the progress evaluation. For example, it will evaluate the progress based on changes in emotion and suggest areas for improvement. In this way, by analyzing changes in emotion as the meeting progresses and reflecting them in the progress evaluation, it will be possible to grasp the atmosphere and progress of the meeting in more detail.

[0062] The progress evaluation unit can compare the progress evaluation of a meeting with meetings in different industries or fields and provide a benchmark. For example, the progress evaluation unit builds a system in which a generation AI compares the progress evaluation of a meeting with meetings in different industries or fields and provides a benchmark. For example, it compares and evaluates the progress of meetings at companies in the same industry. In addition, data on meetings from different industries or fields is collected, and the generation AI evaluates the progress based on that data. For example, it compares and provides a benchmark by comparing the progress of meetings in different industries. In addition, a system is developed in which a generation AI compares the progress evaluation of a meeting with meetings in different industries or fields and provides a benchmark. For example, it compares and evaluates the progress of meetings in different industries and suggests areas for improvement. This makes it possible to objectively evaluate the progress of a meeting by comparing it with meetings in different industries or fields.

[0063] The progress evaluation unit can visualize the progress of the meeting and display it visually in graphs and charts. For example, the progress evaluation unit will build a system in which the generation AI visualizes the progress of the meeting and displays it visually in graphs and charts. For example, the progress of the meeting will be displayed in a pie chart or bar graph. The generation AI will also analyze the progress of the meeting and visualize it based on that data. For example, the progress of each agenda item will be displayed in a timeline or chart. A system will also be developed in which the generation AI will visualize the progress of the meeting in real time and display it visually in graphs and charts. For example, the progress will be displayed in graphs and charts as the meeting progresses. This will allow the progress of the meeting to be visually displayed, making it possible to grasp the progress at a glance.

[0064] The progress evaluation unit can use the emotion estimation function to monitor the emotional reactions of participants in real time as the meeting progresses and reflect them in the progress evaluation. The progress evaluation unit, for example, builds a system in which a generation AI monitors the emotional reactions of participants in real time as the meeting progresses and reflects them in the progress evaluation. For example, it reflects increases and decreases in emotion in the progress evaluation. In addition, it uses the emotion estimation function to analyze the emotional reactions of participants in real time as the meeting progresses and reflects them in the progress evaluation. For example, it evaluates the progress based on changes in emotion. In addition, it develops a system in which a generation AI analyzes the emotional reactions of participants as the meeting progresses and reflects them in the progress evaluation. For example, it evaluates the progress based on changes in emotion and suggests areas for improvement. In this way, by monitoring the emotional reactions of participants in real time, it is possible to grasp the progress of the meeting in more detail.

[0065] The utterance evaluation unit can develop an algorithm that analyzes the content of each utterance and evaluates its quality and usefulness. For example, the utterance evaluation unit develops an algorithm in which the generation AI analyzes the content of each utterance and evaluates its quality and usefulness. For example, it evaluates based on the specificity and logic of the utterance. The utterance content is also analyzed, and the generation AI evaluates the quality and usefulness of the utterance based on that data. For example, it evaluates based on the originality and novelty of the utterance. The generation AI also analyzes the content of each utterance in real time and develops an algorithm that evaluates its quality and usefulness. For example, it evaluates based on the influence and feasibility of the utterance. In this way, the efficiency of meetings is improved by evaluating the quality and usefulness of utterances.

[0066] The speech evaluation unit can automatically record the frequency and length of each statement and evaluate the balance of statements. For example, the speech evaluation unit will build a system in which the generation AI automatically records the frequency and length of each statement and evaluates the balance of statements. For example, it will tally up the number of times each participant spoke and the length of time they spoke. It will also analyze the frequency and length of statements and have the generation AI evaluate the balance of statements based on that data. For example, it will detect bias and imbalance in statements. It will also develop a system in which the generation AI records the frequency and length of each statement in real time and evaluates the balance of statements. For example, it will identify participants who speak a lot and participants who speak a little. This will improve the fairness of meetings by recording the frequency and length of statements and evaluating the balance.

[0067] The statement evaluation unit can use the emotion estimation function to analyze the emotional tone of each statement and reflect it in the statement evaluation. The statement evaluation unit, for example, builds a system in which a generation AI analyzes the emotional tone of each statement and reflects it in the statement evaluation. For example, it evaluates the emotional strength and tone of the statement. In addition, it uses the emotion estimation function to analyze the emotional tone of each statement in real time and reflects it in the statement evaluation. For example, it evaluates the positivity or negativity of the statement. In addition, a system is developed in which a generation AI analyzes the emotional tone of each statement and reflects it in the statement evaluation. For example, it evaluates the emotional impact and influence of the statement. In this way, the influence of a statement can be evaluated by analyzing the emotional tone of a statement and reflecting it in the statement evaluation.

[0068] The speech evaluation unit can compare each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, the speech evaluation unit constructs a system in which a generation AI compares each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, it compares and evaluates with past meeting data. In addition, speech data from different meetings and projects is collected, and the generation AI evaluates speech based on that data. For example, it compares and evaluates the content of speech made by the same participant. In addition, a system can be developed in which the generation AI compares each individual's speech evaluation with speech made in different meetings or projects to improve performance. For example, it compares and evaluates the content of speech made in different projects and suggests areas for improvement. In this way, individual performance can be improved by comparing speech made in different meetings and projects.

[0069] The speech evaluation unit can summarize the content of each speech and extract the important points to evaluate it. For example, the speech evaluation unit will build a system in which a generation AI summarizes the content of each speech and extracts the important points to evaluate it. For example, it will extract the main points and keywords of the speech. The content of the speech will also be analyzed, and the generation AI will summarize based on that data and evaluate the important points. For example, it will extract the core parts and conclusions of the speech. In addition, a system will be developed in which the generation AI will summarize the content of each speech in real time, extract the important points, and evaluate them. For example, it will automatically summarize the main points of the speech and reflect them in the evaluation. In this way, the efficiency of meetings will be improved by extracting and evaluating the main points of speech.

[0070] The statement evaluation unit can use the emotion estimation function to analyze the emotional impact of each statement and reflect it in the statement evaluation. The statement evaluation unit, for example, builds a system in which a generation AI analyzes the emotional impact of each statement and reflects it in the statement evaluation. For example, it evaluates the emotional strength and influence of the statement. In addition, it uses the emotion estimation function to analyze the emotional impact of each statement in real time and reflects it in the statement evaluation. For example, it evaluates the emotional tone and nuance of the statement. In addition, it develops a system in which a generation AI analyzes the emotional impact of each statement and reflects it in the statement evaluation. For example, it evaluates the emotional influence and empathy of the statement. In this way, the influence of a statement can be evaluated by analyzing the emotional impact of a statement and reflecting it in the statement evaluation.

[0071] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.

[0072] The system for improving meeting management can further include a concentration measurement unit that measures the concentration level of participants. The concentration measurement unit, for example, uses a generation AI to analyze the gaze and posture of participants during a meeting to evaluate their concentration level. For example, it measures the concentration level based on the amount of time participants spend looking at the screen and changes in their posture. The concentration measurement unit can also analyze the frequency and content of participants' comments to evaluate their concentration level. For example, it measures the concentration level based on the frequency of comments and the relevance of the content. The concentration measurement unit can also analyze the biometric data of participants during a meeting to evaluate their concentration level. For example, it measures the concentration level based on changes in heart rate and breathing. In this way, the efficiency of meetings can be improved by evaluating the concentration level of participants.

[0073] The system for improving conference management can further include an opinion collection unit that anonymously collects participants' opinions. The opinion collection unit, for example, provides an interface that allows participants to anonymously post opinions during the conference. For example, participants use a smartphone or tablet to input their opinions anonymously, and the system collects them. The opinion collection unit can also include a function that allows participants to provide anonymous feedback after the conference. For example, opinions are collected in the form of a questionnaire after the conference ends. The opinion collection unit can also analyze the collected opinions and suggest areas for improvement in the conference. For example, the progress and content of the conference can be evaluated based on the participants' opinions, and areas for improvement can be suggested. In this way, the quality of the conference can be improved by anonymously collecting participants' opinions.

[0074] The meeting management improvement system can further include an agenda progress visualization unit that visualizes the progress for each agenda item. The agenda progress visualization unit, for example, uses a generation AI to analyze the progress of each agenda item in real time and display it in a graph or chart. For example, it displays the progress time for each agenda item and the number of speakers in a pie chart or bar graph. The agenda progress visualization unit can also display the progress of each agenda item on a timeline. For example, it can display the start and end times of each agenda item on a timeline, allowing the progress to be understood at a glance. The agenda progress visualization unit can also analyze the progress of each agenda item and issue an alert if progress is behind schedule. For example, it can issue an alert if the scheduled time is exceeded. In this way, by visualizing the progress for each agenda item, the progress of the meeting can be managed efficiently.

[0075] The meeting management improvement system can further include an emotion display unit that estimates the emotional state of participants and displays emotional changes in real time. The emotion display unit, for example, uses a generation AI to analyze participants' voice data and facial expression data to estimate their emotional state. For example, it estimates emotions based on changes in speech tone and facial expressions and displays them in real time. The emotion display unit can also display changes in emotions during the meeting in graphs or charts. For example, it can display increases and decreases in emotions in a line graph. The emotion display unit can also adjust the progress of the meeting based on changes in emotions. For example, it can suggest changing the agenda if emotions are rising. In this way, displaying participants' emotional states in real time makes it easier to grasp the atmosphere of the meeting.

[0076] The system for improving meeting management can further include a speech order adjustment unit that estimates the emotional state of participants and adjusts the order of speech based on their emotions. The speech order adjustment unit, for example, uses a generation AI to analyze the participants' voice data and facial expression data to estimate their emotional state. For example, it estimates emotions based on changes in the tone of speech and facial expressions and adjusts the order of speech. The speech order adjustment unit can also change the order of speech in real time based on the emotional state. For example, it can prioritize speech from participants with heightened emotions. The speech order adjustment unit can also optimize the order of speech based on changes in emotions. For example, it can postpone speech from participants with calmer emotions. In this way, adjusting the order of speech based on emotions can smooth the progress of the meeting.

[0077] The meeting management improvement system can further include a progress adjustment unit that estimates the emotional state of participants and adjusts the progress of the meeting based on their emotions. The progress adjustment unit, for example, uses a generative AI to analyze the participants' voice data and facial expression data to estimate their emotional state. For example, it estimates emotions based on changes in the tone of speech and facial expressions and adjusts the progress of the meeting. The progress adjustment unit can also change the progress of the meeting in real time based on the emotional state. For example, it can change the agenda if emotions are high. The progress adjustment unit can also optimize the progress of the meeting based on changes in emotions. For example, it can move on to the agenda if emotions are calm. In this way, adjusting the progress of the meeting based on emotions can maintain a smooth atmosphere in the meeting.

[0078] The system for improving meeting management can further include a summarization unit that estimates the emotional state of participants and summarizes the content of the meeting based on their emotions. The summarization unit, for example, uses a generation AI to analyze the participants' voice data and facial expression data to estimate their emotional state. For example, it estimates emotions based on the tone of their remarks and changes in their facial expressions and summarizes the content of the meeting. The summarization unit can also extract the main points of the meeting based on their emotional state. For example, it can focus on summarizing statements that express high emotions. The summarization unit can also summarize the content of the meeting based on changes in emotions. For example, it can concisely summarize statements that express calm emotions. In this way, summarizing the content of the meeting based on emotions makes it easier to grasp the important points.

[0079] The system for improving meeting management can further include a statement classification unit that automatically classifies the content of participants' statements. The statement classification unit, for example, uses a generation AI to analyze the content of each statement and classify it by category. For example, it may classify it into categories such as proposals, questions, and opinions. The statement classification unit can also analyze the content of statements and classify them by related agenda. For example, it may display each statement linked to the related agenda. The statement classification unit can also analyze the content of statements and classify them according to importance. For example, it may display important statements preferentially. This makes it easier to organize the content of the meeting by automatically classifying the content of statements.

[0080] The system for improving meeting management can further include a speech summarization unit that automatically summarizes what participants say. The speech summarization unit, for example, uses a generation AI to analyze the content of each speech and extract and summarize the key points. For example, it may concisely summarize the main points of the speech. The speech summarization unit can also analyze the content of speech and extract and summarize important information. For example, it may summarize the conclusions or proposals of the speech. The speech summarization unit can also summarize the content of speech in real time and display it during the meeting. For example, it may display a summary as soon as a speech is made. This allows the content of the meeting to be grasped efficiently by automatically summarizing the content of speech.

[0081] The system for improving meeting management can further include a comment evaluation unit that automatically evaluates the content of participants' comments. The comment evaluation unit, for example, uses a generation AI to analyze the content of each comment and evaluate it based on evaluation criteria. For example, it may evaluate based on the specificity and logic of the comment. The comment evaluation unit can also analyze the content of comments and evaluate the usefulness of the comment. For example, it may evaluate based on the originality and novelty of the comment. The comment evaluation unit can also evaluate the content of comments in real time and provide feedback during the meeting. For example, it may display the evaluation as soon as a comment is made. This makes it possible to improve the quality of meetings by automatically evaluating the content of comments.

[0082] The processing flow of the second embodiment will be briefly explained below.

[0083] Step 1: The transcription unit transcribes the audio data of the meeting. For example, the recording of the meeting is input into the generation AI, which then transcribes it using speech recognition technology. The transcription unit can also transcribe the audio data of the meeting in real time. For example, as someone speaks during the meeting, the content of that speech is displayed as text data at the same time. Step 2: The progress evaluation section evaluates the progress of the entire meeting based on the transcription data generated by the transcription section. For example, the generation AI analyzes the transcription data and calculates the ratio of the number of speakers to the number of participants. It also analyzes the duration of the meeting and evaluates how effectively the time was used. It also evaluates whether the main points were made clear at each stage: explanation, discussion, and summary. Step 3: The speech evaluation unit evaluates the content of each individual's speech based on the transcription data. For example, the generation AI analyzes the transcription data and measures the length of each speech. It also detects and points out redundant expressions, emotional statements, and notable catchphrases.

[0084] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0085] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0086] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0087] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0088] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0089] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0090] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0091] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0092] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0093] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0094] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0095] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0096] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0097] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0098] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0099] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0100] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0101] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0102] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0103] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0104] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0105] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0106] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0107] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0108] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0109] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0110] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0111] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0112] In the headset type terminal 314, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.

[0113] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0114] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0115] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0116] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0117] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0118] 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0119] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.

[0120] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0121] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.

[0122] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0123] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0124] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0125] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0126] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0127] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.

[0128] In the robot 414, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The robot 414 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.

[0129] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.

[0130] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0131] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI ​​other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI ​​may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.

[0132] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.

[0133] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0134] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0135] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0136] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[0137] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.

[0138] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[0139] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[0140] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.

[0141] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0142] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0143] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0144] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[0145] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0146] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0147] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[0148] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.

[0149] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[0150] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference. [Explanation of symbols]

[0151] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot

Claims

1. a transcription unit that transcribes the audio data of the meeting; a progress evaluation unit that evaluates the progress of the entire conference based on the transcription data generated by the transcription unit; a comment evaluation unit that evaluates the content of each individual's comment based on the transcription data. A system characterized by:

2. The transcription unit The audio data is transcribed in real time and text data is provided instantly during the meeting.

2. The system of claim 1.

3. The transcription unit Automatically remove background noise from the audio data to generate more accurate transcription data 2. The system of claim 1.

4. The transcription unit Estimate the speaker's emotional state and assign emotion tags to the transcription data.

2. The system of claim 1.

5. The transcription unit The voice data and video data are analyzed, and the speaker's facial expressions and gestures are reflected in the text data.

2. The system of claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A