system
The system addresses the inefficiencies in real-time textification, report generation, and visualization of meeting content by using AI to transcribe, generate, and visualize meeting data, enhancing work efficiency and decision-making speed.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
The conventional process of textifying meeting content in real time, generating reports, and visualizing the content requires labor and time, leading to inefficiencies in information dissemination and decision-making.
A system comprising an audio analysis unit, report generation unit, and visualization unit, utilizing speech recognition and generative AI to transcribe, generate reports, and visualize meeting content in real time, respectively.
Enables real-time transcription, report generation, and visualization of meeting content, improving work efficiency by automating information provision and facilitating rapid decision-making.
Smart Images

Figure 2026073090000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional technology, there is a problem that the process of textifying the content of a meeting in real time, generating a report, and further aggregating and visualizing the content requires labor and time.
[0005] The system according to the embodiment aims to textify the content of a meeting in real time, generate a report, and further aggregate and visualize the content.
Means for Solving the Problems
[0006] The system according to this embodiment comprises an audio analysis unit, a report generation unit, an information provision unit, and a visualization unit. The audio analysis unit transcribes meeting audio data into text in real time. The report generation unit generates a report of the meeting content based on the data transcribed by the audio analysis unit. The information provision unit automatically aggregates the contents of the report generated by the report generation unit and provides it to the administrator. The visualization unit visualizes the information provided by the information provision unit. [Effects of the Invention]
[0007] The system according to this embodiment can transcribe meeting content into text in real time, generate reports, and further aggregate and visualize that content. [Brief explanation of the drawing]
[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]
[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0010] First, let's explain the terminology used in the following explanation.
[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).
[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0014] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between a plurality of computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.
[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.
[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.
[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example of form 1) The system according to an embodiment of the present invention is a system in which a generating AI transcribes discussions during pre-construction toolbox meetings into text in real time and generates a report of the meeting content. This system uses a generating AI to transcribe meeting audio data into text in real time and generates a report of the meeting content. For example, the generating AI uses speech recognition technology to convert the content of the discussion into text data. Next, the generating AI generates a report of the meeting content based on the transcribed data. This report includes the main points of the discussion, decisions made, action items, etc. Furthermore, the generating AI automatically aggregates the contents of the report and provides it to the administrator. The administrator can understand the progress of the meeting and any problems based on the aggregated data. This system automates the reporting of meeting content and the provision of information to the administrator, improving work efficiency. In addition, real-time transcription and report generation allow for rapid sharing of content, thus improving the speed of decision-making. This system can solve similar problems not only in the telecommunications industry but also in the general construction industry and many other industries involving on-site work. The market size is vast, and it is expected to be a beneficial tool for many companies. This allows the system to transcribe meeting audio data into text in real time, generate reports, and provide them to administrators, thereby improving work efficiency.
[0029] The system according to the embodiment comprises a voice analysis unit, a report generation unit, an information provision unit, and a visualization unit. The voice analysis unit transcribes meeting audio data into text in real time. The voice analysis unit converts meeting audio data into text data using, for example, speech recognition technology. The voice analysis unit can analyze and transcribe audio data using a generation AI. For example, the generation AI transcribes the content of the discussion into text in real time using speech recognition technology. The report generation unit generates a report of the meeting content based on the data transcribed by the voice analysis unit. For example, the report generation unit generates a report that includes key points of the discussion, decisions made, and action items based on the transcribed data. The report generation unit can analyze text data and generate a report using a generation AI. For example, the generation AI extracts key points of the discussion, decisions made, and action items based on the text data and generates a report. The information provision unit automatically aggregates the contents of the report generated by the report generation unit and provides it to the administrator. For example, the information provision unit automatically aggregates the contents of the report and provides it to the administrator. The Information Provision Department can analyze and summarize the contents of reports using a generation AI. For example, the generation AI can automatically summarize the progress and issues of meetings based on the contents of the reports and provide them to administrators. The Visualization Department visualizes the information provided by the Information Provision Department. For example, the Visualization Department visualizes the summarized data using methods such as dashboards and graphs. The Visualization Department can analyze and visualize the summarized data using a generation AI. For example, the generation AI visualizes the summarized data using methods such as dashboards and graphs. As a result, the system improves work efficiency by transcribing meeting audio data into text in real time, generating reports, and providing them to administrators.
[0030] The audio analysis unit transcribes meeting audio data into text in real time. Specifically, it uses high-precision speech recognition technology to sequentially convert speech during the meeting into text data. This speech recognition technology includes noise cancellation and speaker separation functions, enabling accurate transcription even when multiple speakers are speaking simultaneously. Furthermore, by using generative AI, it can understand the context of the audio data and apply appropriate punctuation and paragraph breaks. For example, the generative AI analyzes the audio data, understands the speaker's intentions and emotions, and transcribes it using appropriate expressions. This ensures that the meeting content is transcribed accurately and clearly. The audio analysis unit can also be customized to recognize specific technical terms and abbreviations and transcribe them appropriately. This ensures that industry-specific terminology and company abbreviations are accurately transcribed, improving accuracy in subsequent analysis and report generation. The audio analysis unit also has the function to immediately send the transcribed data to the report generation unit, so report generation begins immediately after the meeting ends.
[0031] The report generation unit generates a report of the meeting content based on the data transcribed into text by the voice analysis unit. Specifically, the report generation unit uses generation AI to analyze the text data and extract key points of the discussion, decisions made, and action items. The generation AI utilizes natural language processing technology to automatically identify important information from the text data and compile it into a report in an appropriate format. For example, the generation AI extracts important keywords and phrases from the speech and uses them to organize the key points of the discussion. Furthermore, for decisions and action items, it analyzes the context and tone of the speech to ensure they are accurately reflected in the report. The report generation unit also has a function to automatically format the generated report and output it in an easy-to-read layout. This ensures that the report is provided in a consistent format, making it easy for readers to understand the content. In addition, the report generation unit has a function to save the generated report to cloud storage, making it accessible to stakeholders at any time. This makes it easy to share and refer to the report, and facilitates smooth information transmission. Furthermore, the report generation unit also has a function to compare with past reports and perform trend analysis, allowing for long-term evaluation of the progress and results of meetings.
[0032] The Information Provision Department automatically aggregates the contents of reports generated by the Report Generation Department and provides them to administrators. Specifically, the Information Provision Department uses a generation AI to analyze the contents of reports and automatically aggregate meeting progress and issues. The generation AI extracts important data points from the reports and generates statistical information and graphs based on them. For example, the generation AI aggregates the number of decisions and action items and their progress from each meeting and provides this information to administrators. It can also automatically extract problems and risk factors by analyzing the contents of the reports and issue warnings to administrators. The Information Provision Department displays these aggregated results in a dashboard format, allowing administrators to grasp the situation at a glance. Furthermore, the Information Provision Department has a function to regularly update the aggregated results and provide the latest information. This allows administrators to always understand the latest situation and respond quickly. The Information Provision Department also has a function to link the aggregated results with other systems and departments, promoting information sharing throughout the organization. For example, linking the aggregated results with project management systems and risk management systems enables efficient operation of the entire organization.
[0033] The Visualization Unit visualizes the information provided by the Information Provision Unit. Specifically, the Visualization Unit uses a generation AI to analyze aggregated data and visualize it using methods such as dashboards and graphs. The generation AI analyzes the characteristics and trends of the data and selects the optimal visualization method. For example, the generation AI can generate timelines showing the progress of meetings or heatmaps showing problems. This makes it easier for administrators and stakeholders to visually grasp the information and enables quick decision-making. The Visualization Unit also has a function to provide users with customizable dashboards, allowing them to see the necessary information at a glance. For example, users can select the type of data and graphs to display according to their role and interests. The Visualization Unit also has a function to update the displayed content in real time in response to data updates, always providing the latest information. Furthermore, the Visualization Unit has a function to link generated graphs and dashboards with other systems and devices, promoting information sharing and collaboration. For example, by displaying visualized data on a projector or large display, information can be shared in real time during meetings. In this way, the Visualization Unit supports the visual understanding of information and supports quick and effective decision-making.
[0034] The audio analysis unit can transcribe meeting audio data into text in real time using speech recognition technology. For example, the audio analysis unit can convert meeting audio data into text data using speech recognition technology. The audio analysis unit can also analyze and transcribe audio data using generative AI. For example, the generative AI transcribes the content of the discussion into text in real time using speech recognition technology. This allows for accurate transcription of meeting audio data into text using speech recognition technology. Speech recognition technology includes, but is not limited to, deep learning and HMM (Hidden Markov Model). Some or all of the above-described processes in the audio analysis unit may be performed using or without the generative AI. For example, the audio analysis unit can input audio data into the generative AI and have the generative AI generate text data.
[0035] The report generation unit can generate a report containing the key points of the discussion, decisions made, and action items based on transcribed data. For example, the report generation unit can generate a report containing the key points of the discussion, decisions made, and action items based on transcribed data. The report generation unit can use a generation AI to analyze text data and generate a report. For example, the generation AI can extract the key points of the discussion, decisions made, and action items based on text data and generate a report. This allows for a clear record of the meeting content by generating a report containing the key points of the discussion, decisions made, and action items. Extraction of key points of the discussion includes, but is not limited to, methods such as keyword extraction and importance evaluation. Extraction of decisions includes, but is not limited to, methods such as voting results and consensus building. Extraction of action items includes, but is not limited to, methods such as task lists and deadlines. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input text data into the generation AI and have the generation AI perform report generation.
[0036] The information provision department can automatically aggregate the contents of reports and provide them to administrators. For example, the information provision department can automatically aggregate the contents of reports and provide them to administrators. The information provision department can analyze and aggregate the contents of reports using a generation AI. For example, the generation AI can automatically aggregate the progress and problems of meetings based on the contents of reports and provide them to administrators. This makes it easier for administrators to understand the progress and problems of meetings by automatically aggregating and providing the contents of reports to administrators. Aggregation includes, but is not limited to, the items to be aggregated and the aggregation method. Some or all of the above processing in the information provision department may be performed using a generation AI or not. For example, the information provision department can input the contents of reports into a generation AI and have the generation AI perform the aggregation.
[0037] The visualization unit can visualize aggregated data using methods such as dashboards and graph displays. For example, the visualization unit visualizes aggregated data using methods such as dashboards and graph displays. The visualization unit can analyze and visualize aggregated data using a generation AI. For example, the generation AI visualizes aggregated data using methods such as dashboards and graph displays. This allows administrators to intuitively understand the data by visualizing it. Visualization includes, but is not limited to, graphs, dashboards, and charts. Some or all of the above-described processes in the visualization unit may be performed using the generation AI, or they may be performed without the generation AI. For example, the visualization unit can input aggregated data into the generation AI and have the generation AI perform the visualization.
[0038] The speech analysis unit can identify speakers during speech analysis and distinguish and transcribe the content of each speaker's speech. For example, the speech analysis unit can use a generating AI to analyze the speaker's voiceprint and distinguish and transcribe each speaker's speech. The speech analysis unit can also use a generating AI to analyze the timing of speakers' speech and separate the text for each speaker. Furthermore, the speech analysis unit can use a generating AI to learn the speech characteristics of speakers and classify the content of their speech. This allows for more accurate recording of meeting content by distinguishing and transcribing the content of each speaker's speech. Speaker identification includes, but is not limited to, speech feature analysis and speaker recognition technology. Some or all of the above-described processes in the speech analysis unit may be performed using a generating AI or not. For example, the speech analysis unit can input speaker speech data into a generating AI and have the generating AI perform speaker identification and distinction of speech content.
[0039] The audio analysis unit can automatically remove background noise during audio analysis, thereby improving the quality of the audio data. For example, the audio analysis unit can use a generation AI to filter out background noise in real time and obtain clear audio data. The audio analysis unit can also use a generation AI to remove unwanted sounds from the audio data using noise reduction technology. Furthermore, the audio analysis unit can use a generation AI to analyze the frequency characteristics of the audio data and reduce noise components. This improves the quality of the audio data by removing background noise. Background noise removal includes, but is not limited to, noise filtering technology and noise cancellation. Some or all of the above processing in the audio analysis unit may be performed using a generation AI or not. For example, the audio analysis unit can input audio data into a generation AI and have the generation AI perform background noise removal.
[0040] The speech analysis unit can simultaneously analyze speech data in multiple languages and perform multilingual text conversion in real time. For example, the speech analysis unit's generation AI can simultaneously analyze speech data in multiple languages and convert each language into text. The speech analysis unit can also have the generation AI automatically determine the language of the speech data and apply an appropriate language model for text conversion. Furthermore, the speech analysis unit can have the generation AI perform real-time text conversion using multilingual speech recognition technology. This enables multilingual text conversion by simultaneously analyzing speech data in multiple languages. Multilingual support includes, but is not limited to, the types of languages supported and translation accuracy. Some or all of the above-described processes in the speech analysis unit may be performed using the generation AI or not. For example, the speech analysis unit can input speech data in multiple languages into the generation AI and have the generation AI perform multilingual text conversion.
[0041] The speech analysis unit can analyze the speaker's voice characteristics during speech analysis and estimate the emotion and intent of the utterance. For example, the speech analysis unit can use a generative AI to analyze the speaker's voice characteristics and estimate the emotion of the utterance. The speech analysis unit can also use a generative AI to analyze the speaker's voice intonation and speed and estimate the intent of the utterance. Furthermore, the speech analysis unit can use a generative AI to learn from the speaker's voice data and estimate emotions and intentions with high accuracy. As a result, by analyzing the speaker's voice characteristics, the emotion and intent of the utterance can be estimated with high accuracy. Analysis of voice characteristics includes, but is not limited to, pitch, tone, and rhythm. Estimation of emotions and intentions includes, but is not limited to, voice tone analysis and contextual analysis. Some or all of the above-described processes in the speech analysis unit may be performed using a generative AI or not. For example, the speech analysis unit can input the speaker's voice data into a generative AI and have the generative AI perform the estimation of emotions and intentions.
[0042] The report generation unit can adjust the level of detail in a report based on the importance of the discussions during report generation. For example, the report generation unit can generate a report with detailed explanations for important discussions. It can also generate a concise report for general discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the importance of a discussion and generate a report with an appropriate level of detail. This allows important information to be recorded in detail by adjusting the level of detail in the report based on the importance of the discussions. The evaluation of the importance of a discussion includes, but is not limited to, voting results and participants' opinions. The adjustment of the level of detail includes, but is not limited to, the depth of information and the level of detail in the description. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input discussion importance data into the generation AI and have the generation AI perform the adjustment of the level of detail in the report.
[0043] The report generation unit can apply different report formats depending on the category of the discussion when generating a report. For example, for technical discussions, the report generation unit can generate a report in a technical report format. It can also generate a report in a business report format for management discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the category of the discussion and generate a report in the appropriate format. This makes it easier to organize information by generating reports in the appropriate format according to the category of the discussion. The classification of discussion categories includes, but is not limited to, technical discussions and business discussions. The application of report formats includes, but is not limited to, templates and layouts. Some or all of the above processing in the report generation unit may be performed using the generation AI, or it may be performed without the generation AI. For example, the report generation unit can input discussion category data into the generation AI and have the generation AI apply the report format.
[0044] The report generation unit can determine the priority of reports based on the progress of discussions when generating reports. For example, the report generation unit can generate a preliminary report for ongoing discussions. It can also generate a detailed report for completed discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the progress of discussions and generate reports with appropriate priorities. This allows for the rapid provision of important information by determining the priority of reports based on the progress of discussions. The evaluation of the progress of discussions includes, but is not limited to, the progress of agenda items and the number of contributions. The determination of priorities includes, but is not limited to, the importance assessment and urgency assessment. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input discussion progress data into the generation AI and have the generation AI perform the determination of report priorities.
[0045] The report generation unit can adjust the order of the reports based on the relevance of the discussions during report generation. For example, the report generation unit may list important discussions first and less relevant discussions later. The report generation unit can also use a generation AI to automatically determine the relevance of discussions and generate the report in an appropriate order. Furthermore, the report generation unit can group highly relevant discussions and list them together in the report. This makes it easier to organize information by adjusting the order of the reports based on the relevance of the discussions. The evaluation of the relevance of discussions includes, but is not limited to, commonalities in topics or causal relationships. The adjustment of the order includes, but is not limited to, chronological order or order of importance. Some or all of the above processing in the report generation unit may be performed using a generation AI or not. For example, the report generation unit can input discussion relevance data into a generation AI and have the generation AI perform the adjustment of the report order.
[0046] The information provision unit can improve the accuracy of the information it provides by referring to past report data when providing information. For example, the information provision unit can use a generating AI to analyze past report data and improve the accuracy of the information provided. The information provision unit can also use a generating AI to provide highly relevant information based on past report data. Furthermore, the information provision unit can use a generating AI to learn from past report data and improve the accuracy of the information provided. As a result, the accuracy of the information provided is improved by referring to past report data. Referring to past report data includes, but is not limited to, database searches and filtering. Some or all of the above processing in the information provision unit may be performed using a generating AI or not. For example, the information provision unit can input past report data into a generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0047] The information provision unit can customize the information provided based on the administrator's attribute information. For example, the information provision unit can adjust the level of detail of the information provided according to the administrator's position. The information provision unit can also provide highly relevant information according to the administrator's area of expertise. Furthermore, the information provision unit can provide optimal information based on the administrator's past behavioral history. In this way, more appropriate information can be provided by customizing the information provided based on the administrator's attribute information. The administrator's attribute information includes, but is not limited to, positions and areas of expertise. Some or all of the above processing in the information provision unit may be performed using a generation AI, or it may be performed without a generation AI. For example, the information provision unit can input the administrator's attribute information into a generation AI and have the generation AI perform the customization of the information provided.
[0048] The information provision department can provide optimal information by considering the administrator's geographical location when providing information. For example, if the administrator is on-site, the information provision department will prioritize providing information relevant to the site. Furthermore, if the administrator is in the office, the information provision department can also provide overall progress updates. In addition, the information provision department can provide optimal information based on the administrator's location. This improves the usefulness of the information by providing optimal information based on the administrator's geographical location. Geographical location information includes, but is not limited to, GPS data and IP addresses. Some or all of the above processing in the information provision department may be performed using or without a generating AI. For example, the information provision department can input the administrator's geographical location information into a generating AI and have the generating AI provide optimal information.
[0049] The information provision unit can improve the accuracy of the information it provides by referring to relevant external data when providing information. For example, the information provision unit's generating AI can refer to an external database to improve the accuracy of the information provided. The information provision unit can also have the generating AI provide optimal information based on relevant external data. Furthermore, the information provision unit can have the generating AI learn from external data to improve the accuracy of the information provided. As a result, the accuracy of the information provided is improved by referring to relevant external data. Referencing external data includes, but is not limited to, public databases and API integration. Some or all of the above processing in the information provision unit may be performed using the generating AI or not. For example, the information provision unit can input external data into the generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0050] The visualization unit can optimize the displayed content by referring to past visualization data during visualization. For example, the visualization unit's generating AI can analyze past visualization data and provide optimal displayed content. The visualization unit can also have the generating AI display highly relevant information based on past visualization data. Furthermore, the visualization unit can improve the accuracy of the displayed content by having the generating AI learn from past visualization data. This improves the accuracy of the displayed content by referring to past visualization data. Referring to past visualization data includes, but is not limited to, database searches and filtering. Some or all of the above-described processes in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input past visualization data into the generating AI and have the generating AI perform the optimization of the displayed content.
[0051] The visualization unit can apply different visualization methods to each data category during visualization. For example, the visualization unit can provide a visualization in the format of a technical report for technical data. It can also provide a visualization in the format of a business report for management data. Furthermore, the visualization unit can have a generating AI automatically determine the data category and apply an appropriate visualization method. This makes it easier to understand the information by applying the appropriate visualization method according to the data category. Examples of data category classifications include, but are not limited to, technical data and business data. Examples of visualization method applications include, but are not limited to, heatmaps and treemaps. Some or all of the above-described processes in the visualization unit may be performed using the generating AI, or they may be performed without the generating AI. For example, the visualization unit can input the data category into the generating AI and have the generating AI apply the visualization method.
[0052] The visualization unit can adjust the displayed content based on the data submission date during visualization. For example, the visualization unit may prioritize the visualization of the latest data and display older data concisely. The visualization unit can also have the generation AI automatically determine the data submission date and provide appropriate displayed content. Furthermore, the visualization unit can highlight and display important data based on the submission date. This allows for the prioritization of the display of the latest information by adjusting the displayed content based on the data submission date. The acquisition of the submission date includes, but is not limited to, timestamps and submission dates. The adjustment of the displayed content includes, but is not limited to, prioritizing the display of the latest information and filtering out older information. Some or all of the above processing in the visualization unit may be performed using the generation AI or not. For example, the visualization unit can input the data submission date into the generation AI and have the generation AI perform the adjustment of the displayed content.
[0053] The visualization unit can improve the accuracy of the displayed content by referencing relevant external data during visualization. For example, the visualization unit's generating AI can reference an external database to improve the accuracy of the displayed content. The visualization unit can also have the generating AI provide optimal displayed content based on relevant external data. Furthermore, the visualization unit can have the generating AI learn from external data to improve the accuracy of the displayed content. As a result, the accuracy of the displayed content is improved by referencing relevant external data. Reference to external data includes, but is not limited to, public databases and API integration. Some or all of the above-described processes in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input external data into the generating AI and have the generating AI perform the task of improving the accuracy of the displayed content.
[0054] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0055] The speech analysis unit can estimate the speaker's intent during speech data analysis and classify the analysis results based on the estimated intent. For example, if the speaker is asking a question, the generation AI can classify the statement as a question and transcribe it appropriately. If the speaker is giving an instruction, the generation AI can classify the statement as an instruction and transcribe it quickly. Furthermore, if the speaker is expressing an opinion, the generation AI can classify the statement as an opinion and transcribe it in detail. This improves the accuracy of speech data analysis by classifying the analysis results according to the speaker's intent. Intent estimation is achieved, for example, using an intent estimation engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the speech analysis unit may be performed using a generation AI or not. For example, the speech analysis unit can input the speaker's speech data into a generation AI and have the generation AI perform the classification of the analysis results.
[0056] The report generation unit can adjust the level of detail in the generated report based on the importance of the discussion. For example, for important discussions, the generating AI can generate a report with detailed explanations. For general discussions, the generating AI can also generate a concise report. Furthermore, the generating AI can automatically determine the importance of a discussion and generate a report with an appropriate level of detail. This allows important information to be recorded in detail by adjusting the level of detail in the report based on the importance of the discussion. The evaluation of the importance of a discussion includes, but is not limited to, voting results and participants' opinions. The adjustment of the level of detail includes, but is not limited to, the depth of information and the level of detail in the description. Some or all of the above processing in the report generation unit may be performed using the generating AI or not. For example, the report generation unit can input discussion importance data into the generating AI and have the generating AI perform the adjustment of the level of detail in the report.
[0057] The information provision department can improve the accuracy of the information it provides by referring to past report data. For example, a generating AI can analyze past report data to improve the accuracy of the information provided. The generating AI can also provide highly relevant information based on past report data. Furthermore, the generating AI can learn from past report data to improve the accuracy of the information provided. In this way, the accuracy of the information provided is improved by referring to past report data. Referring to past report data includes, but is not limited to, database searches and filtering. Some or all of the above processing in the information provision department may be performed using the generating AI or not. For example, the information provision department can input past report data into the generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0058] The visualization unit can apply different visualization methods to the display of visualized data for each data category. For example, for technical data, the generating AI can provide visualization in the format of a technical report. Similarly, for management data, the generating AI can provide visualization in the format of a business report. Furthermore, the generating AI can automatically determine the data category and apply an appropriate visualization method. This makes it easier to understand the information by applying the appropriate visualization method according to the data category. The classification of data categories includes, but is not limited to, technical data and business data. The application of visualization methods includes, but is not limited to, heatmaps and treemaps. Some or all of the above processing in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input the data category into the generating AI and have the generating AI execute the application of the visualization method.
[0059] The audio analysis unit can automatically remove background noise during audio data analysis, thereby improving the quality of the audio data. For example, the generation AI can filter background noise in real time to obtain clear audio data. The generation AI can also remove unwanted sounds from the audio data using noise reduction technology. Furthermore, the generation AI can analyze the frequency characteristics of the audio data and reduce noise components. This improves the quality of the audio data by removing background noise. Background noise removal includes, but is not limited to, noise filtering technology and noise cancellation. Some or all of the above-described processes in the audio analysis unit may be performed using the generation AI or not. For example, the audio analysis unit can input audio data into the generation AI and have the generation AI perform background noise removal.
[0060] The following briefly describes the processing flow for example form 1.
[0061] Step 1: The audio analysis unit transcribes the meeting's audio data into text in real time. For example, it uses speech recognition technology or generative AI to analyze the audio data and convert it into text data. Step 2: The report generation unit generates a report of the meeting content based on the data transcribed into text by the voice analysis unit. For example, it uses a generation AI to extract key points of the discussion, decisions made, and action items, and then generates the report. Step 3: The Information Provision Department automatically compiles the contents of the reports generated by the Report Generation Department and provides them to the administrator. For example, it uses generation AI to analyze the contents of the reports, automatically compiles the progress of meetings and problems, and provides them to the administrator. Step 4: The Visualization Unit visualizes the information provided by the Information Provision Unit. For example, it visualizes aggregated data using methods such as dashboards and graphs.
[0062] (Example of form 2) The system according to an embodiment of the present invention is a system in which a generating AI transcribes discussions during pre-construction toolbox meetings into text in real time and generates a report of the meeting content. This system uses a generating AI to transcribe meeting audio data into text in real time and generates a report of the meeting content. For example, the generating AI uses speech recognition technology to convert the content of the discussion into text data. Next, the generating AI generates a report of the meeting content based on the transcribed data. This report includes the main points of the discussion, decisions made, action items, etc. Furthermore, the generating AI automatically aggregates the contents of the report and provides it to the administrator. The administrator can understand the progress of the meeting and any problems based on the aggregated data. This system automates the reporting of meeting content and the provision of information to the administrator, improving work efficiency. In addition, real-time transcription and report generation allow for rapid sharing of content, thus improving the speed of decision-making. This system can solve similar problems not only in the telecommunications industry but also in the general construction industry and many other industries involving on-site work. The market size is vast, and it is expected to be a beneficial tool for many companies. This allows the system to transcribe meeting audio data into text in real time, generate reports, and provide them to administrators, thereby improving work efficiency.
[0063] The system according to the embodiment comprises a voice analysis unit, a report generation unit, an information provision unit, and a visualization unit. The voice analysis unit transcribes meeting audio data into text in real time. The voice analysis unit converts meeting audio data into text data using, for example, speech recognition technology. The voice analysis unit can analyze and transcribe audio data using a generation AI. For example, the generation AI transcribes the content of the discussion into text in real time using speech recognition technology. The report generation unit generates a report of the meeting content based on the data transcribed by the voice analysis unit. For example, the report generation unit generates a report that includes key points of the discussion, decisions made, and action items based on the transcribed data. The report generation unit can analyze text data and generate a report using a generation AI. For example, the generation AI extracts key points of the discussion, decisions made, and action items based on the text data and generates a report. The information provision unit automatically aggregates the contents of the report generated by the report generation unit and provides it to the administrator. For example, the information provision unit automatically aggregates the contents of the report and provides it to the administrator. The Information Provision Department can analyze and summarize the contents of reports using a generation AI. For example, the generation AI can automatically summarize the progress and issues of meetings based on the contents of the reports and provide them to administrators. The Visualization Department visualizes the information provided by the Information Provision Department. For example, the Visualization Department visualizes the summarized data using methods such as dashboards and graphs. The Visualization Department can analyze and visualize the summarized data using a generation AI. For example, the generation AI visualizes the summarized data using methods such as dashboards and graphs. As a result, the system improves work efficiency by transcribing meeting audio data into text in real time, generating reports, and providing them to administrators.
[0064] The audio analysis unit transcribes meeting audio data into text in real time. Specifically, it uses high-precision speech recognition technology to sequentially convert speech during the meeting into text data. This speech recognition technology includes noise cancellation and speaker separation functions, enabling accurate transcription even when multiple speakers are speaking simultaneously. Furthermore, by using generative AI, it can understand the context of the audio data and apply appropriate punctuation and paragraph breaks. For example, the generative AI analyzes the audio data, understands the speaker's intentions and emotions, and transcribes it using appropriate expressions. This ensures that the meeting content is transcribed accurately and clearly. The audio analysis unit can also be customized to recognize specific technical terms and abbreviations and transcribe them appropriately. This ensures that industry-specific terminology and company abbreviations are accurately transcribed, improving accuracy in subsequent analysis and report generation. The audio analysis unit also has the function to immediately send the transcribed data to the report generation unit, so report generation begins immediately after the meeting ends.
[0065] The report generation unit generates a report of the meeting content based on the data transcribed into text by the voice analysis unit. Specifically, the report generation unit uses generation AI to analyze the text data and extract key points of the discussion, decisions made, and action items. The generation AI utilizes natural language processing technology to automatically identify important information from the text data and compile it into a report in an appropriate format. For example, the generation AI extracts important keywords and phrases from the speech and uses them to organize the key points of the discussion. Furthermore, for decisions and action items, it analyzes the context and tone of the speech to ensure they are accurately reflected in the report. The report generation unit also has a function to automatically format the generated report and output it in an easy-to-read layout. This ensures that the report is provided in a consistent format, making it easy for readers to understand the content. In addition, the report generation unit has a function to save the generated report to cloud storage, making it accessible to stakeholders at any time. This makes it easy to share and refer to the report, and facilitates smooth information transmission. Furthermore, the report generation unit also has a function to compare with past reports and perform trend analysis, allowing for long-term evaluation of the progress and results of meetings.
[0066] The Information Provision Department automatically aggregates the contents of reports generated by the Report Generation Department and provides them to administrators. Specifically, the Information Provision Department uses a generation AI to analyze the contents of reports and automatically aggregate meeting progress and issues. The generation AI extracts important data points from the reports and generates statistical information and graphs based on them. For example, the generation AI aggregates the number of decisions and action items and their progress from each meeting and provides this information to administrators. It can also automatically extract problems and risk factors by analyzing the contents of the reports and issue warnings to administrators. The Information Provision Department displays these aggregated results in a dashboard format, allowing administrators to grasp the situation at a glance. Furthermore, the Information Provision Department has a function to regularly update the aggregated results and provide the latest information. This allows administrators to always understand the latest situation and respond quickly. The Information Provision Department also has a function to link the aggregated results with other systems and departments, promoting information sharing throughout the organization. For example, linking the aggregated results with project management systems and risk management systems enables efficient operation of the entire organization.
[0067] The Visualization Unit visualizes the information provided by the Information Provision Unit. Specifically, the Visualization Unit uses a generation AI to analyze aggregated data and visualize it using methods such as dashboards and graphs. The generation AI analyzes the characteristics and trends of the data and selects the optimal visualization method. For example, the generation AI can generate timelines showing the progress of meetings or heatmaps showing problems. This makes it easier for administrators and stakeholders to visually grasp the information and enables quick decision-making. The Visualization Unit also has a function to provide users with customizable dashboards, allowing them to see the necessary information at a glance. For example, users can select the type of data and graphs to display according to their role and interests. The Visualization Unit also has a function to update the displayed content in real time in response to data updates, always providing the latest information. Furthermore, the Visualization Unit has a function to link generated graphs and dashboards with other systems and devices, promoting information sharing and collaboration. For example, by displaying visualized data on a projector or large display, information can be shared in real time during meetings. In this way, the Visualization Unit supports the visual understanding of information and supports quick and effective decision-making.
[0068] The audio analysis unit can transcribe meeting audio data into text in real time using speech recognition technology. For example, the audio analysis unit can convert meeting audio data into text data using speech recognition technology. The audio analysis unit can also analyze and transcribe audio data using generative AI. For example, the generative AI transcribes the content of the discussion into text in real time using speech recognition technology. This allows for accurate transcription of meeting audio data into text using speech recognition technology. Speech recognition technology includes, but is not limited to, deep learning and HMM (Hidden Markov Model). Some or all of the above-described processes in the audio analysis unit may be performed using or without the generative AI. For example, the audio analysis unit can input audio data into the generative AI and have the generative AI generate text data.
[0069] The report generation unit can generate a report containing the key points of the discussion, decisions made, and action items based on transcribed data. For example, the report generation unit can generate a report containing the key points of the discussion, decisions made, and action items based on transcribed data. The report generation unit can use a generation AI to analyze text data and generate a report. For example, the generation AI can extract the key points of the discussion, decisions made, and action items based on text data and generate a report. This allows for a clear record of the meeting content by generating a report containing the key points of the discussion, decisions made, and action items. Extraction of key points of the discussion includes, but is not limited to, methods such as keyword extraction and importance evaluation. Extraction of decisions includes, but is not limited to, methods such as voting results and consensus building. Extraction of action items includes, but is not limited to, methods such as task lists and deadlines. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input text data into the generation AI and have the generation AI perform report generation.
[0070] The information provision department can automatically aggregate the contents of reports and provide them to administrators. For example, the information provision department can automatically aggregate the contents of reports and provide them to administrators. The information provision department can analyze and aggregate the contents of reports using a generation AI. For example, the generation AI can automatically aggregate the progress and problems of meetings based on the contents of reports and provide them to administrators. This makes it easier for administrators to understand the progress and problems of meetings by automatically aggregating and providing the contents of reports to administrators. Aggregation includes, but is not limited to, the items to be aggregated and the aggregation method. Some or all of the above processing in the information provision department may be performed using a generation AI or not. For example, the information provision department can input the contents of reports into a generation AI and have the generation AI perform the aggregation.
[0071] The visualization unit can visualize aggregated data using methods such as dashboards and graph displays. For example, the visualization unit visualizes aggregated data using methods such as dashboards and graph displays. The visualization unit can analyze and visualize aggregated data using a generation AI. For example, the generation AI visualizes aggregated data using methods such as dashboards and graph displays. This allows administrators to intuitively understand the data by visualizing it. Visualization includes, but is not limited to, graphs, dashboards, and charts. Some or all of the above-described processes in the visualization unit may be performed using the generation AI, or they may be performed without the generation AI. For example, the visualization unit can input aggregated data into the generation AI and have the generation AI perform the visualization.
[0072] The voice analysis unit can estimate the user's emotions and adjust the accuracy of the voice data analysis based on the estimated emotions. For example, if the user is nervous, the voice analysis unit's generating AI can increase the accuracy of the voice data analysis and reduce misrecognition. If the user is relaxed, the voice analysis unit can also prioritize processing speed by having the generating AI return the analysis accuracy to normal. Furthermore, if the user is in a hurry, the voice analysis unit can increase the analysis accuracy and quickly transcribe the text. This improves the accuracy of the voice data analysis by adjusting the analysis accuracy according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generating AI. The generating AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the voice analysis unit may be performed using or without a generating AI. For example, the voice analysis unit can input user emotion data into a generating AI and have the generating AI adjust the analysis accuracy.
[0073] The speech analysis unit can identify speakers during speech analysis and distinguish and transcribe the content of each speaker's speech. For example, the speech analysis unit can use a generating AI to analyze the speaker's voiceprint and distinguish and transcribe each speaker's speech. The speech analysis unit can also use a generating AI to analyze the timing of speakers' speech and separate the text for each speaker. Furthermore, the speech analysis unit can use a generating AI to learn the speech characteristics of speakers and classify the content of their speech. This allows for more accurate recording of meeting content by distinguishing and transcribing the content of each speaker's speech. Speaker identification includes, but is not limited to, speech feature analysis and speaker recognition technology. Some or all of the above-described processes in the speech analysis unit may be performed using a generating AI or not. For example, the speech analysis unit can input speaker speech data into a generating AI and have the generating AI perform speaker identification and distinction of speech content.
[0074] The audio analysis unit can automatically remove background noise during audio analysis, thereby improving the quality of the audio data. For example, the audio analysis unit can use a generation AI to filter out background noise in real time and obtain clear audio data. The audio analysis unit can also use a generation AI to remove unwanted sounds from the audio data using noise reduction technology. Furthermore, the audio analysis unit can use a generation AI to analyze the frequency characteristics of the audio data and reduce noise components. This improves the quality of the audio data by removing background noise. Background noise removal includes, but is not limited to, noise filtering technology and noise cancellation. Some or all of the above processing in the audio analysis unit may be performed using a generation AI or not. For example, the audio analysis unit can input audio data into a generation AI and have the generation AI perform background noise removal.
[0075] The voice analysis unit can estimate the user's emotions and determine the priority of the analysis results based on the estimated emotions. For example, if the user is nervous, the voice analysis unit can prioritize analyzing and transcribing important statements. If the user is relaxed, the voice analysis unit can also analyze and transcribing all statements equally. Furthermore, if the user is in a hurry, the voice analysis unit can prioritize analyzing key points and transcribing them quickly. This allows for the priority of obtaining important information by determining the priority of the analysis results according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the voice analysis unit may be performed using or without the generative AI. For example, the voice analysis unit can input user emotion data into the generative AI and have the generative AI determine the priority of the analysis results.
[0076] The speech analysis unit can simultaneously analyze speech data in multiple languages and perform multilingual text conversion in real time. For example, the speech analysis unit's generation AI can simultaneously analyze speech data in multiple languages and convert each language into text. The speech analysis unit can also have the generation AI automatically determine the language of the speech data and apply an appropriate language model for text conversion. Furthermore, the speech analysis unit can have the generation AI perform real-time text conversion using multilingual speech recognition technology. This enables multilingual text conversion by simultaneously analyzing speech data in multiple languages. Multilingual support includes, but is not limited to, the types of languages supported and translation accuracy. Some or all of the above-described processes in the speech analysis unit may be performed using the generation AI or not. For example, the speech analysis unit can input speech data in multiple languages into the generation AI and have the generation AI perform multilingual text conversion.
[0077] The speech analysis unit can analyze the speaker's voice characteristics during speech analysis and estimate the emotion and intent of the utterance. For example, the speech analysis unit can use a generative AI to analyze the speaker's voice characteristics and estimate the emotion of the utterance. The speech analysis unit can also use a generative AI to analyze the speaker's voice intonation and speed and estimate the intent of the utterance. Furthermore, the speech analysis unit can use a generative AI to learn from the speaker's voice data and estimate emotions and intentions with high accuracy. As a result, by analyzing the speaker's voice characteristics, the emotion and intent of the utterance can be estimated with high accuracy. Analysis of voice characteristics includes, but is not limited to, pitch, tone, and rhythm. Estimation of emotions and intentions includes, but is not limited to, voice tone analysis and contextual analysis. Some or all of the above-described processes in the speech analysis unit may be performed using a generative AI or not. For example, the speech analysis unit can input the speaker's voice data into a generative AI and have the generative AI perform the estimation of emotions and intentions.
[0078] The report generation unit can estimate the user's emotions and adjust the way the report is presented based on the estimated emotions. For example, if the user is nervous, the report generation unit can generate a report using concise and clear language. If the user is relaxed, the report generation unit can also generate a report that includes detailed explanations. Furthermore, if the user is in a hurry, the report generation unit can generate a report that emphasizes the key points. In this way, by adjusting the way the report is presented according to the user's emotions, a more appropriate report can be generated. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the report generation unit may be performed using the generative AI or not. For example, the report generation unit can input user emotion data into the generative AI and have the generative AI adjust the way the report is presented.
[0079] The report generation unit can adjust the level of detail in a report based on the importance of the discussions during report generation. For example, the report generation unit can generate a report with detailed explanations for important discussions. It can also generate a concise report for general discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the importance of a discussion and generate a report with an appropriate level of detail. This allows important information to be recorded in detail by adjusting the level of detail in the report based on the importance of the discussions. The evaluation of the importance of a discussion includes, but is not limited to, voting results and participants' opinions. The adjustment of the level of detail includes, but is not limited to, the depth of information and the level of detail in the description. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input discussion importance data into the generation AI and have the generation AI perform the adjustment of the level of detail in the report.
[0080] The report generation unit can apply different report formats depending on the category of the discussion when generating a report. For example, for technical discussions, the report generation unit can generate a report in a technical report format. It can also generate a report in a business report format for management discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the category of the discussion and generate a report in the appropriate format. This makes it easier to organize information by generating reports in the appropriate format according to the category of the discussion. The classification of discussion categories includes, but is not limited to, technical discussions and business discussions. The application of report formats includes, but is not limited to, templates and layouts. Some or all of the above processing in the report generation unit may be performed using the generation AI, or it may be performed without the generation AI. For example, the report generation unit can input discussion category data into the generation AI and have the generation AI apply the report format.
[0081] The report generation unit can estimate the user's emotions and adjust the length of the report based on the estimated emotions. For example, if the user is stressed, the report generation unit can generate a short, concise report. If the user is relaxed, the report generation unit can also generate a longer report with more detailed explanations. Furthermore, if the user is in a hurry, the report generation unit can generate a concise, quick-to-read report. By adjusting the length of the report according to the user's emotions, a more appropriate report can be generated. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the report generation unit may be performed using or without a generative AI. For example, the report generation unit can input user emotion data into a generative AI and have the generative AI adjust the length of the report.
[0082] The report generation unit can determine the priority of reports based on the progress of discussions when generating reports. For example, the report generation unit can generate a preliminary report for ongoing discussions. It can also generate a detailed report for completed discussions. Furthermore, the report generation unit can use a generation AI to automatically determine the progress of discussions and generate reports with appropriate priorities. This allows for the rapid provision of important information by determining the priority of reports based on the progress of discussions. The evaluation of the progress of discussions includes, but is not limited to, the progress of agenda items and the number of contributions. The determination of priorities includes, but is not limited to, the importance assessment and urgency assessment. Some or all of the above processing in the report generation unit may be performed using the generation AI or not. For example, the report generation unit can input discussion progress data into the generation AI and have the generation AI perform the determination of report priorities.
[0083] The report generation unit can adjust the order of the reports based on the relevance of the discussions during report generation. For example, the report generation unit may list important discussions first and less relevant discussions later. The report generation unit can also use a generation AI to automatically determine the relevance of discussions and generate the report in an appropriate order. Furthermore, the report generation unit can group highly relevant discussions and list them together in the report. This makes it easier to organize information by adjusting the order of the reports based on the relevance of the discussions. The evaluation of the relevance of discussions includes, but is not limited to, commonalities in topics or causal relationships. The adjustment of the order includes, but is not limited to, chronological order or order of importance. Some or all of the above processing in the report generation unit may be performed using a generation AI or not. For example, the report generation unit can input discussion relevance data into a generation AI and have the generation AI perform the adjustment of the report order.
[0084] The information provider can estimate the user's emotions and determine the priority of the information to be provided based on the estimated emotions. For example, if the user is stressed, the information provider can prioritize providing important information. If the user is relaxed, the information provider can also provide all information equally. Furthermore, if the user is in a hurry, the information provider can prioritize providing key points. In this way, by determining the priority of information provided according to the user's emotions, important information can be provided preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the information provider may be performed using or without a generative AI. For example, the information provider can input user emotion data into a generative AI and have the generative AI determine the priority of the information to be provided.
[0085] The information provision unit can improve the accuracy of the information it provides by referring to past report data when providing information. For example, the information provision unit can use a generating AI to analyze past report data and improve the accuracy of the information provided. The information provision unit can also use a generating AI to provide highly relevant information based on past report data. Furthermore, the information provision unit can use a generating AI to learn from past report data and improve the accuracy of the information provided. As a result, the accuracy of the information provided is improved by referring to past report data. Referring to past report data includes, but is not limited to, database searches and filtering. Some or all of the above processing in the information provision unit may be performed using a generating AI or not. For example, the information provision unit can input past report data into a generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0086] The information provision unit can customize the information provided based on the administrator's attribute information. For example, the information provision unit can adjust the level of detail of the information provided according to the administrator's position. The information provision unit can also provide highly relevant information according to the administrator's area of expertise. Furthermore, the information provision unit can provide optimal information based on the administrator's past behavioral history. In this way, more appropriate information can be provided by customizing the information provided based on the administrator's attribute information. The administrator's attribute information includes, but is not limited to, positions and areas of expertise. Some or all of the above processing in the information provision unit may be performed using a generation AI, or it may be performed without a generation AI. For example, the information provision unit can input the administrator's attribute information into a generation AI and have the generation AI perform the customization of the information provided.
[0087] The information provider can estimate the user's emotions and adjust the display method of the information based on the estimated emotions. For example, if the user is nervous, the information provider can provide a simple and highly visible display method. If the user is relaxed, the information provider can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the information provider can provide a display method that emphasizes the key points. By adjusting the display method of the information according to the user's emotions, visibility is improved. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the information provider may be performed using the generative AI or not. For example, the information provider can input user emotion data into the generative AI and have the generative AI perform the adjustment of the display method.
[0088] The information provision department can provide optimal information by considering the administrator's geographical location when providing information. For example, if the administrator is on-site, the information provision department will prioritize providing information relevant to the site. Furthermore, if the administrator is in the office, the information provision department can also provide overall progress updates. In addition, the information provision department can provide optimal information based on the administrator's location. This improves the usefulness of the information by providing optimal information based on the administrator's geographical location. Geographical location information includes, but is not limited to, GPS data and IP addresses. Some or all of the above processing in the information provision department may be performed using or without a generating AI. For example, the information provision department can input the administrator's geographical location information into a generating AI and have the generating AI provide optimal information.
[0089] The information provision unit can improve the accuracy of the information it provides by referring to relevant external data when providing information. For example, the information provision unit's generating AI can refer to an external database to improve the accuracy of the information provided. The information provision unit can also have the generating AI provide optimal information based on relevant external data. Furthermore, the information provision unit can have the generating AI learn from external data to improve the accuracy of the information provided. As a result, the accuracy of the information provided is improved by referring to relevant external data. Referencing external data includes, but is not limited to, public databases and API integration. Some or all of the above processing in the information provision unit may be performed using the generating AI or not. For example, the information provision unit can input external data into the generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0090] The visualization unit can estimate the user's emotions and adjust the display method of the visualization based on the estimated user emotions. For example, if the user is tense, the visualization unit can provide a simple and highly visible display method. If the user is relaxed, the visualization unit can also provide a display method that includes detailed information. Furthermore, if the user is in a hurry, the visualization unit can provide a display method that emphasizes key points. This improves visibility by adjusting the display method of the visualization according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the visualization unit may be performed using or without a generative AI. For example, the visualization unit can input user emotion data into a generative AI and have the generative AI adjust the display method.
[0091] The visualization unit can optimize the displayed content by referring to past visualization data during visualization. For example, the visualization unit's generating AI can analyze past visualization data and provide optimal displayed content. The visualization unit can also have the generating AI display highly relevant information based on past visualization data. Furthermore, the visualization unit can improve the accuracy of the displayed content by having the generating AI learn from past visualization data. This improves the accuracy of the displayed content by referring to past visualization data. Referring to past visualization data includes, but is not limited to, database searches and filtering. Some or all of the above-described processes in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input past visualization data into the generating AI and have the generating AI perform the optimization of the displayed content.
[0092] The visualization unit can apply different visualization methods to each data category during visualization. For example, the visualization unit can provide a visualization in the format of a technical report for technical data. It can also provide a visualization in the format of a business report for management data. Furthermore, the visualization unit can have a generating AI automatically determine the data category and apply an appropriate visualization method. This makes it easier to understand the information by applying the appropriate visualization method according to the data category. Examples of data category classifications include, but are not limited to, technical data and business data. Examples of visualization method applications include, but are not limited to, heatmaps and treemaps. Some or all of the above-described processes in the visualization unit may be performed using the generating AI, or they may be performed without the generating AI. For example, the visualization unit can input the data category into the generating AI and have the generating AI apply the visualization method.
[0093] The visualization unit can estimate the user's emotions and determine the visualization priority based on the estimated emotions. For example, if the user is stressed, the visualization unit will prioritize the visualization of important information. If the user is relaxed, the visualization unit can also visualize all information equally. Furthermore, if the user is in a hurry, the visualization unit can prioritize the visualization of key points. In this way, by determining the visualization priority according to the user's emotions, important information can be displayed preferentially. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the visualization unit may be performed using the generative AI or not. For example, the visualization unit can input user emotion data into the generative AI and have the generative AI perform the determination of visualization priority.
[0094] The visualization unit can adjust the displayed content based on the data submission date during visualization. For example, the visualization unit may prioritize the visualization of the latest data and display older data concisely. The visualization unit can also have the generation AI automatically determine the data submission date and provide appropriate displayed content. Furthermore, the visualization unit can highlight and display important data based on the submission date. This allows for the prioritization of the display of the latest information by adjusting the displayed content based on the data submission date. The acquisition of the submission date includes, but is not limited to, timestamps and submission dates. The adjustment of the displayed content includes, but is not limited to, prioritizing the display of the latest information and filtering out older information. Some or all of the above processing in the visualization unit may be performed using the generation AI or not. For example, the visualization unit can input the data submission date into the generation AI and have the generation AI perform the adjustment of the displayed content.
[0095] The visualization unit can improve the accuracy of the displayed content by referencing relevant external data during visualization. For example, the visualization unit's generating AI can reference an external database to improve the accuracy of the displayed content. The visualization unit can also have the generating AI provide optimal displayed content based on relevant external data. Furthermore, the visualization unit can have the generating AI learn from external data to improve the accuracy of the displayed content. As a result, the accuracy of the displayed content is improved by referencing relevant external data. Reference to external data includes, but is not limited to, public databases and API integration. Some or all of the above-described processes in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input external data into the generating AI and have the generating AI perform the task of improving the accuracy of the displayed content.
[0096] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.
[0097] The speech analysis unit can estimate the speaker's emotions during the analysis of speech data and filter the analysis results based on the estimated emotions. For example, if the speaker is angry, the generating AI will analyze the utterance with particular care to reduce misrecognition. If the speaker is happy, the generating AI can analyze the utterance in a positive context and transcribe it appropriately. Furthermore, if the speaker is sad, the generating AI can carefully analyze the utterance and transcribe it while preserving the nuances of the emotion. This improves the accuracy of speech data analysis by filtering the analysis results according to the speaker's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generating AI. The generating AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the speech analysis unit may be performed using the generating AI or not. For example, the speech analysis unit can input speaker emotion data into the generating AI and have the generating AI perform filtering of the analysis results.
[0098] The report generation unit can estimate the user's emotions in the generated report and adjust the tone and style of the report based on the estimated emotions. For example, if the user is stressed, the generating AI can make the tone of the report calmer to improve readability. If the user is relaxed, the generating AI can make the style of the report more casual and use more approachable language. Furthermore, if the user is in a hurry, the generating AI can highlight the key points of the report to facilitate quick understanding. This improves the readability of the report by adjusting the tone and style according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generating AI. The generating AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above processing in the report generation unit may be performed using or without a generating AI. For example, the report generation unit can input user emotion data into a generating AI and have the generating AI adjust the tone and style of the report.
[0099] The information provision unit can estimate the user's emotions in response to the information it provides and adjust the presentation method based on the estimated emotions. For example, if the user is nervous, the generative AI can provide the information in a simple, highly visual format to make it easier to understand. If the user is relaxed, the generative AI can provide the information in a format that includes detailed information to promote deeper understanding. Furthermore, if the user is in a hurry, the generative AI can provide the information in a format that emphasizes the key points, allowing the user to quickly obtain the necessary information. In this way, adjusting the presentation method of information according to the user's emotions improves the receptivity of the information. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the information provision unit may be performed using the generative AI or not. For example, the information provision unit can input user emotion data into the generative AI and have the generative AI adjust the presentation method of the information.
[0100] The visualization unit can estimate the user's emotions in response to the display of visualized data and adjust the displayed content based on the estimated emotions. For example, if the user is stressed, the generating AI can provide simple, highly visible graphs and charts to make them easier to understand. If the user is relaxed, the generating AI can provide complex graphs and charts containing detailed data to encourage deeper analysis. Furthermore, if the user is in a hurry, the generating AI can provide graphs and charts that highlight key points, allowing the user to quickly obtain the necessary information. This improves data readability by adjusting the displayed content according to the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generating AI. The generating AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the visualization unit may be performed using or without a generating AI. For example, the visualization unit can input user emotion data into a generating AI and have the generating AI adjust the displayed content.
[0101] The speech analysis unit can estimate the speaker's emotions during the analysis of speech data and classify the analysis results based on the estimated emotions. For example, if the speaker is angry, the generating AI will analyze the utterance with particular care to reduce misrecognition. If the speaker is happy, the generating AI can analyze the utterance in a positive context and perform appropriate text conversion. Furthermore, if the speaker is sad, the generating AI can carefully analyze the utterance and convert it into text while preserving the nuances of the emotion. This improves the accuracy of speech data analysis by classifying the analysis results according to the speaker's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generating AI. The generating AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the speech analysis unit may be performed using the generating AI or not. For example, the speech analysis unit can input speaker emotion data into the generating AI and have the generating AI perform the classification of the analysis results.
[0102] The speech analysis unit can estimate the speaker's intent during speech data analysis and classify the analysis results based on the estimated intent. For example, if the speaker is asking a question, the generation AI can classify the statement as a question and transcribe it appropriately. If the speaker is giving an instruction, the generation AI can classify the statement as an instruction and transcribe it quickly. Furthermore, if the speaker is expressing an opinion, the generation AI can classify the statement as an opinion and transcribe it in detail. This improves the accuracy of speech data analysis by classifying the analysis results according to the speaker's intent. Intent estimation is achieved, for example, using an intent estimation engine or a generation AI. The generation AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processes in the speech analysis unit may be performed using a generation AI or not. For example, the speech analysis unit can input the speaker's speech data into a generation AI and have the generation AI perform the classification of the analysis results.
[0103] The report generation unit can adjust the level of detail in the generated report based on the importance of the discussion. For example, for important discussions, the generating AI can generate a report with detailed explanations. For general discussions, the generating AI can also generate a concise report. Furthermore, the generating AI can automatically determine the importance of a discussion and generate a report with an appropriate level of detail. This allows important information to be recorded in detail by adjusting the level of detail in the report based on the importance of the discussion. The evaluation of the importance of a discussion includes, but is not limited to, voting results and participants' opinions. The adjustment of the level of detail includes, but is not limited to, the depth of information and the level of detail in the description. Some or all of the above processing in the report generation unit may be performed using the generating AI or not. For example, the report generation unit can input discussion importance data into the generating AI and have the generating AI perform the adjustment of the level of detail in the report.
[0104] The information provision department can improve the accuracy of the information it provides by referring to past report data. For example, a generating AI can analyze past report data to improve the accuracy of the information provided. The generating AI can also provide highly relevant information based on past report data. Furthermore, the generating AI can learn from past report data to improve the accuracy of the information provided. In this way, the accuracy of the information provided is improved by referring to past report data. Referring to past report data includes, but is not limited to, database searches and filtering. Some or all of the above processing in the information provision department may be performed using the generating AI or not. For example, the information provision department can input past report data into the generating AI and have the generating AI perform the task of improving the accuracy of the information provided.
[0105] The visualization unit can apply different visualization methods to the display of visualized data for each data category. For example, for technical data, the generating AI can provide visualization in the format of a technical report. Similarly, for management data, the generating AI can provide visualization in the format of a business report. Furthermore, the generating AI can automatically determine the data category and apply an appropriate visualization method. This makes it easier to understand the information by applying the appropriate visualization method according to the data category. The classification of data categories includes, but is not limited to, technical data and business data. The application of visualization methods includes, but is not limited to, heatmaps and treemaps. Some or all of the above processing in the visualization unit may be performed using the generating AI or not. For example, the visualization unit can input the data category into the generating AI and have the generating AI execute the application of the visualization method.
[0106] The audio analysis unit can automatically remove background noise during audio data analysis, thereby improving the quality of the audio data. For example, the generation AI can filter background noise in real time to obtain clear audio data. The generation AI can also remove unwanted sounds from the audio data using noise reduction technology. Furthermore, the generation AI can analyze the frequency characteristics of the audio data and reduce noise components. This improves the quality of the audio data by removing background noise. Background noise removal includes, but is not limited to, noise filtering technology and noise cancellation. Some or all of the above-described processes in the audio analysis unit may be performed using the generation AI or not. For example, the audio analysis unit can input audio data into the generation AI and have the generation AI perform background noise removal.
[0107] The following briefly describes the processing flow for example form 2.
[0108] Step 1: The audio analysis unit transcribes the meeting's audio data into text in real time. For example, it uses speech recognition technology or generative AI to analyze the audio data and convert it into text data. Step 2: The report generation unit generates a report of the meeting content based on the data transcribed into text by the voice analysis unit. For example, it uses a generation AI to extract key points of the discussion, decisions made, and action items, and then generates the report. Step 3: The Information Provision Department automatically compiles the contents of the reports generated by the Report Generation Department and provides them to the administrator. For example, using generation AI, it analyzes the contents of the reports, automatically compiles the progress of meetings and problems, and provides them to the administrator. Step 4: The visualization unit visualizes the information provided by the information provision unit. For example, it visualizes aggregated data using methods such as dashboards and graphs.
[0109] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0110] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.
[0111] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0112] Each of the multiple elements described above, including the voice analysis unit, report generation unit, information provision unit, and visualization unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the voice analysis unit acquires meeting audio data using the microphone 38B of the smart device 14 and transcribes it into text in real time using the control unit 46A. The report generation unit generates a report based on the text data using the specific processing unit 290 of the data processing unit 12. The information provision unit aggregates the contents of the report using the specific processing unit 290 of the data processing unit 12 and provides it to the administrator. The visualization unit visualizes the aggregated data using the display 40A of the smart device 14 in a dashboard or graph display. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0113] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0114] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0115] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0116] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0117] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0118] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0119] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0120] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.
[0121] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0122] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0123] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0124] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0125] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0126] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0127] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0128] Each of the multiple elements described above, including the voice analysis unit, report generation unit, information provision unit, and visualization unit, is implemented, for example, by at least one of the smart glasses 214 and the data processing unit 12. For example, the voice analysis unit acquires meeting audio data using the microphone 238 of the smart glasses 214 and transcribes it into text in real time by the control unit 46A. The report generation unit generates a report based on the text data using the specific processing unit 290 of the data processing unit 12. The information provision unit aggregates the contents of the report using the specific processing unit 290 of the data processing unit 12 and provides it to the administrator. The visualization unit visualizes the aggregated data using the display of the smart glasses 214 in a dashboard or graph display. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0129] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0130] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0131] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0132] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0133] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0134] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0135] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0136] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0137] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0138] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0139] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0140] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0141] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0142] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0143] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0144] Each of the multiple elements described above, including the voice analysis unit, report generation unit, information provision unit, and visualization unit, is implemented by, for example, at least one of the headset terminal 314 and the data processing unit 12. For example, the voice analysis unit acquires meeting audio data using the microphone 238 of the headset terminal 314 and transcribes it into text in real time using the control unit 46A. The report generation unit generates a report based on the text data using, for example, the specific processing unit 290 of the data processing unit 12. The information provision unit aggregates the contents of the report using, for example, the specific processing unit 290 of the data processing unit 12 and provides it to the administrator. The visualization unit visualizes the aggregated data using, for example, the display 343 of the headset terminal 314 in a dashboard or graph display. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0145] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0146] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0147] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.
[0148] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0149] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0150] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).
[0151] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0152] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0153] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0154] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0155] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.
[0156] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.
[0157] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).
[0158] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0159] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.
[0160] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0161] Each of the multiple elements described above, including the voice analysis unit, report generation unit, information provision unit, and visualization unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the voice analysis unit acquires meeting audio data using the microphone 238 of the robot 414 and transcribes it into text in real time using the control unit 46A. The report generation unit generates a report based on the text data using the specific processing unit 290 of the data processing unit 12. The information provision unit aggregates the contents of the report using the specific processing unit 290 of the data processing unit 12 and provides it to the administrator. The visualization unit visualizes the aggregated data using the display of the robot 414 in a dashboard or graph display. The correspondence between each unit and the device or control unit is not limited to the example described above and can be modified in various ways.
[0162] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0163] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0164] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0165] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0166] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0167] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0168] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0169] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.
[0170] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0171] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0172] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0173] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0174] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0175] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0176] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0177] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.
[0178] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0179] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0180] (Note 1) The audio analysis unit transcribes meeting audio data into text in real time, A report generation unit generates a report of the meeting content based on the data converted into text by the aforementioned voice analysis unit, The information provision unit automatically aggregates the contents of the reports generated by the aforementioned report generation unit and provides them to the administrator, The system includes a visualization unit that visualizes the information provided by the information provision unit. A system characterized by the following features. (Note 2) The aforementioned voice analysis unit, Using speech recognition technology, audio data from meetings is converted to text in real time. The system described in Appendix 1, characterized by the features described herein. (Note 3) The report generation unit, Based on the transcribed data, generate a report containing key discussion points, decisions, and action items. The system described in Appendix 1, characterized by the features described herein. (Note 4) The aforementioned information provision unit, The report contents are automatically compiled and provided to the administrator. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned visualization unit, Visualize the aggregated data using methods such as dashboards and graphs. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned voice analysis unit, It estimates the user's emotions and adjusts the accuracy of the voice data analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned voice analysis unit, During speech analysis, the speaker is identified, and the content of each speaker's speech is distinguished and converted into text. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned voice analysis unit, During audio analysis, background noise is automatically removed to improve the quality of the audio data. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned voice analysis unit, It estimates the user's emotions and prioritizes the analysis results based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned voice analysis unit, During speech analysis, audio data in multiple languages is analyzed simultaneously, and multilingual text is generated in real time. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned voice analysis unit, During speech analysis, the speaker's vocal characteristics are analyzed to estimate the emotion and intent behind their statements. The system described in Appendix 1, characterized by the features described herein. (Note 12) The report generation unit, We estimate the user's emotions and adjust the way the report is presented based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 13) The report generation unit, When generating a report, adjust the level of detail in the report based on the importance of the discussion. The system described in Appendix 1, characterized by the features described herein. (Note 14) The report generation unit, When generating reports, apply different report formats depending on the discussion category. The system described in Appendix 1, characterized by the features described herein. (Note 15) The report generation unit, The system estimates the user's sentiment and adjusts the length of the report based on the estimated sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 16) The report generation unit, When generating reports, prioritize reports based on the progress of the discussions. The system described in Appendix 1, characterized by the features described herein. (Note 17) The report generation unit, When generating reports, adjust the order of reports based on the relevance of the discussions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned information provision unit, It estimates the user's emotions and prioritizes the information provided based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned information provision unit, When providing information, we refer to past report data to improve the accuracy of the information provided. The system described in Appendix 1, characterized by the features described herein. (Note 20) The aforementioned information provision unit, When providing information, customize the information provided based on the administrator's attribute information. The system described in Appendix 1, characterized by the features described herein. (Note 21) The aforementioned information provision unit, The system estimates the user's emotions and adjusts how the information is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 22) The aforementioned information provision unit, When providing information, we will consider the administrator's geographical location to provide the most relevant information. The system described in Appendix 1, characterized by the features described herein. (Note 23) The aforementioned information provision unit, When providing information, we refer to relevant external data to improve the accuracy of the information provided. The system described in Appendix 1, characterized by the features described herein. (Note 24) The aforementioned visualization unit, It estimates the user's emotions and adjusts the display method of the visualization based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned visualization unit, When visualizing data, the display content is optimized by referring to past visualization data. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned visualization unit, When visualizing data, different visualization methods are applied to each data category. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned visualization unit, It estimates the user's emotions and determines the visualization priority based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned visualization unit, When visualizing the data, adjust the displayed content based on when the data was submitted. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned visualization unit, When visualizing, the accuracy of the displayed content is improved by referencing relevant external data. The system described in Appendix 1, characterized by the features described herein. [Explanation of symbols]
[0181] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots
Claims
1. The audio analysis unit transcribes meeting audio data into text in real time, A report generation unit generates a report of the meeting content based on the data converted into text by the aforementioned voice analysis unit, The information provision unit automatically aggregates the contents of the reports generated by the aforementioned report generation unit and provides them to the administrator, The system includes a visualization unit that visualizes the information provided by the information provision unit. A system characterized by the following features.
2. The aforementioned voice analysis unit, Using speech recognition technology, audio data from meetings is converted to text in real time. The system according to feature 1.
3. The report generation unit, Based on the transcribed data, generate a report containing key discussion points, decisions, and action items. The system according to feature 1.
4. The aforementioned information provision unit, The report contents are automatically compiled and provided to the administrator. The system according to feature 1.
5. The aforementioned visualization unit, Visualize the aggregated data using methods such as dashboards and graphs. The system according to feature 1.
6. The aforementioned voice analysis unit, It estimates the user's emotions and adjusts the accuracy of the voice data analysis based on the estimated user emotions. The system according to feature 1.
7. The aforementioned voice analysis unit, During speech analysis, the speaker is identified, and the content of each speaker's speech is distinguished and converted into text. The system according to feature 1.
8. The aforementioned voice analysis unit, During audio analysis, background noise is automatically removed to improve the quality of the audio data. The system according to feature 1.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A