system
A system using speech and image recognition automates the creation of meeting minutes, reducing manual labor and ensuring high accuracy and efficiency in generating meeting records.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional methods for creating meeting minutes require significant manual labor for text conversion of voice recordings and visual materials, leading to time consumption and human errors.
A system integrating speech recognition to convert audio to text and image recognition to extract visual data, automatically generating accurate meeting minutes by combining and organizing this information.
Significantly reduces workload and ensures highly accurate, timely generation of meeting minutes that reflect the meeting content accurately.
Smart Images

Figure 2026074977000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional creation of minutes of an online meeting, a great deal of labor is required for manual text conversion of voice recordings and sequential arrangement and recording of visual materials displayed during the meeting, resulting in problems such as time consumption and human errors. An object of the present invention is to provide a system that can automatically generate highly accurate minutes of a meeting while reducing such labor.
Means for Solving the Problems
[0005] This invention provides a system comprising speech recognition means for acquiring audio recordings in real time and converting the audio into text, and image recognition means for capturing visual materials displayed during a meeting and extracting text and graph information. By integrating the information and automatically generating and outputting highly accurate meeting minutes, it is possible to significantly reduce the workload that has been a challenge in the past.
[0006] "Audio recording" refers to the collection of audio data of speech spoken during a meeting.
[0007] "Speech recognition means" refers to a technology or device that converts acquired speech recordings into text data.
[0008] "Visual materials" refer to information such as documents, slides, and graphs displayed using projectors, screens, or computers during a meeting.
[0009] "Means of capturing" refers to techniques or devices for acquiring digital images of visual materials.
[0010] "Image recognition means" refers to a technology or device that reads character information and graphic data from captured visual materials and extracts them as text data.
[0011] "Text information" refers to string data obtained through speech recognition and image recognition means, and is used for creating meeting minutes.
[0012] "Graph information" refers to the numerical data and interpretation information contained in charts and graphs within visual materials.
[0013] "Meeting minutes information" refers to meeting record data compiled by integrating and organizing information extracted from audio recordings and visual materials. [Brief explanation of the drawing]
[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention is implemented as a system for automatically creating highly accurate meeting minutes in online meetings. The system is primarily composed of speech recognition and image recognition means, and enables rapid and accurate meeting minute generation by efficiently processing audio data and visual information.
[0036] During a meeting, the terminal first captures the participants' voices in real time and records them digitally. Next, the server transfers this audio data to a speech recognition engine, which automatically converts it into text. The speech recognition process also includes speaker identification, ensuring clear distinction between speakers.
[0037] Simultaneously, the terminal captures screen images of visual materials such as slides and documents shared during the meeting. These image data are sent to a server, where image recognition algorithms are used to extract text and chart data. This information becomes crucial for later meeting minute generation.
[0038] Next, the server integrates the text data generated by speech recognition with the information extracted by image recognition. In this integration process, duplicate and inconsistent information is removed, and the data is processed to accurately reflect the meeting content. As a result, a well-organized meeting transcript is created, aligned with the timeline.
[0039] Ultimately, the generated meeting minutes are provided to the user and become immediately available for review and sharing right after the meeting ends. For example, in a project progress meeting, the minutes will be generated accurately, including not only the statement "The project is progressing smoothly" but also the data shown on the slide, such as "Progress: 70% complete."
[0040] This system eliminates the need for manual data recording and supports decision-making based on fast, highly accurate data.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] As soon as the online meeting begins, the device captures participants' voices in real time via the microphone and records them as digital data. This data is then prepared to be transferred to a server over the network.
[0044] Step 2:
[0045] The server inputs the received audio data into a speech recognition engine, performs noise reduction, and converts it into accurate text information. During this process, the speaker is identified, and the content of the speech is organized by speaker.
[0046] Step 3:
[0047] The device automatically takes screenshots of slides and documents shared during the meeting. This process is performed each time a new document is projected, and the captured image data is sent to the server.
[0048] Step 4:
[0049] The server processes the transmitted image data through an image recognition engine and uses OCR to extract text information as text data. It also identifies charts and graphs and analyzes numerical data if necessary.
[0050] Step 5:
[0051] The server performs a process of integrating text data obtained from speech recognition with text and numerical data obtained from image recognition. The data is scrutinized to avoid duplication and inconsistencies, and comprehensive, time-series data is generated.
[0052] Step 6:
[0053] The server organizes the integrated data and generates a final document as meeting minutes. It checks the structure of the document, summarizes it as needed, and formats it in a clear and user-friendly format.
[0054] Step 7:
[0055] Users receive the generated meeting minutes and review their contents. After review, they can share them with all participants as needed, enabling quick feedback and decision-making.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] In online meetings, it is difficult to individually manage participants' audio and visual materials, and there is a need to create meeting minutes accurately and quickly. However, conventional methods require manual recording, resulting in insufficient accuracy and efficiency. Therefore, a system is needed that effectively integrates audio data and visual materials and automatically generates meeting minutes that accurately reflect the content of the meeting.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text information, means for acquiring visual media displayed during a meeting, means for extracting text information and chart information from the visual media, means for integrating the text information obtained by the audio recognition means with the text information and chart information obtained by the image processing means and processing it as meeting minutes information, and means for immediately providing the generated meeting minutes information to users. This makes it possible to quickly and accurately generate meeting minutes from the audio and visual materials of an online meeting and immediately share the information with all participants.
[0061] "Audio data" refers to a digital recording of what meeting participants said.
[0062] "Acoustic recognition means" refers to processing means that analyze acquired audio data and convert it into textual information.
[0063] "Visual media" refers to visual materials such as slides and documents used during a meeting.
[0064] "Image processing means" refers to a processing device for extracting textual information and diagrammatic information from a visual medium.
[0065] "Textual information" refers to string data extracted from audio data or visual media.
[0066] "Graphical information" refers to data in the form of shapes or tables extracted from visual media.
[0067] "Integration" refers to the process of combining text information obtained through speech recognition with text information and chart information obtained through image processing into a single document.
[0068] "Meeting minutes information" refers to information in document form that records the content of a meeting.
[0069] "User" refers to an individual or organization that is able to receive and use the generated meeting minutes information.
[0070] This invention relates to a system for automatically generating accurate meeting minutes during online meetings. This system uses speech recognition and image processing means to extract and integrate information from participants' statements and visual materials, thereby generating meeting minutes that accurately reflect the content of the meeting.
[0071] The terminal captures the voice of each meeting participant in real time via a microphone device and records it as digital audio data. The recorded audio data is transmitted to a server via a communication network. This server has the capability to convert the audio data into text information, for example, using a speech recognition API. In this process, speaker identification is also performed, making it possible to clearly distinguish between speakers.
[0072] Simultaneously, the terminal captures visual media such as slides and documents used during the meeting using screen capture software. This image data is also sent to the server. The server uses image processing algorithms to extract text and diagram information from the visual media. For example, an image recognition service might be used for this process.
[0073] The generated data is integrated by the server, removing duplicates and inconsistencies to create organized meeting minutes. Finally, the generated minutes are provided to the user, allowing for immediate review and sharing immediately after the meeting ends. As a concrete example, in a project progress meeting, the minutes would include both the statement "The project is progressing smoothly" and the data displayed on the slide, such as "Progress: 70% complete."
[0074] Example prompt: "Please explain how to integrate speech recognition and image recognition for creating meeting minutes during online meetings."
[0075] Therefore, this invention is an innovative system that automates conventional manual recording work and can provide meeting minutes that reflect the content of meetings quickly and with high accuracy.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The terminal captures the voices of meeting participants in real time using a microphone device and converts them into digital audio data. The input is a raw audio signal, and the output is digital audio data (e.g., WAV format). This conversion process involves converting from analog to digital signals, specifically using a microphone device and audio capture software.
[0079] Step 2:
[0080] The terminal uses screen capture software to capture visual materials such as slides and documents used during a meeting. The input is the visual material displayed on the terminal, and the output is image data (e.g., PNG format). Specifically, the screen capture button is pressed to save the current screen.
[0081] Step 3:
[0082] The server receives audio data sent from the terminal. Next, it inputs this data into a speech recognition API, converting the audio data into text information. The input is digital audio data, and the output is text data. This process includes breaking down the audio waveform into phonemes and converting those phonemes into strings.
[0083] Step 4:
[0084] The server receives image data sent from the terminal and extracts text and diagram information from the image using an image recognition service. The input is image data, and the output is extracted text and diagram information. The image data is analyzed, and text is extracted using OCR (Optical Character Recognition) technology.
[0085] Step 5:
[0086] The server integrates text information obtained from audio data with information extracted from image data. During this process, it checks for duplicate or inconsistent information and generates meeting minutes in an organized format. Input is text information from speech recognition and image recognition, and output is integrated meeting minutes information. Information is arranged considering its relevance, and unnecessary data is removed.
[0087] Step 6:
[0088] The server immediately provides the generated meeting minutes information to the user. The input is integrated meeting minutes information, and the output is meeting minutes in document format delivered to the user. Specifically, the server converts the meeting minutes data to PDF or DOCX format and sends a download link to the user's device.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] In online consultations at virtual stores, there is a need to efficiently record customer conversations and integrate them with product data. However, conventional methods often lack consistency and accuracy because audio and visual information are processed separately. To solve this problem, it is necessary to develop a system that integrates customer conversations and product information and outputs them as a record immediately.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for acquiring acoustic information, means for acoustic analysis for converting the acoustic information into text information, and means for acquiring visual data presented during negotiations. This makes it possible to accurately acquire the content of conversations with customers as text information and immediately integrate it with product data.
[0094] "Acoustic information" refers to audio data and signals, and is the information used for audio recording and speech recognition.
[0095] "Acoustic analysis means" refers to technology that interprets acoustic information and converts it into text information. This may include a speech recognition engine.
[0096] "Visual data" refers to information presented as visual materials or images. This may include charts, graphs, and text.
[0097] "Image analysis means" refers to the process of extracting necessary information, such as text and illustrations, from visual data.
[0098] "Integrated processing means" refers to a technology that combines text information obtained by acoustic analysis means with information obtained by image analysis means and processes them in an integrated manner.
[0099] "Recorded information" refers to documents and files created based on analyzed text and image information.
[0100] This invention is a system that acquires and analyzes acoustic information in real time to effectively record the content of consultations with customers. Specifically, the server converts the audio data into text information using a speech recognition engine. The audio data is acquired by the terminal via the microphone of a smartphone or PC. The server further acquires visual data presented during negotiations and extracts text information and diagram data using image analysis means. Image recognition algorithms and OCR technology are used for this analysis. The image information is captured through a camera-equipped terminal and transmitted to the server.
[0101] This data is integrated on the server and output as a consistent consultation record. This ensures that customer conversations and product data are organized and recorded without any loss. For example, if a customer asks online, "How can I use this product?", the voice data is recognized in real time, and usage data is simultaneously extracted from product specifications and images taken by the device and automatically added to the consultation record. By utilizing a generative AI model, this integration process can be performed quickly and efficiently.
[0102] Examples of prompt messages include, "Please summarize the customer's consultation based on the following audio and images." This system can improve the quality of customer service in virtual stores.
[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0104] Step 1:
[0105] The user initiates an online consultation with a customer using a smartphone or computer. The user's device begins acquiring audio data from the customer in real time via the microphone. The input is audio data, which is then sent to the server for subsequent processing.
[0106] Step 2:
[0107] The server receives audio data and converts it into text information using a speech recognition engine. This process utilizes a generative AI model to analyze acoustic features and perform data processing and calculations to convert them into strings. The output is the textual information of the conversation with the customer.
[0108] Step 3:
[0109] The device uses its camera to photograph visual data presented during negotiations, such as product brochures or instruction manuals. The input is image data, which is then sent to the server.
[0110] Step 4:
[0111] The server acquires image data and extracts textual and illustrative information using image analysis tools. Specifically, it uses OCR technology to identify characters in the image and converts them into text using a generative AI model. The output is textual data of the visual information.
[0112] Step 5:
[0113] The server integrates the textual information from speech recognition results and the textual information from image analysis into a single consultation record using an integrated processing mechanism. This process associates the spoken words with the image information in chronological order, maintaining consistency throughout the integration. The output is the final consultation record document.
[0114] Step 6:
[0115] The server provides the user with the generated consultation record. The user reviews this record and makes corrections or additions as needed. They may also send the record to other systems or applications using prompts. For example, a prompt such as, "Please summarize the customer consultation based on the following audio and images," can be used to request the generating AI model to summarize the output data.
[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0117] This invention is implemented as an online meeting system that, in addition to speech recognition and image recognition, utilizes an emotion engine to capture the user's emotional state and reflect it in the meeting minutes. This system generates meeting minutes that not only record information but also take into account the emotional reactions of participants during the meeting.
[0118] During the meeting, the terminal captures participants' audio and video with high precision, and the audio data is recorded in digital format. Simultaneously, the video data is transmitted along with visual materials for analysis by the emotion sensor.
[0119] The server uses a speech recognition engine to convert audio data into text data in real time. This conversion incorporates speaker identification, and the speaker's identity is also recorded. In addition, screen capture data of shared documents is analyzed by an image recognition engine to extract text and graph information from the documents.
[0120] In parallel with these recognition processes, the server analyzes the user's emotions from the audio and video through an emotion engine. The emotion engine analyzes the tone, speed, and body language of the voice, and generates qualitative data on the participant's emotional state (e.g., joy, surprise, confusion, etc.).
[0121] Next, the server integrates the text data with the emotional state generated by the emotion engine and produces a document edited as meeting minutes. During this process, the user's reactions are recorded chronologically based on the content of the conversation, making significant emotional changes during the meeting visible.
[0122] For example, in a new product launch meeting, if a user says, "The new feature is groundbreaking," and the emotion engine detects surprise or excitement on the user's face, that emotional response will be appropriately added to the meeting minutes. As a result, information that takes into account the emotions of the meeting participants is recorded, which can function as useful information for decision-making after the meeting.
[0123] Users can immediately review and utilize meeting minutes containing this sentiment information after the meeting ends, leading to deeper insights into the meeting content and providing higher-quality support.
[0124] The following describes the processing flow.
[0125] Step 1:
[0126] The device captures participants' audio and video data via microphone and camera at the start of the meeting and records it as digital data. This data is then prepared to be sent to the server in real time.
[0127] Step 2:
[0128] The server inputs the received audio data into a speech recognition engine and converts it into text data. During this process, noise reduction and speaker identification are performed simultaneously to clearly record who spoke.
[0129] Step 3:
[0130] The server extracts visual materials displayed during the meeting from the video data and analyzes them using an image recognition engine. Here, it identifies text information and graphs within the materials and extracts the necessary information as text data.
[0131] Step 4:
[0132] The server uses an emotion engine to analyze participants' facial expressions and body language based on video data, identifying their emotional state. This includes changes in voice tone and speed, allowing for a detailed capture of the user's emotions.
[0133] Step 5:
[0134] The server integrates text data obtained from speech recognition, text and graph information obtained from image recognition, and sentiment data obtained from the sentiment engine. This integration process organizes the information chronologically and generates a unified meeting record.
[0135] Step 6:
[0136] The server delivers the generated meeting minutes to the dashboard, making it possible to verify the accuracy of the content. Here, emotional responses to individual statements are visualized and adjusted to make the meeting content easier to understand.
[0137] Step 7:
[0138] Users can review these meeting minutes after the meeting ends, enabling them to efficiently make decisions and take follow-up actions based on detailed meeting analysis information, including sentiment data.
[0139] (Example 2)
[0140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0141] Traditional online meeting systems could record meeting content as text, but they did not take into account participants' emotional responses. As a result, important emotional shifts and participant reactions during meetings were not utilized in decision-making. Furthermore, it was difficult to comprehensively integrate information from visual materials.
[0142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0143] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data for speech recognition, means for acquiring visual materials, means for image processing for extracting information from the visual materials, and means for analyzing the emotional state of the participants. This makes it possible to generate a meeting record that integrates audio information, visual material information, and the emotional responses of the participants into a single record.
[0144] "Audio data" refers to information recorded in digital format from what participants said during a meeting.
[0145] "Text data" refers to information in character format converted by speech recognition technology, representing the content of a meeting in written form.
[0146] "Visual materials" refer to visual information such as slides and graphs that are shared during a meeting.
[0147] "Image processing means" refers to techniques for extracting text information and graph information from visual materials.
[0148] "Emotional analysis means" refers to a technology that analyzes participants' emotions from audio and video data and generates qualitative data from it.
[0149] "Integration" is the process of combining audio data, information obtained from visual materials, and emotional data into a single meeting record.
[0150] "Meeting record information" refers to a document that is generated by integrating the content of a meeting with the emotional reactions of the participants at that time.
[0151] This invention is a system that improves the recording and analysis of information in online meetings. Specifically, a terminal uses a microphone and camera to capture participants' audio and video data with high accuracy and transmit it to a server. This system first uses a speech recognition engine to convert the audio data into text data in real time. This speech recognition process includes a speaker identification function, and speaker information is added to each statement.
[0152] In addition, the server acquires visual materials and uses an image processing engine to extract necessary information from the materials. This image processing engine has the ability to efficiently extract and analyze text and graph information from screen captures.
[0153] Furthermore, an emotion analysis engine analyzes participants' emotional states from audio and video data. Specifically, it analyzes voice tone and speed, facial expressions, and body language to qualitatively represent emotions such as joy and surprise.
[0154] This data is integrated and generated as meeting minutes that include emotional information. These minutes include key emotional shifts and highlights of topics during the meeting, helping users accurately understand the flow and key points of the meeting.
[0155] For example, when a user says "This feature is innovative" during a new product announcement, the server detects the user's emotional response, such as surprise or excitement, and records the details in the meeting minutes.
[0156] By using generative AI models, users can perform specific analyses and visualizations using prompts. For example, a prompt such as "Highlight the parts that surprised the participants" allows users to easily identify points of interest.
[0157] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0158] Step 1:
[0159] The terminal uses a microphone and camera at the start of the meeting to capture participants' audio and video data. This process takes real-time audio and video as input and generates audio and video files in digital format as output. These files are immediately transferred to the server.
[0160] Step 2:
[0161] The server converts the audio data received from the terminal into text data using a speech recognition engine. It receives an audio file as input and outputs text data, including speaker identification, through data processing. Specifically, it analyzes phonological patterns and maps each utterance to a specific speaker.
[0162] Step 3:
[0163] The server uses an image processing engine to acquire visual materials and extracts text and graph information from them. Here, screen-captured video data is used as input, and important text elements and graphics are extracted through data calculations and output as structured information.
[0164] Step 4:
[0165] The server uses an emotion analysis engine to analyze participants' emotional states from audio and video data. It receives audio tone and facial expression data from the video as input, and uses this data analysis to generate emotional indicators. The output presents each participant's emotional state as qualitative data.
[0166] Step 5:
[0167] The server integrates data obtained through speech recognition, image processing, and sentiment analysis to generate meeting minutes. This step takes text data, visual information, and sentiment data as input, and combines and edits them to output detailed meeting minutes based on the flow of the meeting.
[0168] Step 6:
[0169] Users can use a generative AI model to evaluate meeting minutes generated with prompts and provide feedback for improvement. For example, they can instruct the model to emphasize specific emotional responses, resulting in the minutes being highlighted as part of the output.
[0170] (Application Example 2)
[0171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0172] In brick-and-mortar stores, it is crucial to use customer interest and satisfaction to improve services. However, traditional methods have made it difficult to accurately capture customer emotions and reactions, making it challenging to obtain practical feedback for service improvement.
[0173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0174] In this invention, the server includes means for acquiring audio recordings, means for performing emotion analysis, integrating emotion numerical information obtained from audio and video, and generating customer response evaluations, and means for estimating customer interest and satisfaction using the emotion numerical information and utilizing that information in commercial activities. This makes it possible to obtain valuable feedback based on customers' real-time reactions.
[0175] "Audio recording" is the process of saving audio as digital data.
[0176] "Speech recognition means" refers to a technical device for converting speech data into text data.
[0177] "Visual materials" refer to digital content, including images and text information, that are displayed during meetings or commercial activities.
[0178] "Image recognition means" refers to a technological device that analyzes and extracts text information and graph information from visual materials.
[0179] "Meeting minutes" refers to information that records the content of meetings or business negotiations in written form.
[0180] "Emotional analysis" is a technology that analyzes emotional responses from audio and video data.
[0181] "Emotional numerical information" refers to emotional state data expressed in specific numerical values, obtained through emotion analysis.
[0182] "Response evaluation" refers to evaluation data generated based on customer emotions and reactions.
[0183] "Interest level" is a quantitative indicator that shows the degree of interest a customer is showing.
[0184] "Customer satisfaction" is an indicator that shows the degree of customer satisfaction with the services or products provided.
[0185] "Commercial activity" refers to all business activities related to the sale of goods or services.
[0186] The system for realizing this invention considers customer service in a physical store as an application example using speech recognition, image recognition, and emotion analysis technologies. The hardware and software configuration of the system will be described below.
[0187] The device uses smart glasses to capture audio and video during customer interactions. The smart glasses' camera records the customer's facial expressions and gestures in real time, and the microphone captures the customer's voice. This audio data is converted into text on a server using a speech recognition engine such as Google® Speech Recognition. Speech recognition includes speaker identification to determine which customer is speaking.
[0188] Simultaneously, the server uses the OpenCV library to analyze the captured video data and evaluates the customer's emotional state using an emotion analysis engine called Emotion Recognizer. Emotion analysis extracts numerical emotional information from the customer's voice tone, speed, and facial expressions to calculate levels of interest and satisfaction. This data provides valuable information for understanding customer needs and reactions, with the aim of utilizing it in commercial activities.
[0189] Store employees, who are also users of the smart glasses, can make appropriate product recommendations to customers based on the feedback they receive. By aggregating this kind of information, the overall service quality of the store can be improved.
[0190] For example, if a customer asks a question about a new product, sentiment analysis can instantly determine whether or not the customer is interested. As a result, employees can provide customized service, such as introducing related products that might pique the customer's interest.
[0191] Examples of prompts for a generative AI model are as follows:
[0192] "Based on the conversational text and customer sentiment data detected by the system installed in the smart glasses, please suggest what kind of product recommendations should be made."
[0193] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0194] Step 1:
[0195] The device uses the camera and microphone of smart glasses to capture the customer's video and audio in real time. The input for this step is the customer's face and conversational audio, and the output generates raw video and audio data.
[0196] Step 2:
[0197] The server receives audio data sent from the terminal and converts it into text using the Google Speech Recognition engine. The input is the captured audio data, and the output is text data in which the audio content has been converted into written information. Speaker identification information is also added simultaneously through speech recognition.
[0198] Step 3:
[0199] The server analyzes video data using the OpenCV library. The input is video data sent from the terminal, and the output is numerical emotion information. In this process, Emotion Recognizer is used to quantify the customer's emotional state from the video.
[0200] Step 4:
[0201] The server integrates text data and sentiment numerical information obtained from speech recognition to generate a customer response evaluation. The input is the text data and sentiment numerical information obtained from the previous step, and the output is evaluation data that assesses the customer's level of interest and satisfaction.
[0202] Step 5:
[0203] The server provides feedback to the user regarding the generated response evaluation. The user, i.e., the store employee, then uses this information to immediately make the most appropriate product recommendations to the customer. The input is the customer's response evaluation, and the output is a customized product recommendation tailored to the customer's needs. This feedback process enables employees to provide more effective customer service.
[0204] Step 6:
[0205] The user inputs the generated prompt sentences into the generating AI model to obtain further suggestions and responses. At this stage, the prompt sentences shown previously are used, and further service improvements are made based on the output from the generating AI model.
[0206] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0207] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0208] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0209] [Second Embodiment]
[0210] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0211] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0212] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0213] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0214] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0215] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0216] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0217] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0218] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0219] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0220] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0221] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0222] This invention is implemented as a system for automatically creating highly accurate meeting minutes in online meetings. The system is primarily composed of speech recognition and image recognition means, and enables rapid and accurate meeting minute generation by efficiently processing audio data and visual information.
[0223] During a meeting, the terminal first captures the participants' voices in real time and records them digitally. Next, the server transfers this audio data to a speech recognition engine, which automatically converts it into text. The speech recognition process also includes speaker identification, ensuring clear distinction between speakers.
[0224] Simultaneously, the terminal captures screen images of visual materials such as slides and documents shared during the meeting. These image data are sent to a server, where image recognition algorithms are used to extract text and chart data. This information becomes crucial for later meeting minute generation.
[0225] Next, the server integrates the text data generated by speech recognition with the information extracted by image recognition. In this integration process, duplicate and inconsistent information is removed, and the data is processed to accurately reflect the meeting content. As a result, a well-organized meeting transcript is created, aligned with the timeline.
[0226] Ultimately, the generated meeting minutes are provided to the user and become immediately available for review and sharing right after the meeting ends. For example, in a project progress meeting, the minutes will be generated accurately, including not only the statement "The project is progressing smoothly" but also the data shown on the slide, such as "Progress: 70% complete."
[0227] This system eliminates the need for manual data recording and supports decision-making based on fast, highly accurate data.
[0228] The following describes the processing flow.
[0229] Step 1:
[0230] As soon as the online meeting begins, the device captures participants' voices in real time via the microphone and records them as digital data. This data is then prepared to be transferred to a server over the network.
[0231] Step 2:
[0232] The server inputs the received audio data into a speech recognition engine, performs noise reduction, and converts it into accurate text information. During this process, the speaker is identified, and the content of the speech is organized by speaker.
[0233] Step 3:
[0234] The device automatically takes screenshots of slides and documents shared during the meeting. This process is performed each time a new document is projected, and the captured image data is sent to the server.
[0235] Step 4:
[0236] The server processes the transmitted image data through an image recognition engine and uses OCR to extract text information as text data. It also identifies charts and graphs and analyzes numerical data if necessary.
[0237] Step 5:
[0238] The server performs a process of integrating text data obtained from speech recognition with text and numerical data obtained from image recognition. The data is scrutinized to avoid duplication and inconsistencies, and comprehensive, time-series data is generated.
[0239] Step 6:
[0240] The server organizes the integrated data and generates a final document as meeting minutes. It checks the structure of the document, summarizes it as needed, and formats it in a clear and user-friendly format.
[0241] Step 7:
[0242] Users receive the generated meeting minutes and review their contents. After review, they can share them with all participants as needed, enabling quick feedback and decision-making.
[0243] (Example 1)
[0244] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0245] In online meetings, it is difficult to individually manage participants' audio and visual materials, and there is a need to create meeting minutes accurately and quickly. However, conventional methods require manual recording, resulting in insufficient accuracy and efficiency. Therefore, a system is needed that effectively integrates audio data and visual materials and automatically generates meeting minutes that accurately reflect the content of the meeting.
[0246] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0247] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text information, means for acquiring visual media displayed during a meeting, means for extracting text information and chart information from the visual media, means for integrating the text information obtained by the audio recognition means with the text information and chart information obtained by the image processing means and processing it as meeting minutes information, and means for immediately providing the generated meeting minutes information to users. This makes it possible to quickly and accurately generate meeting minutes from the audio and visual materials of an online meeting and immediately share the information with all participants.
[0248] "Audio data" refers to a digital recording of what meeting participants said.
[0249] "Acoustic recognition means" refers to processing means that analyze acquired audio data and convert it into textual information.
[0250] "Visual media" refers to visual materials such as slides and documents used during a meeting.
[0251] "Image processing means" refers to a processing device for extracting textual information and diagrammatic information from a visual medium.
[0252] "Textual information" refers to string data extracted from audio data or visual media.
[0253] "Graphical information" refers to data in the form of shapes or tables extracted from visual media.
[0254] "Integration" refers to the process of combining text information obtained through speech recognition with text information and chart information obtained through image processing into a single document.
[0255] "Meeting minutes information" refers to information in document form that records the content of a meeting.
[0256] "User" refers to an individual or organization that is able to receive and use the generated meeting minutes information.
[0257] This invention relates to a system for automatically generating accurate meeting minutes during online meetings. This system uses speech recognition and image processing means to extract and integrate information from participants' statements and visual materials, thereby generating meeting minutes that accurately reflect the content of the meeting.
[0258] The terminal captures the voice of each meeting participant in real time via a microphone device and records it as digital audio data. The recorded audio data is transmitted to a server via a communication network. This server has the capability to convert the audio data into text information, for example, using a speech recognition API. In this process, speaker identification is also performed, making it possible to clearly distinguish between speakers.
[0259] Simultaneously, the terminal captures visual media such as slides and documents used during the meeting using screen capture software. This image data is also sent to the server. The server uses image processing algorithms to extract text and diagram information from the visual media. For example, an image recognition service might be used for this process.
[0260] The generated data is integrated by the server, removing duplicates and inconsistencies to create organized meeting minutes. Finally, the generated minutes are provided to the user, allowing for immediate review and sharing immediately after the meeting ends. As a concrete example, in a project progress meeting, the minutes would include both the statement "The project is progressing smoothly" and the data displayed on the slide, such as "Progress: 70% complete."
[0261] Example prompt: "Please explain how to integrate speech recognition and image recognition for creating meeting minutes during online meetings."
[0262] Therefore, this invention is an innovative system that automates conventional manual recording work and can provide meeting minutes that reflect the content of meetings quickly and with high accuracy.
[0263] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0264] Step 1:
[0265] The terminal captures the voices of meeting participants in real time using a microphone device and converts them into digital audio data. The input is a raw audio signal, and the output is digital audio data (e.g., WAV format). This conversion process involves converting from analog to digital signals, specifically using a microphone device and audio capture software.
[0266] Step 2:
[0267] The terminal uses screen capture software to capture visual materials such as slides and documents used during a meeting. The input is the visual material displayed on the terminal, and the output is image data (e.g., PNG format). Specifically, the screen capture button is pressed to save the current screen.
[0268] Step 3:
[0269] The server receives audio data sent from the terminal. Next, it inputs this data into a speech recognition API, converting the audio data into text information. The input is digital audio data, and the output is text data. This process includes breaking down the audio waveform into phonemes and converting those phonemes into strings.
[0270] Step 4:
[0271] The server receives image data sent from the terminal and extracts text and diagram information from the image using an image recognition service. The input is image data, and the output is extracted text and diagram information. The image data is analyzed, and text is extracted using OCR (Optical Character Recognition) technology.
[0272] Step 5:
[0273] The server integrates text information obtained from audio data with information extracted from image data. During this process, it checks for duplicate or inconsistent information and generates meeting minutes in an organized format. Input is text information from speech recognition and image recognition, and output is integrated meeting minutes information. Information is arranged considering its relevance, and unnecessary data is removed.
[0274] Step 6:
[0275] The server immediately provides the generated meeting minutes information to the user. The input is integrated meeting minutes information, and the output is meeting minutes in document format delivered to the user. Specifically, the server converts the meeting minutes data to PDF or DOCX format and sends a download link to the user's device.
[0276] (Application Example 1)
[0277] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0278] In online consultations in virtual stores, it is required to efficiently record the content of conversations with customers and integrate it with product data. However, in conventional methods, since voice and visual information are processed separately, the consistency and accuracy of information are often lacking. To solve this problem, it is necessary to develop a system that integrally processes conversations with customers and product information and immediately outputs it as a record.
[0279] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following respective means.
[0280] In this invention, the server includes means for acquiring acoustic information, acoustic analysis means for converting the acoustic information into character information, and means for acquiring visual data presented during negotiation. Thereby, it becomes possible to accurately acquire the content of conversations with customers as character information and immediately integrate it with product data.
[0281] "Acoustic information" refers to data and signals of voice, and is information used for voice recording and voice recognition.
[0282] "Acoustic analysis means" refers to a technology for interpreting acoustic information and converting it into text information. This may include a voice recognition engine.
[0283] "Visual data" refers to information presented as visual materials or images. This may include charts and characters.
[0284] "Image analysis means" refers to a process of extracting necessary information, such as characters and illustrated information, from visual data.
[0285] "Integrated processing means" refers to a technology for combining text information by acoustic analysis means and information by image analysis means and integrally processing them.
[0286] "Recorded information" refers to documents and files created based on the analyzed text information and image information.
[0287] This invention is a system that acquires and analyzes acoustic information in real time to effectively record the content of consultations with customers. Specifically, the server converts the audio data into text information using a speech recognition engine. The audio data is acquired by the terminal via the microphone of a smartphone or PC. The server further acquires visual data presented during negotiations and extracts text information and diagram data using image analysis means. Image recognition algorithms and OCR technology are used for this analysis. The image information is captured through a camera-equipped terminal and transmitted to the server.
[0288] This data is integrated on the server and output as a consistent consultation record. This ensures that customer conversations and product data are organized and recorded without any loss. For example, if a customer asks online, "How can I use this product?", the voice data is recognized in real time, and usage data is simultaneously extracted from product specifications and images taken by the device and automatically added to the consultation record. By utilizing a generative AI model, this integration process can be performed quickly and efficiently.
[0289] Examples of prompt messages include, "Please summarize the customer's consultation based on the following audio and images." This system can improve the quality of customer service in virtual stores.
[0290] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0291] Step 1:
[0292] The user initiates an online consultation with a customer using a smartphone or computer. The user's device begins acquiring audio data from the customer in real time via the microphone. The input is audio data, which is then sent to the server for subsequent processing.
[0293] Step 2:
[0294] The server receives audio data and converts it into text information using a speech recognition engine. This process utilizes a generative AI model to analyze acoustic features and perform data processing and calculations to convert them into strings. The output is the textual information of the conversation with the customer.
[0295] Step 3:
[0296] The device uses its camera to photograph visual data presented during negotiations, such as product brochures or instruction manuals. The input is image data, which is then sent to the server.
[0297] Step 4:
[0298] The server acquires image data and extracts textual and illustrative information using image analysis tools. Specifically, it uses OCR technology to identify characters in the image and converts them into text using a generative AI model. The output is textual data of the visual information.
[0299] Step 5:
[0300] The server integrates the textual information from speech recognition results and the textual information from image analysis into a single consultation record using an integrated processing mechanism. This process associates the spoken words with the image information in chronological order, maintaining consistency throughout the integration. The output is the final consultation record document.
[0301] Step 6:
[0302] The server provides the user with the generated consultation record. The user reviews this record and makes corrections or additions as needed. They may also send the record to other systems or applications using prompts. For example, a prompt such as, "Please summarize the customer consultation based on the following audio and images," can be used to request the generating AI model to summarize the output data.
[0303] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion recognition model 59 and perform specific processing using the user's emotions.
[0304] In addition to speech recognition and image recognition, the present invention is implemented as an online meeting system that utilizes an emotion engine to capture the user's emotional state and reflects this in the minutes. With this system, minutes are generated that not only record information but also take into account the emotional reactions of participants during the meeting.
[0305] During the meeting, the terminal captures the voices and videos of the participants with high precision, and the voice data is recorded in digital format. At the same time, the video data is transmitted so that it can be analyzed by an emotion sensor together with visual materials.
[0306] The server uses a speech recognition engine to convert the voice data into text data in real time. This conversion incorporates a speaker identification function, and who the speaker is is also recorded. In addition, the screen capture data of the shared materials is analyzed by an image recognition engine, and the text information and graph information within the materials are extracted.
[0307] In parallel with these recognition processes, the server analyzes the user's emotions from the voice and video through an emotion engine. The emotion engine analyzes the tone, speed, body language, etc. of the voice and generates the emotional state of the participants (e.g., joy, surprise, confusion, etc.) as qualitative data.
[0308] Next, the server integrates the text data with the emotional state obtained by the emotion engine and generates a document edited as minutes. At this time, the user's reactions are recorded in chronological order based on the content of the conversation, and important emotional changes during the meeting are visualized.
[0309] For example, in a new product launch meeting, if a user says, "The new feature is groundbreaking," and the emotion engine detects surprise or excitement on the user's face, that emotional response will be appropriately added to the meeting minutes. As a result, information that takes into account the emotions of the meeting participants is recorded, which can function as useful information for decision-making after the meeting.
[0310] Users can immediately review and utilize meeting minutes containing this sentiment information after the meeting ends, leading to deeper insights into the meeting content and providing higher-quality support.
[0311] The following describes the processing flow.
[0312] Step 1:
[0313] The device captures participants' audio and video data via microphone and camera at the start of the meeting and records it as digital data. This data is then prepared to be sent to the server in real time.
[0314] Step 2:
[0315] The server inputs the received audio data into a speech recognition engine and converts it into text data. During this process, noise reduction and speaker identification are performed simultaneously to clearly record who spoke.
[0316] Step 3:
[0317] The server extracts visual materials displayed during the meeting from the video data and analyzes them using an image recognition engine. Here, it identifies text information and graphs within the materials and extracts the necessary information as text data.
[0318] Step 4:
[0319] The server uses an emotion engine to analyze participants' facial expressions and body language based on video data, identifying their emotional state. This includes changes in voice tone and speed, allowing for a detailed capture of the user's emotions.
[0320] Step 5:
[0321] The server integrates text data obtained from speech recognition, text and graph information obtained from image recognition, and sentiment data obtained from the sentiment engine. This integration process organizes the information chronologically and generates a unified meeting record.
[0322] Step 6:
[0323] The server delivers the generated meeting minutes to the dashboard, making it possible to verify the accuracy of the content. Here, emotional responses to individual statements are visualized and adjusted to make the meeting content easier to understand.
[0324] Step 7:
[0325] Users can review these meeting minutes after the meeting ends, enabling them to efficiently make decisions and take follow-up actions based on detailed meeting analysis information, including sentiment data.
[0326] (Example 2)
[0327] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0328] Traditional online meeting systems could record meeting content as text, but they did not take into account participants' emotional responses. As a result, important emotional shifts and participant reactions during meetings were not utilized in decision-making. Furthermore, it was difficult to comprehensively integrate information from visual materials.
[0329] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0330] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data for speech recognition, means for acquiring visual materials, means for image processing for extracting information from the visual materials, and means for analyzing the emotional state of the participants. This makes it possible to generate a meeting record that integrates audio information, visual material information, and the emotional responses of the participants into a single record.
[0331] "Audio data" refers to information recorded in digital format from what participants said during a meeting.
[0332] "Text data" refers to information in character format converted by speech recognition technology, representing the content of a meeting in written form.
[0333] "Visual materials" refer to visual information such as slides and graphs that are shared during a meeting.
[0334] "Image processing means" refers to techniques for extracting text information and graph information from visual materials.
[0335] "Emotional analysis means" refers to a technology that analyzes participants' emotions from audio and video data and generates qualitative data from it.
[0336] "Integration" is the process of combining audio data, information obtained from visual materials, and emotional data into a single meeting record.
[0337] "Meeting record information" refers to a document that is generated by integrating the content of a meeting with the emotional reactions of the participants at that time.
[0338] This invention is a system that improves the recording and analysis of information in online meetings. Specifically, a terminal uses a microphone and camera to capture participants' audio and video data with high accuracy and transmit it to a server. This system first uses a speech recognition engine to convert the audio data into text data in real time. This speech recognition process includes a speaker identification function, and speaker information is added to each statement.
[0339] In addition, the server acquires visual materials and uses an image processing engine to extract necessary information from the materials. This image processing engine has the ability to efficiently extract and analyze text and graph information from screen captures.
[0340] Furthermore, an emotion analysis engine analyzes participants' emotional states from audio and video data. Specifically, it analyzes voice tone and speed, facial expressions, and body language to qualitatively represent emotions such as joy and surprise.
[0341] This data is integrated and generated as meeting minutes that include emotional information. These minutes include key emotional shifts and highlights of topics during the meeting, helping users accurately understand the flow and key points of the meeting.
[0342] For example, when a user says "This feature is innovative" during a new product announcement, the server detects the user's emotional response, such as surprise or excitement, and records the details in the meeting minutes.
[0343] By using generative AI models, users can perform specific analyses and visualizations using prompts. For example, a prompt such as "Highlight the parts that surprised the participants" allows users to easily identify points of interest.
[0344] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0345] Step 1:
[0346] The terminal uses a microphone and camera at the start of the meeting to capture participants' audio and video data. This process takes real-time audio and video as input and generates audio and video files in digital format as output. These files are immediately transferred to the server.
[0347] Step 2:
[0348] The server converts the audio data received from the terminal into text data using a speech recognition engine. It receives an audio file as input and outputs text data, including speaker identification, through data processing. Specifically, it analyzes phonological patterns and maps each utterance to a specific speaker.
[0349] Step 3:
[0350] The server uses an image processing engine to acquire visual materials and extracts text and graph information from them. Here, screen-captured video data is used as input, and important text elements and graphics are extracted through data calculations and output as structured information.
[0351] Step 4:
[0352] The server uses an emotion analysis engine to analyze participants' emotional states from audio and video data. It receives audio tone and facial expression data from the video as input, and uses this data analysis to generate emotional indicators. The output presents each participant's emotional state as qualitative data.
[0353] Step 5:
[0354] The server integrates data obtained through speech recognition, image processing, and sentiment analysis to generate meeting minutes. This step takes text data, visual information, and sentiment data as input, and combines and edits them to output detailed meeting minutes based on the flow of the meeting.
[0355] Step 6:
[0356] Users can use a generative AI model to evaluate meeting minutes generated with prompts and provide feedback for improvement. For example, they can instruct the model to emphasize specific emotional responses, resulting in the minutes being highlighted as part of the output.
[0357] (Application Example 2)
[0358] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0359] In brick-and-mortar stores, it is crucial to use customer interest and satisfaction to improve services. However, traditional methods have made it difficult to accurately capture customer emotions and reactions, making it challenging to obtain practical feedback for service improvement.
[0360] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0361] In this invention, the server includes means for acquiring audio recordings, means for performing emotion analysis, integrating emotion numerical information obtained from audio and video, and generating customer response evaluations, and means for estimating customer interest and satisfaction using the emotion numerical information and utilizing that information in commercial activities. This makes it possible to obtain valuable feedback based on customers' real-time reactions.
[0362] "Audio recording" is the process of saving audio as digital data.
[0363] "Speech recognition means" refers to a technical device for converting speech data into text data.
[0364] "Visual materials" refer to digital content, including images and text information, that are displayed during meetings or commercial activities.
[0365] "Image recognition means" refers to a technological device that analyzes and extracts text information and graph information from visual materials.
[0366] "Meeting minutes" refers to information that records the content of meetings or business negotiations in written form.
[0367] "Emotional analysis" is a technology that analyzes emotional responses from audio and video data.
[0368] "Emotional numerical information" refers to emotional state data expressed in specific numerical values, obtained through emotion analysis.
[0369] "Response evaluation" refers to evaluation data generated based on customer emotions and reactions.
[0370] "Interest level" is a quantitative indicator that shows the degree of interest a customer is showing.
[0371] "Customer satisfaction" is an indicator that shows the degree of customer satisfaction with the services or products provided.
[0372] "Commercial activity" refers to all business activities related to the sale of goods or services.
[0373] The system for realizing this invention considers customer service in a physical store as an application example using speech recognition, image recognition, and emotion analysis technologies. The hardware and software configuration of the system will be described below.
[0374] The device uses smart glasses to capture audio and video during customer interactions. The smart glasses' camera records the customer's facial expressions and gestures in real time, and the microphone captures the customer's voice. This audio data is converted into text on a server using a speech recognition engine such as Google Speech Recognition. Speech recognition includes speaker identification to determine which customer is speaking.
[0375] Simultaneously, the server uses the OpenCV library to analyze the captured video data and evaluates the customer's emotional state using an emotion analysis engine called Emotion Recognizer. Emotion analysis extracts numerical emotional information from the customer's voice tone, speed, and facial expressions to calculate levels of interest and satisfaction. This data provides valuable information for understanding customer needs and reactions, with the aim of utilizing it in commercial activities.
[0376] Store employees, who are also users of the smart glasses, can make appropriate product recommendations to customers based on the feedback they receive. By aggregating this kind of information, the overall service quality of the store can be improved.
[0377] For example, if a customer asks a question about a new product, sentiment analysis can instantly determine whether or not the customer is interested. As a result, employees can provide customized service, such as introducing related products that might pique the customer's interest.
[0378] Examples of prompts for a generative AI model are as follows:
[0379] "Based on the conversational text and customer sentiment data detected by the system installed in the smart glasses, please suggest what kind of product recommendations should be made."
[0380] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0381] Step 1:
[0382] The device uses the camera and microphone of smart glasses to capture the customer's video and audio in real time. The input for this step is the customer's face and conversational audio, and the output generates raw video and audio data.
[0383] Step 2:
[0384] The server receives audio data sent from the terminal and converts it into text using the Google Speech Recognition engine. The input is the captured audio data, and the output is text data in which the audio content has been converted into written information. Speaker identification information is also added simultaneously through speech recognition.
[0385] Step 3:
[0386] The server analyzes video data using the OpenCV library. The input is video data sent from the terminal, and the output is numerical emotion information. In this process, Emotion Recognizer is used to quantify the customer's emotional state from the video.
[0387] Step 4:
[0388] The server integrates text data and sentiment numerical information obtained from speech recognition to generate a customer response evaluation. The input is the text data and sentiment numerical information obtained from the previous step, and the output is evaluation data that assesses the customer's level of interest and satisfaction.
[0389] Step 5:
[0390] The server provides feedback to the user regarding the generated response evaluation. The user, i.e., the store employee, then uses this information to immediately make the most appropriate product recommendations to the customer. The input is the customer's response evaluation, and the output is a customized product recommendation tailored to the customer's needs. This feedback process enables employees to provide more effective customer service.
[0391] Step 6:
[0392] The user inputs the generated prompt sentences into the generating AI model to obtain further suggestions and responses. At this stage, the prompt sentences shown previously are used, and further service improvements are made based on the output from the generating AI model.
[0393] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0394] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0395] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0396] [Third Embodiment]
[0397] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0398] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0399] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0400] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0401] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0402] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0403] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0404] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0405] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0406] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0407] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0408] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0409] This invention is implemented as a system for automatically creating highly accurate meeting minutes in online meetings. The system is primarily composed of speech recognition and image recognition means, and enables rapid and accurate meeting minute generation by efficiently processing audio data and visual information.
[0410] During a meeting, the terminal first captures the participants' voices in real time and records them digitally. Next, the server transfers this audio data to a speech recognition engine, which automatically converts it into text. The speech recognition process also includes speaker identification, ensuring clear distinction between speakers.
[0411] Simultaneously, the terminal captures screen images of visual materials such as slides and documents shared during the meeting. These image data are sent to a server, where image recognition algorithms are used to extract text and chart data. This information becomes crucial for later meeting minute generation.
[0412] Next, the server integrates the text data generated by speech recognition with the information extracted by image recognition. In this integration process, duplicate and inconsistent information is removed, and the data is processed to accurately reflect the meeting content. As a result, a well-organized meeting transcript is created, aligned with the timeline.
[0413] Ultimately, the generated meeting minutes are provided to the user and become immediately available for review and sharing right after the meeting ends. For example, in a project progress meeting, the minutes will be generated accurately, including not only the statement "The project is progressing smoothly" but also the data shown on the slide, such as "Progress: 70% complete."
[0414] This system eliminates the need for manual data recording and supports decision-making based on fast, highly accurate data.
[0415] The following describes the processing flow.
[0416] Step 1:
[0417] As soon as the online meeting begins, the device captures participants' voices in real time via the microphone and records them as digital data. This data is then prepared to be transferred to a server over the network.
[0418] Step 2:
[0419] The server inputs the received audio data into a speech recognition engine, performs noise reduction, and converts it into accurate text information. During this process, the speaker is identified, and the content of the speech is organized by speaker.
[0420] Step 3:
[0421] The device automatically takes screenshots of slides and documents shared during the meeting. This process is performed each time a new document is projected, and the captured image data is sent to the server.
[0422] Step 4:
[0423] The server processes the transmitted image data through an image recognition engine and uses OCR to extract text information as text data. It also identifies charts and graphs and analyzes numerical data if necessary.
[0424] Step 5:
[0425] The server performs a process of integrating text data obtained from speech recognition with text and numerical data obtained from image recognition. The data is scrutinized to avoid duplication and inconsistencies, and comprehensive, time-series data is generated.
[0426] Step 6:
[0427] The server organizes the integrated data and generates a final document as meeting minutes. It checks the structure of the document, summarizes it as needed, and formats it in a clear and user-friendly format.
[0428] Step 7:
[0429] Users receive the generated meeting minutes and review their contents. After review, they can share them with all participants as needed, enabling quick feedback and decision-making.
[0430] (Example 1)
[0431] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0432] In online meetings, it is difficult to individually manage participants' audio and visual materials, and there is a need to create meeting minutes accurately and quickly. However, conventional methods require manual recording, resulting in insufficient accuracy and efficiency. Therefore, a system is needed that effectively integrates audio data and visual materials and automatically generates meeting minutes that accurately reflect the content of the meeting.
[0433] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0434] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text information, means for acquiring visual media displayed during a meeting, means for extracting text information and chart information from the visual media, means for integrating the text information obtained by the audio recognition means with the text information and chart information obtained by the image processing means and processing it as meeting minutes information, and means for immediately providing the generated meeting minutes information to users. This makes it possible to quickly and accurately generate meeting minutes from the audio and visual materials of an online meeting and immediately share the information with all participants.
[0435] "Audio data" refers to a digital recording of what meeting participants said.
[0436] "Acoustic recognition means" refers to processing means that analyze acquired audio data and convert it into textual information.
[0437] "Visual media" refers to visual materials such as slides and documents used during a meeting.
[0438] "Image processing means" refers to a processing device for extracting textual information and diagrammatic information from a visual medium.
[0439] "Textual information" refers to string data extracted from audio data or visual media.
[0440] "Graphical information" refers to data in the form of shapes or tables extracted from visual media.
[0441] "Integration" refers to the process of combining text information obtained through speech recognition with text information and chart information obtained through image processing into a single document.
[0442] "Meeting minutes information" refers to information in document form that records the content of a meeting.
[0443] "User" refers to an individual or organization that is able to receive and use the generated meeting minutes information.
[0444] This invention relates to a system for automatically generating accurate meeting minutes during online meetings. This system uses speech recognition and image processing means to extract and integrate information from participants' statements and visual materials, thereby generating meeting minutes that accurately reflect the content of the meeting.
[0445] The terminal captures the voice of each meeting participant in real time via a microphone device and records it as digital audio data. The recorded audio data is transmitted to a server via a communication network. This server has the capability to convert the audio data into text information, for example, using a speech recognition API. In this process, speaker identification is also performed, making it possible to clearly distinguish between speakers.
[0446] Simultaneously, the terminal captures visual media such as slides and documents used during the meeting using screen capture software. This image data is also sent to the server. The server uses image processing algorithms to extract text and diagram information from the visual media. For example, an image recognition service might be used for this process.
[0447] The generated data is integrated by the server, removing duplicates and inconsistencies to create organized meeting minutes. Finally, the generated minutes are provided to the user, allowing for immediate review and sharing immediately after the meeting ends. As a concrete example, in a project progress meeting, the minutes would include both the statement "The project is progressing smoothly" and the data displayed on the slide, such as "Progress: 70% complete."
[0448] Example prompt: "Please explain how to integrate speech recognition and image recognition for creating meeting minutes during online meetings."
[0449] Therefore, this invention is an innovative system that automates conventional manual recording work and can provide meeting minutes that reflect the content of meetings quickly and with high accuracy.
[0450] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0451] Step 1:
[0452] The terminal captures the voices of meeting participants in real time using a microphone device and converts them into digital audio data. The input is a raw audio signal, and the output is digital audio data (e.g., WAV format). This conversion process involves converting from analog to digital signals, specifically using a microphone device and audio capture software.
[0453] Step 2:
[0454] The terminal uses screen capture software to capture visual materials such as slides and documents used during a meeting. The input is the visual material displayed on the terminal, and the output is image data (e.g., PNG format). Specifically, the screen capture button is pressed to save the current screen.
[0455] Step 3:
[0456] The server receives audio data sent from the terminal. Next, it inputs this data into a speech recognition API, converting the audio data into text information. The input is digital audio data, and the output is text data. This process includes breaking down the audio waveform into phonemes and converting those phonemes into strings.
[0457] Step 4:
[0458] The server receives image data sent from the terminal and extracts text and diagram information from the image using an image recognition service. The input is image data, and the output is extracted text and diagram information. The image data is analyzed, and text is extracted using OCR (Optical Character Recognition) technology.
[0459] Step 5:
[0460] The server integrates text information obtained from audio data with information extracted from image data. During this process, it checks for duplicate or inconsistent information and generates meeting minutes in an organized format. Input is text information from speech recognition and image recognition, and output is integrated meeting minutes information. Information is arranged considering its relevance, and unnecessary data is removed.
[0461] Step 6:
[0462] The server immediately provides the generated meeting minutes information to the user. The input is integrated meeting minutes information, and the output is meeting minutes in document format delivered to the user. Specifically, the server converts the meeting minutes data to PDF or DOCX format and sends a download link to the user's device.
[0463] (Application Example 1)
[0464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0465] In online consultations at virtual stores, there is a need to efficiently record customer conversations and integrate them with product data. However, conventional methods often lack consistency and accuracy because audio and visual information are processed separately. To solve this problem, it is necessary to develop a system that integrates customer conversations and product information and outputs them as a record immediately.
[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0467] In this invention, the server includes means for acquiring acoustic information, means for acoustic analysis for converting the acoustic information into text information, and means for acquiring visual data presented during negotiations. This makes it possible to accurately acquire the content of conversations with customers as text information and immediately integrate it with product data.
[0468] "Acoustic information" refers to audio data and signals, and is the information used for audio recording and speech recognition.
[0469] "Acoustic analysis means" refers to technology that interprets acoustic information and converts it into text information. This may include a speech recognition engine.
[0470] "Visual data" refers to information presented as visual materials or images. This may include charts, graphs, and text.
[0471] "Image analysis means" refers to the process of extracting necessary information, such as text and illustrations, from visual data.
[0472] "Integrated processing means" refers to a technology that combines text information obtained by acoustic analysis means with information obtained by image analysis means and processes them in an integrated manner.
[0473] "Recorded information" refers to documents and files created based on analyzed text and image information.
[0474] This invention is a system that acquires and analyzes acoustic information in real time to effectively record the content of consultations with customers. Specifically, the server converts the audio data into text information using a speech recognition engine. The audio data is acquired by the terminal via the microphone of a smartphone or PC. The server further acquires visual data presented during negotiations and extracts text information and diagram data using image analysis means. Image recognition algorithms and OCR technology are used for this analysis. The image information is captured through a camera-equipped terminal and transmitted to the server.
[0475] This data is integrated on the server and output as a consistent consultation record. This ensures that customer conversations and product data are organized and recorded without any loss. For example, if a customer asks online, "How can I use this product?", the voice data is recognized in real time, and usage data is simultaneously extracted from product specifications and images taken by the device and automatically added to the consultation record. By utilizing a generative AI model, this integration process can be performed quickly and efficiently.
[0476] Examples of prompt messages include, "Please summarize the customer's consultation based on the following audio and images." This system can improve the quality of customer service in virtual stores.
[0477] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0478] Step 1:
[0479] The user initiates an online consultation with a customer using a smartphone or computer. The user's device begins acquiring audio data from the customer in real time via the microphone. The input is audio data, which is then sent to the server for subsequent processing.
[0480] Step 2:
[0481] The server receives audio data and converts it into text information using a speech recognition engine. This process utilizes a generative AI model to analyze acoustic features and perform data processing and calculations to convert them into strings. The output is the textual information of the conversation with the customer.
[0482] Step 3:
[0483] The device uses its camera to photograph visual data presented during negotiations, such as product brochures or instruction manuals. The input is image data, which is then sent to the server.
[0484] Step 4:
[0485] The server acquires image data and extracts textual and illustrative information using image analysis tools. Specifically, it uses OCR technology to identify characters in the image and converts them into text using a generative AI model. The output is textual data of the visual information.
[0486] Step 5:
[0487] The server integrates the textual information from speech recognition results and the textual information from image analysis into a single consultation record using an integrated processing mechanism. This process associates the spoken words with the image information in chronological order, maintaining consistency throughout the integration. The output is the final consultation record document.
[0488] Step 6:
[0489] The server provides the user with the generated consultation record. The user reviews this record and makes corrections or additions as needed. They may also send the record to other systems or applications using prompts. For example, a prompt such as, "Please summarize the customer consultation based on the following audio and images," can be used to request the generating AI model to summarize the output data.
[0490] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0491] This invention is implemented as an online meeting system that, in addition to speech recognition and image recognition, utilizes an emotion engine to capture the user's emotional state and reflect it in the meeting minutes. This system generates meeting minutes that not only record information but also take into account the emotional reactions of participants during the meeting.
[0492] During the meeting, the terminal captures participants' audio and video with high precision, and the audio data is recorded in digital format. Simultaneously, the video data is transmitted along with visual materials for analysis by the emotion sensor.
[0493] The server uses a speech recognition engine to convert audio data into text data in real time. This conversion incorporates speaker identification, and the speaker's identity is also recorded. In addition, screen capture data of shared documents is analyzed by an image recognition engine to extract text and graph information from the documents.
[0494] In parallel with these recognition processes, the server analyzes the user's emotions from the audio and video through an emotion engine. The emotion engine analyzes the tone, speed, and body language of the voice, and generates qualitative data on the participant's emotional state (e.g., joy, surprise, confusion, etc.).
[0495] Next, the server integrates the text data with the emotional state generated by the emotion engine and produces a document edited as meeting minutes. During this process, the user's reactions are recorded chronologically based on the content of the conversation, making significant emotional changes during the meeting visible.
[0496] For example, in a new product launch meeting, if a user says, "The new feature is groundbreaking," and the emotion engine detects surprise or excitement on the user's face, that emotional response will be appropriately added to the meeting minutes. As a result, information that takes into account the emotions of the meeting participants is recorded, which can function as useful information for decision-making after the meeting.
[0497] Users can immediately review and utilize meeting minutes containing this sentiment information after the meeting ends, leading to deeper insights into the meeting content and providing higher-quality support.
[0498] The following describes the processing flow.
[0499] Step 1:
[0500] The device captures participants' audio and video data via microphone and camera at the start of the meeting and records it as digital data. This data is then prepared to be sent to the server in real time.
[0501] Step 2:
[0502] The server inputs the received audio data into a speech recognition engine and converts it into text data. During this process, noise reduction and speaker identification are performed simultaneously to clearly record who spoke.
[0503] Step 3:
[0504] The server extracts visual materials displayed during the meeting from the video data and analyzes them using an image recognition engine. Here, it identifies text information and graphs within the materials and extracts the necessary information as text data.
[0505] Step 4:
[0506] The server uses an emotion engine to analyze participants' facial expressions and body language based on video data, identifying their emotional state. This includes changes in voice tone and speed, allowing for a detailed capture of the user's emotions.
[0507] Step 5:
[0508] The server integrates text data obtained from speech recognition, text and graph information obtained from image recognition, and sentiment data obtained from the sentiment engine. This integration process organizes the information chronologically and generates a unified meeting record.
[0509] Step 6:
[0510] The server delivers the generated meeting minutes to the dashboard, making it possible to verify the accuracy of the content. Here, emotional responses to individual statements are visualized and adjusted to make the meeting content easier to understand.
[0511] Step 7:
[0512] Users can review these meeting minutes after the meeting ends, enabling them to efficiently make decisions and take follow-up actions based on detailed meeting analysis information, including sentiment data.
[0513] (Example 2)
[0514] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0515] Traditional online meeting systems could record meeting content as text, but they did not take into account participants' emotional responses. As a result, important emotional shifts and participant reactions during meetings were not utilized in decision-making. Furthermore, it was difficult to comprehensively integrate information from visual materials.
[0516] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0517] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data for speech recognition, means for acquiring visual materials, means for image processing for extracting information from the visual materials, and means for analyzing the emotional state of the participants. This makes it possible to generate a meeting record that integrates audio information, visual material information, and the emotional responses of the participants into a single record.
[0518] "Audio data" refers to information recorded in digital format from what participants said during a meeting.
[0519] "Text data" refers to information in character format converted by speech recognition technology, representing the content of a meeting in written form.
[0520] "Visual materials" refer to visual information such as slides and graphs that are shared during a meeting.
[0521] "Image processing means" refers to techniques for extracting text information and graph information from visual materials.
[0522] "Emotional analysis means" refers to a technology that analyzes participants' emotions from audio and video data and generates qualitative data from it.
[0523] "Integration" is the process of combining audio data, information obtained from visual materials, and emotional data into a single meeting record.
[0524] "Meeting record information" refers to a document that is generated by integrating the content of a meeting with the emotional reactions of the participants at that time.
[0525] This invention is a system that improves the recording and analysis of information in online meetings. Specifically, a terminal uses a microphone and camera to capture participants' audio and video data with high accuracy and transmit it to a server. This system first uses a speech recognition engine to convert the audio data into text data in real time. This speech recognition process includes a speaker identification function, and speaker information is added to each statement.
[0526] In addition, the server acquires visual materials and uses an image processing engine to extract necessary information from the materials. This image processing engine has the ability to efficiently extract and analyze text and graph information from screen captures.
[0527] Furthermore, an emotion analysis engine analyzes participants' emotional states from audio and video data. Specifically, it analyzes voice tone and speed, facial expressions, and body language to qualitatively represent emotions such as joy and surprise.
[0528] This data is integrated and generated as meeting minutes that include emotional information. These minutes include key emotional shifts and highlights of topics during the meeting, helping users accurately understand the flow and key points of the meeting.
[0529] For example, when a user says "This feature is innovative" during a new product announcement, the server detects the user's emotional response, such as surprise or excitement, and records the details in the meeting minutes.
[0530] By using generative AI models, users can perform specific analyses and visualizations using prompts. For example, a prompt such as "Highlight the parts that surprised the participants" allows users to easily identify points of interest.
[0531] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0532] Step 1:
[0533] The terminal uses a microphone and camera at the start of the meeting to capture participants' audio and video data. This process takes real-time audio and video as input and generates audio and video files in digital format as output. These files are immediately transferred to the server.
[0534] Step 2:
[0535] The server converts the audio data received from the terminal into text data using a speech recognition engine. It receives an audio file as input and outputs text data, including speaker identification, through data processing. Specifically, it analyzes phonological patterns and maps each utterance to a specific speaker.
[0536] Step 3:
[0537] The server uses an image processing engine to acquire visual materials and extracts text and graph information from them. Here, screen-captured video data is used as input, and important text elements and graphics are extracted through data calculations and output as structured information.
[0538] Step 4:
[0539] The server uses an emotion analysis engine to analyze participants' emotional states from audio and video data. It receives audio tone and facial expression data from the video as input, and uses this data analysis to generate emotional indicators. The output presents each participant's emotional state as qualitative data.
[0540] Step 5:
[0541] The server integrates data obtained through speech recognition, image processing, and sentiment analysis to generate meeting minutes. This step takes text data, visual information, and sentiment data as input, and combines and edits them to output detailed meeting minutes based on the flow of the meeting.
[0542] Step 6:
[0543] Users can use a generative AI model to evaluate meeting minutes generated with prompts and provide feedback for improvement. For example, they can instruct the model to emphasize specific emotional responses, resulting in the minutes being highlighted as part of the output.
[0544] (Application Example 2)
[0545] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0546] In brick-and-mortar stores, it is crucial to use customer interest and satisfaction to improve services. However, traditional methods have made it difficult to accurately capture customer emotions and reactions, making it challenging to obtain practical feedback for service improvement.
[0547] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0548] In this invention, the server includes means for acquiring audio recordings, means for performing emotion analysis, integrating emotion numerical information obtained from audio and video, and generating customer response evaluations, and means for estimating customer interest and satisfaction using the emotion numerical information and utilizing that information in commercial activities. This makes it possible to obtain valuable feedback based on customers' real-time reactions.
[0549] "Audio recording" is the process of saving audio as digital data.
[0550] "Speech recognition means" refers to a technical device for converting speech data into text data.
[0551] "Visual materials" refer to digital content, including images and text information, that are displayed during meetings or commercial activities.
[0552] "Image recognition means" refers to a technological device that analyzes and extracts text information and graph information from visual materials.
[0553] "Meeting minutes" refers to information that records the content of meetings or business negotiations in written form.
[0554] "Emotional analysis" is a technology that analyzes emotional responses from audio and video data.
[0555] "Emotional numerical information" refers to emotional state data expressed in specific numerical values, obtained through emotion analysis.
[0556] "Response evaluation" refers to evaluation data generated based on customer emotions and reactions.
[0557] "Interest level" is a quantitative indicator that shows the degree of interest a customer is showing.
[0558] "Customer satisfaction" is an indicator that shows the degree of customer satisfaction with the services or products provided.
[0559] "Commercial activity" refers to all business activities related to the sale of goods or services.
[0560] The system for realizing this invention considers customer service in a physical store as an application example using speech recognition, image recognition, and emotion analysis technologies. The hardware and software configuration of the system will be described below.
[0561] The device uses smart glasses to capture audio and video during customer interactions. The smart glasses' camera records the customer's facial expressions and gestures in real time, and the microphone captures the customer's voice. This audio data is converted into text on a server using a speech recognition engine such as Google Speech Recognition. Speech recognition includes speaker identification to determine which customer is speaking.
[0562] Simultaneously, the server uses the OpenCV library to analyze the captured video data and evaluates the customer's emotional state using an emotion analysis engine called Emotion Recognizer. Emotion analysis extracts numerical emotional information from the customer's voice tone, speed, and facial expressions to calculate levels of interest and satisfaction. This data provides valuable information for understanding customer needs and reactions, with the aim of utilizing it in commercial activities.
[0563] Store employees, who are also users of the smart glasses, can make appropriate product recommendations to customers based on the feedback they receive. By aggregating this kind of information, the overall service quality of the store can be improved.
[0564] For example, if a customer asks a question about a new product, sentiment analysis can instantly determine whether or not the customer is interested. As a result, employees can provide customized service, such as introducing related products that might pique the customer's interest.
[0565] Examples of prompts for a generative AI model are as follows:
[0566] "Based on the conversational text and customer sentiment data detected by the system installed in the smart glasses, please suggest what kind of product recommendations should be made."
[0567] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0568] Step 1:
[0569] The device uses the camera and microphone of smart glasses to capture the customer's video and audio in real time. The input for this step is the customer's face and conversational audio, and the output generates raw video and audio data.
[0570] Step 2:
[0571] The server receives audio data sent from the terminal and converts it into text using the Google Speech Recognition engine. The input is the captured audio data, and the output is text data in which the audio content has been converted into written information. Speaker identification information is also added simultaneously through speech recognition.
[0572] Step 3:
[0573] The server analyzes video data using the OpenCV library. The input is video data sent from the terminal, and the output is numerical emotion information. In this process, Emotion Recognizer is used to quantify the customer's emotional state from the video.
[0574] Step 4:
[0575] The server integrates text data and sentiment numerical information obtained from speech recognition to generate a customer response evaluation. The input is the text data and sentiment numerical information obtained from the previous step, and the output is evaluation data that assesses the customer's level of interest and satisfaction.
[0576] Step 5:
[0577] The server provides feedback to the user regarding the generated response evaluation. The user, i.e., the store employee, then uses this information to immediately make the most appropriate product recommendations to the customer. The input is the customer's response evaluation, and the output is a customized product recommendation tailored to the customer's needs. This feedback process enables employees to provide more effective customer service.
[0578] Step 6:
[0579] The user inputs the generated prompt sentences into the generating AI model to obtain further suggestions and responses. At this stage, the prompt sentences shown previously are used, and further service improvements are made based on the output from the generating AI model.
[0580] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0581] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0582] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0583] [Fourth Embodiment]
[0584] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0585] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0586] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0587] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0588] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0589] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0590] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0591] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0592] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0593] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0594] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0595] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0596] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0597] This invention is implemented as a system for automatically creating highly accurate meeting minutes in online meetings. The system is primarily composed of speech recognition and image recognition means, and enables rapid and accurate meeting minute generation by efficiently processing audio data and visual information.
[0598] During a meeting, the terminal first captures the participants' voices in real time and records them digitally. Next, the server transfers this audio data to a speech recognition engine, which automatically converts it into text. The speech recognition process also includes speaker identification, ensuring clear distinction between speakers.
[0599] Simultaneously, the terminal captures screen images of visual materials such as slides and documents shared during the meeting. These image data are sent to a server, where image recognition algorithms are used to extract text and chart data. This information becomes crucial for later meeting minute generation.
[0600] Next, the server integrates the text data generated by speech recognition with the information extracted by image recognition. In this integration process, duplicate and inconsistent information is removed, and the data is processed to accurately reflect the meeting content. As a result, a well-organized meeting transcript is created, aligned with the timeline.
[0601] Ultimately, the generated meeting minutes are provided to the user and become immediately available for review and sharing right after the meeting ends. For example, in a project progress meeting, the minutes will be generated accurately, including not only the statement "The project is progressing smoothly" but also the data shown on the slide, such as "Progress: 70% complete."
[0602] This system eliminates the need for manual data recording and supports decision-making based on fast, highly accurate data.
[0603] The following describes the processing flow.
[0604] Step 1:
[0605] As soon as the online meeting begins, the device captures participants' voices in real time via the microphone and records them as digital data. This data is then prepared to be transferred to a server over the network.
[0606] Step 2:
[0607] The server inputs the received audio data into a speech recognition engine, performs noise reduction, and converts it into accurate text information. During this process, the speaker is identified, and the content of the speech is organized by speaker.
[0608] Step 3:
[0609] The device automatically takes screenshots of slides and documents shared during the meeting. This process is performed each time a new document is projected, and the captured image data is sent to the server.
[0610] Step 4:
[0611] The server processes the transmitted image data through an image recognition engine and uses OCR to extract text information as text data. It also identifies charts and graphs and analyzes numerical data if necessary.
[0612] Step 5:
[0613] The server performs a process of integrating text data obtained from speech recognition with text and numerical data obtained from image recognition. The data is scrutinized to avoid duplication and inconsistencies, and comprehensive, time-series data is generated.
[0614] Step 6:
[0615] The server organizes the integrated data and generates a final document as meeting minutes. It checks the structure of the document, summarizes it as needed, and formats it in a clear and user-friendly format.
[0616] Step 7:
[0617] Users receive the generated meeting minutes and review their contents. After review, they can share them with all participants as needed, enabling quick feedback and decision-making.
[0618] (Example 1)
[0619] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0620] In online meetings, it is difficult to individually manage participants' audio and visual materials, and there is a need to create meeting minutes accurately and quickly. However, conventional methods require manual recording, resulting in insufficient accuracy and efficiency. Therefore, a system is needed that effectively integrates audio data and visual materials and automatically generates meeting minutes that accurately reflect the content of the meeting.
[0621] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0622] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text information, means for acquiring visual media displayed during a meeting, means for extracting text information and chart information from the visual media, means for integrating the text information obtained by the audio recognition means with the text information and chart information obtained by the image processing means and processing it as meeting minutes information, and means for immediately providing the generated meeting minutes information to users. This makes it possible to quickly and accurately generate meeting minutes from the audio and visual materials of an online meeting and immediately share the information with all participants.
[0623] "Audio data" refers to a digital recording of what meeting participants said.
[0624] "Acoustic recognition means" refers to processing means that analyze acquired audio data and convert it into textual information.
[0625] "Visual media" refers to visual materials such as slides and documents used during a meeting.
[0626] "Image processing means" refers to a processing device for extracting textual information and diagrammatic information from a visual medium.
[0627] "Textual information" refers to string data extracted from audio data or visual media.
[0628] "Graphical information" refers to data in the form of shapes or tables extracted from visual media.
[0629] "Integration" refers to the process of combining text information obtained through speech recognition with text information and chart information obtained through image processing into a single document.
[0630] "Meeting minutes information" refers to information in document form that records the content of a meeting.
[0631] "User" refers to an individual or organization that is able to receive and use the generated meeting minutes information.
[0632] This invention relates to a system for automatically generating accurate meeting minutes during online meetings. This system uses speech recognition and image processing means to extract and integrate information from participants' statements and visual materials, thereby generating meeting minutes that accurately reflect the content of the meeting.
[0633] The terminal captures the voice of each meeting participant in real time via a microphone device and records it as digital audio data. The recorded audio data is transmitted to a server via a communication network. This server has the capability to convert the audio data into text information, for example, using a speech recognition API. In this process, speaker identification is also performed, making it possible to clearly distinguish between speakers.
[0634] Simultaneously, the terminal captures visual media such as slides and documents used during the meeting using screen capture software. This image data is also sent to the server. The server uses image processing algorithms to extract text and diagram information from the visual media. For example, an image recognition service might be used for this process.
[0635] The generated data is integrated by the server, removing duplicates and inconsistencies to create organized meeting minutes. Finally, the generated minutes are provided to the user, allowing for immediate review and sharing immediately after the meeting ends. As a concrete example, in a project progress meeting, the minutes would include both the statement "The project is progressing smoothly" and the data displayed on the slide, such as "Progress: 70% complete."
[0636] Example prompt: "Please explain how to integrate speech recognition and image recognition for creating meeting minutes during online meetings."
[0637] Therefore, this invention is an innovative system that automates conventional manual recording work and can provide meeting minutes that reflect the content of meetings quickly and with high accuracy.
[0638] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0639] Step 1:
[0640] The terminal captures the voices of meeting participants in real time using a microphone device and converts them into digital audio data. The input is a raw audio signal, and the output is digital audio data (e.g., WAV format). This conversion process involves converting from analog to digital signals, specifically using a microphone device and audio capture software.
[0641] Step 2:
[0642] The terminal uses screen capture software to capture visual materials such as slides and documents used during a meeting. The input is the visual material displayed on the terminal, and the output is image data (e.g., PNG format). Specifically, the screen capture button is pressed to save the current screen.
[0643] Step 3:
[0644] The server receives audio data sent from the terminal. Next, it inputs this data into a speech recognition API, converting the audio data into text information. The input is digital audio data, and the output is text data. This process includes breaking down the audio waveform into phonemes and converting those phonemes into strings.
[0645] Step 4:
[0646] The server receives image data sent from the terminal and extracts text and diagram information from the image using an image recognition service. The input is image data, and the output is extracted text and diagram information. The image data is analyzed, and text is extracted using OCR (Optical Character Recognition) technology.
[0647] Step 5:
[0648] The server integrates text information obtained from audio data with information extracted from image data. During this process, it checks for duplicate or inconsistent information and generates meeting minutes in an organized format. Input is text information from speech recognition and image recognition, and output is integrated meeting minutes information. Information is arranged considering its relevance, and unnecessary data is removed.
[0649] Step 6:
[0650] The server immediately provides the generated meeting minutes information to the user. The input is integrated meeting minutes information, and the output is meeting minutes in document format delivered to the user. Specifically, the server converts the meeting minutes data to PDF or DOCX format and sends a download link to the user's device.
[0651] (Application Example 1)
[0652] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0653] In online consultations at virtual stores, there is a need to efficiently record customer conversations and integrate them with product data. However, conventional methods often lack consistency and accuracy because audio and visual information are processed separately. To solve this problem, it is necessary to develop a system that integrates customer conversations and product information and outputs them as a record immediately.
[0654] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0655] In this invention, the server includes means for acquiring acoustic information, means for acoustic analysis for converting the acoustic information into text information, and means for acquiring visual data presented during negotiations. This makes it possible to accurately acquire the content of conversations with customers as text information and immediately integrate it with product data.
[0656] "Acoustic information" refers to audio data and signals, and is the information used for audio recording and speech recognition.
[0657] "Acoustic analysis means" refers to technology that interprets acoustic information and converts it into text information. This may include a speech recognition engine.
[0658] "Visual data" refers to information presented as visual materials or images. This may include charts, graphs, and text.
[0659] "Image analysis means" refers to the process of extracting necessary information, such as text and illustrations, from visual data.
[0660] "Integrated processing means" refers to a technology that combines text information obtained by acoustic analysis means with information obtained by image analysis means and processes them in an integrated manner.
[0661] "Recorded information" refers to documents and files created based on analyzed text and image information.
[0662] This invention is a system that acquires and analyzes acoustic information in real time to effectively record the content of consultations with customers. Specifically, the server converts the audio data into text information using a speech recognition engine. The audio data is acquired by the terminal via the microphone of a smartphone or PC. The server further acquires visual data presented during negotiations and extracts text information and diagram data using image analysis means. Image recognition algorithms and OCR technology are used for this analysis. The image information is captured through a camera-equipped terminal and transmitted to the server.
[0663] This data is integrated on the server and output as a consistent consultation record. This ensures that customer conversations and product data are organized and recorded without any loss. For example, if a customer asks online, "How can I use this product?", the voice data is recognized in real time, and usage data is simultaneously extracted from product specifications and images taken by the device and automatically added to the consultation record. By utilizing a generative AI model, this integration process can be performed quickly and efficiently.
[0664] Examples of prompt messages include, "Please summarize the customer's consultation based on the following audio and images." This system can improve the quality of customer service in virtual stores.
[0665] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0666] Step 1:
[0667] The user initiates an online consultation with a customer using a smartphone or computer. The user's device begins acquiring audio data from the customer in real time via the microphone. The input is audio data, which is then sent to the server for subsequent processing.
[0668] Step 2:
[0669] The server receives audio data and converts it into text information using a speech recognition engine. This process utilizes a generative AI model to analyze acoustic features and perform data processing and calculations to convert them into strings. The output is the textual information of the conversation with the customer.
[0670] Step 3:
[0671] The device uses its camera to photograph visual data presented during negotiations, such as product brochures or instruction manuals. The input is image data, which is then sent to the server.
[0672] Step 4:
[0673] The server acquires image data and extracts textual and illustrative information using image analysis tools. Specifically, it uses OCR technology to identify characters in the image and converts them into text using a generative AI model. The output is textual data of the visual information.
[0674] Step 5:
[0675] The server integrates the textual information from speech recognition results and the textual information from image analysis into a single consultation record using an integrated processing mechanism. This process associates the spoken words with the image information in chronological order, maintaining consistency throughout the integration. The output is the final consultation record document.
[0676] Step 6:
[0677] The server provides the user with the generated consultation record. The user reviews this record and makes corrections or additions as needed. They may also send the record to other systems or applications using prompts. For example, a prompt such as, "Please summarize the customer consultation based on the following audio and images," can be used to request the generating AI model to summarize the output data.
[0678] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0679] This invention is implemented as an online meeting system that, in addition to speech recognition and image recognition, utilizes an emotion engine to capture the user's emotional state and reflect it in the meeting minutes. This system generates meeting minutes that not only record information but also take into account the emotional reactions of participants during the meeting.
[0680] During the meeting, the terminal captures participants' audio and video with high precision, and the audio data is recorded in digital format. Simultaneously, the video data is transmitted along with visual materials for analysis by the emotion sensor.
[0681] The server uses a speech recognition engine to convert audio data into text data in real time. This conversion incorporates speaker identification, and the speaker's identity is also recorded. In addition, screen capture data of shared documents is analyzed by an image recognition engine to extract text and graph information from the documents.
[0682] In parallel with these recognition processes, the server analyzes the user's emotions from the audio and video through an emotion engine. The emotion engine analyzes the tone, speed, and body language of the voice, and generates qualitative data on the participant's emotional state (e.g., joy, surprise, confusion, etc.).
[0683] Next, the server integrates the text data with the emotional state generated by the emotion engine and produces a document edited as meeting minutes. During this process, the user's reactions are recorded chronologically based on the content of the conversation, making significant emotional changes during the meeting visible.
[0684] For example, in a new product launch meeting, if a user says, "The new feature is groundbreaking," and the emotion engine detects surprise or excitement on the user's face, that emotional response will be appropriately added to the meeting minutes. As a result, information that takes into account the emotions of the meeting participants is recorded, which can function as useful information for decision-making after the meeting.
[0685] Users can immediately review and utilize meeting minutes containing this sentiment information after the meeting ends, leading to deeper insights into the meeting content and providing higher-quality support.
[0686] The following describes the processing flow.
[0687] Step 1:
[0688] The device captures participants' audio and video data via microphone and camera at the start of the meeting and records it as digital data. This data is then prepared to be sent to the server in real time.
[0689] Step 2:
[0690] The server inputs the received audio data into a speech recognition engine and converts it into text data. During this process, noise reduction and speaker identification are performed simultaneously to clearly record who spoke.
[0691] Step 3:
[0692] The server extracts visual materials displayed during the meeting from the video data and analyzes them using an image recognition engine. Here, it identifies text information and graphs within the materials and extracts the necessary information as text data.
[0693] Step 4:
[0694] The server uses an emotion engine to analyze participants' facial expressions and body language based on video data, identifying their emotional state. This includes changes in voice tone and speed, allowing for a detailed capture of the user's emotions.
[0695] Step 5:
[0696] The server integrates text data obtained from speech recognition, text and graph information obtained from image recognition, and sentiment data obtained from the sentiment engine. This integration process organizes the information chronologically and generates a unified meeting record.
[0697] Step 6:
[0698] The server delivers the generated meeting minutes to the dashboard, making it possible to verify the accuracy of the content. Here, emotional responses to individual statements are visualized and adjusted to make the meeting content easier to understand.
[0699] Step 7:
[0700] Users can review these meeting minutes after the meeting ends, enabling them to efficiently make decisions and take follow-up actions based on detailed meeting analysis information, including sentiment data.
[0701] (Example 2)
[0702] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0703] Traditional online meeting systems could record meeting content as text, but they did not take into account participants' emotional responses. As a result, important emotional shifts and participant reactions during meetings were not utilized in decision-making. Furthermore, it was difficult to comprehensively integrate information from visual materials.
[0704] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0705] In this invention, the server includes means for acquiring audio data, means for converting the audio data into text data for speech recognition, means for acquiring visual materials, means for image processing for extracting information from the visual materials, and means for analyzing the emotional state of the participants. This makes it possible to generate a meeting record that integrates audio information, visual material information, and the emotional responses of the participants into a single record.
[0706] "Audio data" refers to information recorded in digital format from what participants said during a meeting.
[0707] "Text data" refers to information in character format converted by speech recognition technology, representing the content of a meeting in written form.
[0708] "Visual materials" refer to visual information such as slides and graphs that are shared during a meeting.
[0709] "Image processing means" refers to techniques for extracting text information and graph information from visual materials.
[0710] "Emotional analysis means" refers to a technology that analyzes participants' emotions from audio and video data and generates qualitative data from it.
[0711] "Integration" is the process of combining audio data, information obtained from visual materials, and emotional data into a single meeting record.
[0712] "Meeting record information" refers to a document that is generated by integrating the content of a meeting with the emotional reactions of the participants at that time.
[0713] This invention is a system that improves the recording and analysis of information in online meetings. Specifically, a terminal uses a microphone and camera to capture participants' audio and video data with high accuracy and transmit it to a server. This system first uses a speech recognition engine to convert the audio data into text data in real time. This speech recognition process includes a speaker identification function, and speaker information is added to each statement.
[0714] In addition, the server acquires visual materials and uses an image processing engine to extract necessary information from the materials. This image processing engine has the ability to efficiently extract and analyze text and graph information from screen captures.
[0715] Furthermore, an emotion analysis engine analyzes participants' emotional states from audio and video data. Specifically, it analyzes voice tone and speed, facial expressions, and body language to qualitatively represent emotions such as joy and surprise.
[0716] This data is integrated and generated as meeting minutes that include emotional information. These minutes include key emotional shifts and highlights of topics during the meeting, helping users accurately understand the flow and key points of the meeting.
[0717] For example, when a user says "This feature is innovative" during a new product announcement, the server detects the user's emotional response, such as surprise or excitement, and records the details in the meeting minutes.
[0718] By using generative AI models, users can perform specific analyses and visualizations using prompts. For example, a prompt such as "Highlight the parts that surprised the participants" allows users to easily identify points of interest.
[0719] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0720] Step 1:
[0721] The terminal uses a microphone and camera at the start of the meeting to capture participants' audio and video data. This process takes real-time audio and video as input and generates audio and video files in digital format as output. These files are immediately transferred to the server.
[0722] Step 2:
[0723] The server converts the audio data received from the terminal into text data using a speech recognition engine. It receives an audio file as input and outputs text data, including speaker identification, through data processing. Specifically, it analyzes phonological patterns and maps each utterance to a specific speaker.
[0724] Step 3:
[0725] The server uses an image processing engine to acquire visual materials and extracts text and graph information from them. Here, screen-captured video data is used as input, and important text elements and graphics are extracted through data calculations and output as structured information.
[0726] Step 4:
[0727] The server uses an emotion analysis engine to analyze participants' emotional states from audio and video data. It receives audio tone and facial expression data from the video as input, and uses this data analysis to generate emotional indicators. The output presents each participant's emotional state as qualitative data.
[0728] Step 5:
[0729] The server integrates data obtained through speech recognition, image processing, and sentiment analysis to generate meeting minutes. This step takes text data, visual information, and sentiment data as input, and combines and edits them to output detailed meeting minutes based on the flow of the meeting.
[0730] Step 6:
[0731] Users can use a generative AI model to evaluate meeting minutes generated with prompts and provide feedback for improvement. For example, they can instruct the model to emphasize specific emotional responses, resulting in the minutes being highlighted as part of the output.
[0732] (Application Example 2)
[0733] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0734] In brick-and-mortar stores, it is crucial to use customer interest and satisfaction to improve services. However, traditional methods have made it difficult to accurately capture customer emotions and reactions, making it challenging to obtain practical feedback for service improvement.
[0735] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0736] In this invention, the server includes means for acquiring audio recordings, means for performing emotion analysis, integrating emotion numerical information obtained from audio and video, and generating customer response evaluations, and means for estimating customer interest and satisfaction using the emotion numerical information and utilizing that information in commercial activities. This makes it possible to obtain valuable feedback based on customers' real-time reactions.
[0737] "Audio recording" is the process of saving audio as digital data.
[0738] "Speech recognition means" refers to a technical device for converting speech data into text data.
[0739] "Visual materials" refer to digital content, including images and text information, that are displayed during meetings or commercial activities.
[0740] "Image recognition means" refers to a technological device that analyzes and extracts text information and graph information from visual materials.
[0741] "Meeting minutes" refers to information that records the content of meetings or business negotiations in written form.
[0742] "Emotional analysis" is a technology that analyzes emotional responses from audio and video data.
[0743] "Emotional numerical information" refers to emotional state data expressed in specific numerical values, obtained through emotion analysis.
[0744] "Response evaluation" refers to evaluation data generated based on customer emotions and reactions.
[0745] "Interest level" is a quantitative indicator that shows the degree of interest a customer is showing.
[0746] "Customer satisfaction" is an indicator that shows the degree of customer satisfaction with the services or products provided.
[0747] "Commercial activity" refers to all business activities related to the sale of goods or services.
[0748] The system for realizing this invention considers customer service in a physical store as an application example using speech recognition, image recognition, and emotion analysis technologies. The hardware and software configuration of the system will be described below.
[0749] The device uses smart glasses to capture audio and video during customer interactions. The smart glasses' camera records the customer's facial expressions and gestures in real time, and the microphone captures the customer's voice. This audio data is converted into text on a server using a speech recognition engine such as Google Speech Recognition. Speech recognition includes speaker identification to determine which customer is speaking.
[0750] Simultaneously, the server uses the OpenCV library to analyze the captured video data and evaluates the customer's emotional state using an emotion analysis engine called Emotion Recognizer. Emotion analysis extracts numerical emotional information from the customer's voice tone, speed, and facial expressions to calculate levels of interest and satisfaction. This data provides valuable information for understanding customer needs and reactions, with the aim of utilizing it in commercial activities.
[0751] Store employees, who are also users of the smart glasses, can make appropriate product recommendations to customers based on the feedback they receive. By aggregating this kind of information, the overall service quality of the store can be improved.
[0752] For example, if a customer asks a question about a new product, sentiment analysis can instantly determine whether or not the customer is interested. As a result, employees can provide customized service, such as introducing related products that might pique the customer's interest.
[0753] Examples of prompts for a generative AI model are as follows:
[0754] "Based on the conversational text and customer sentiment data detected by the system installed in the smart glasses, please suggest what kind of product recommendations should be made."
[0755] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0756] Step 1:
[0757] The device uses the camera and microphone of smart glasses to capture the customer's video and audio in real time. The input for this step is the customer's face and conversational audio, and the output generates raw video and audio data.
[0758] Step 2:
[0759] The server receives audio data sent from the terminal and converts it into text using the Google Speech Recognition engine. The input is the captured audio data, and the output is text data in which the audio content has been converted into written information. Speaker identification information is also added simultaneously through speech recognition.
[0760] Step 3:
[0761] The server analyzes video data using the OpenCV library. The input is video data sent from the terminal, and the output is numerical emotion information. In this process, Emotion Recognizer is used to quantify the customer's emotional state from the video.
[0762] Step 4:
[0763] The server integrates text data and sentiment numerical information obtained from speech recognition to generate a customer response evaluation. The input is the text data and sentiment numerical information obtained from the previous step, and the output is evaluation data that assesses the customer's level of interest and satisfaction.
[0764] Step 5:
[0765] The server provides feedback to the user regarding the generated response evaluation. The user, i.e., the store employee, then uses this information to immediately make the most appropriate product recommendations to the customer. The input is the customer's response evaluation, and the output is a customized product recommendation tailored to the customer's needs. This feedback process enables employees to provide more effective customer service.
[0766] Step 6:
[0767] The user inputs the generated prompt sentences into the generating AI model to obtain further suggestions and responses. At this stage, the prompt sentences shown previously are used, and further service improvements are made based on the output from the generating AI model.
[0768] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0769] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0770] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0771] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0772] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0773] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0774] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0775] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0776] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0777] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0778] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0779] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0780] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0781] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0782] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0783] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0784] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0785] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0786] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0787] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0788] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0789] The following is further disclosed regarding the embodiments described above.
[0790] (Claim 1)
[0791] Means for acquiring audio recordings,
[0792] A speech recognition means for converting the audio recording into text information,
[0793] A means of capturing visual materials displayed during a meeting,
[0794] Image recognition means for extracting text information and graph information from the visual material,
[0795] A means for integrating text information obtained by the speech recognition means with text information and graph information obtained by the image recognition means and outputting it as meeting minutes information,
[0796] A system that includes this.
[0797] (Claim 2)
[0798] The system according to claim 1, characterized in that the speech recognition means identifies the speaker.
[0799] (Claim 3)
[0800] The system according to claim 1, characterized in that the image recognition means detects updates to visual material and performs capture.
[0801] "Example 1"
[0802] (Claim 1)
[0803] Means for acquiring audio data,
[0804] A sound recognition means for converting the audio data into text information,
[0805] Means of acquiring visual media displayed during a meeting,
[0806] Image processing means for extracting textual information and diagrammatic information from the visual medium,
[0807] A means for integrating character information obtained by the aforementioned acoustic recognition means with character information and chart information obtained by the aforementioned image processing means and processing it as meeting minutes information,
[0808] A means of immediately providing the generated meeting minutes information to users,
[0809] A system that includes this.
[0810] (Claim 2)
[0811] The system according to claim 1, characterized in that the sound recognition means identifies the speaker.
[0812] (Claim 3)
[0813] The system according to claim 1, characterized in that the image processing means detects and acquires updates to a visual medium.
[0814] "Application Example 1"
[0815] (Claim 1)
[0816] Means for acquiring acoustic information,
[0817] An acoustic analysis means for converting the acoustic information into textual information,
[0818] Means of obtaining visual data presented during negotiations,
[0819] Image analysis means for extracting textual information and illustrative information from the visual data,
[0820] A means for integrating the character information obtained by the acoustic analysis means and the character information and illustration information obtained by the image analysis means and outputting it as recorded information,
[0821] An integrated processing means that analyzes customer interactions in real time and outputs integrated product data,
[0822] A system that includes this.
[0823] (Claim 2)
[0824] The system according to claim 1, characterized in that the acoustic analysis means identifies the speaker and records the content of the conversation with the customer.
[0825] (Claim 3)
[0826] The system according to claim 1, characterized in that the image analysis means detects and acquires updates to visual data and includes product information in the recorded information.
[0827] "Example 2 of combining an emotion engine"
[0828] (Claim 1)
[0829] Means for acquiring audio data,
[0830] A speech recognition means for converting the audio data into text data,
[0831] Means of acquiring visual information,
[0832] Image processing means for extracting information from the visual material,
[0833] A means of analyzing the emotional state of participants,
[0834] A means for integrating text data from the speech recognition means, information from the image processing means, and emotion data from the emotion analysis means, and outputting it as meeting record information,
[0835] A system that includes this.
[0836] (Claim 2)
[0837] The system according to claim 1, characterized in that the speech recognition means identifies the speaker.
[0838] (Claim 3)
[0839] The system according to claim 1, characterized in that the emotion analysis means analyzes the characteristics of sound and the characteristics of video to generate an emotional state as qualitative data.
[0840] "Application example 2 when combining with an emotional engine"
[0841] (Claim 1)
[0842] Means for acquiring audio recordings,
[0843] A speech recognition means for converting the audio recording into text information,
[0844] A means of capturing visual materials displayed during a meeting,
[0845] Image recognition means for extracting text information and graph information from the visual material,
[0846] A means for integrating text information obtained by the speech recognition means with text information and graph information obtained by the image recognition means and outputting it as meeting minutes information,
[0847] A means for performing emotion analysis, integrating emotional numerical information obtained from audio and video, and generating customer response evaluations,
[0848] A system that includes this.
[0849] (Claim 2)
[0850] The system according to claim 1, characterized in that the speech recognition means identifies the speaker.
[0851] (Claim 3)
[0852] The system according to claim 1, characterized in that the image recognition means detects updates to visual material and performs capture.
[0853] (Claim 4)
[0854] The system according to claim 1, characterized in that it estimates the level of customer interest and satisfaction using the aforementioned emotional numerical information and utilizes that information in commercial activities. [Explanation of symbols]
[0855] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring audio recordings, A speech recognition means for converting the audio recording into text information, A means of capturing visual materials displayed during a meeting, Image recognition means for extracting text information and graph information from the visual material, A means for integrating text information obtained by the speech recognition means with text information and graph information obtained by the image recognition means and outputting it as meeting minutes information, A system that includes this.
2. The system according to claim 1, characterized in that the speech recognition means identifies the speaker.
3. The system according to claim 1, characterized in that the image recognition means detects updates to visual material and performs capture.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A