system

The system addresses the challenges of diverse dialects and languages in meetings by converting audio to text, standardizing dialects, and identifying speakers, enhancing communication and efficiency in meeting minute creation.

JP2026069183APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately transcribing spoken content, identifying speakers, and standardizing language in meetings and business negotiations with diverse dialects and multiple languages, leading to communication losses and inefficient meeting minute creation.

Method used

A system comprising terminal means for audio data collection, speech recognition for text conversion, language conversion for standardization, voiceprint recognition for speaker identification, recording for organized minute generation, and output for real-time viewing and editing, which facilitates real-time transcription, dialect standardization, and speaker-specific minute creation.

Benefits of technology

The system reduces communication losses and significantly decreases the effort required for creating meeting minutes by transcribing audio into text, standardizing dialects, and identifying speakers, enabling efficient and accurate minute generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069183000001_ABST
    Figure 2026069183000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A terminal means for receiving audio data, A speech recognition means that converts received audio data into text data, A language conversion means for converting converted text data into a standard language, A voiceprint recognition means for identifying the speaker of text data, A recording means for generating minutes for each identified speaker, An output method for outputting the generated meeting minutes, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In meetings and business negotiations, it is required to reduce communication losses in situations where different dialects and multiple languages coexist, and to reduce the time and effort required for preparing meeting minutes. Also, in the context of the increasing opportunities for online meetings due to the spread of remote work, there is a need for means to accurately record the content of each speaker's speech and efficiently generate meeting minutes.

Means for Solving the Problems

[0005] The present invention provides a system comprising: terminal means for receiving audio data; speech recognition means for converting the received audio data into text data; language conversion means for converting the converted text data into standard Japanese; voiceprint recognition means for identifying the speaker of the text data; recording means for generating meeting minutes for each identified speaker; and output means for outputting the generated meeting minutes. This makes it possible to transcribe audio from meetings and business negotiations into text in real time and automatically convert dialects into standard Japanese. Furthermore, by utilizing voiceprint recognition technology, it is possible to identify speakers and organize and generate meeting minutes individually. As a result, communication loss can be reduced while significantly reducing the effort required to create meeting minutes.

[0006] "Audio data" refers to information that represents audio in a digital format.

[0007] "Terminal means" refers to devices or equipment used to collect audio data and transmit it to a server or other system for processing.

[0008] "Speech recognition means" refers to technologies and devices that convert speech data into text data.

[0009] "Text data" refers to digital data that can be handled by a computer as character information.

[0010] "Language conversion means" refers to technologies and devices that convert text data expressed in a specific dialect or multiple languages ​​into a standard language or another language.

[0011] "Voiceprint recognition means" refers to technologies and devices that analyze the characteristics of a speaker contained in audio data to identify an individual.

[0012] "Recording means" refers to technologies and devices for saving converted text data in a predetermined format so that it can be used later.

[0013] "Output means" refers to technologies and devices that display generated meeting minutes or other text data to the user or transmit it to other devices.

[0014] "Meeting minutes" are documents or records that document the content of meetings or business negotiations so that they can be referenced later. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13]It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotional engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotional engine is combined.

Embodiment for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention provides a speech recognition system that facilitates communication and improves the efficiency of meeting minute creation in meetings and business negotiations where diverse dialects and multiple languages ​​are present. This system receives speech data, converts it to text in real time, and automatically converts it to standard Japanese as needed. Furthermore, it can identify speakers using voiceprint recognition technology, organize the text for each speaker, and generate meeting minutes.

[0037] Specifically, the terminal collects the voices of participants during meetings and business negotiations. The collected voice data is transmitted to a server via the network. The server converts the received voice data into text using a speech recognition engine. During this process, it refers to a dialect dictionary to replace any dialects in the converted text with standard Japanese.

[0038] Furthermore, the server utilizes voiceprint recognition technology to identify each speaker from the audio data. Using a pre-registered voiceprint database, it can identify individual statements. This makes it clear who said what during an audio conference and allows for the recording of meeting minutes organized by speaker.

[0039] Users can view meeting minutes provided by the server in real time. The minutes include summaries and key action items, and can be edited as needed or shared with other participants. This system reduces communication loss among remote meeting participants and enables effective minute-taking and meeting management.

[0040] As a concrete example, consider a project meeting with a team that includes members from different regions. During the meeting, the terminals simultaneously transmit the participants' voices to the server, which immediately converts the voices into text and standardizes it to a uniform language. For example, if person A, who is participating from Tohoku, says "What should we do today?", this is translated as "What shall we do today?" and displayed in the same format as the other members. In this process, person A's voiceprint is used to automatically classify the statement as A's and reflect it in the meeting minutes. As a result, everyone has a common understanding, and at the end of the meeting, they can obtain an action plan along with detailed meeting minutes.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device captures participants' voices from the microphone during meetings and business negotiations, and converts the audio data into a digital format. The converted digital audio data is then transmitted to the server in real time.

[0044] Step 2:

[0045] The server sequentially processes the audio data received from the terminal and prepares it for input to the speech recognition engine. The server sends the audio data to the processing unit and performs appropriate batch division.

[0046] Step 3:

[0047] The server starts the speech recognition engine and converts speech data into text in real time. Based on the acoustic model and language model, it sequentially converts speech segments into text data.

[0048] Step 4:

[0049] The server applies a dialect dictionary to the text data and automatically converts local dialect expressions to standard Japanese. It detects the relevant dialects in the text and replaces them with standard expressions.

[0050] Step 5:

[0051] The server compares the audio data with a pre-registered voiceprint database and identifies the speaker using voiceprint recognition technology. The server organizes the text data for each speaker and adds information about the speaker.

[0052] Step 6:

[0053] The server generates meeting minutes based on text data organized by speaker. The server arranges the text chronologically and constructs meeting minutes that reflect the content of the meeting.

[0054] Step 7:

[0055] The server extracts summaries and action items from the generated meeting minutes and prepares them for the user. It uses natural language processing to extract and list important information.

[0056] Step 8:

[0057] Users can view meeting minutes provided by the server in real time and make corrections or edits as needed. Users can then share the final meeting minutes with relevant parties.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] In meetings and business negotiations involving diverse dialects and multiple languages, there is a need to create meeting minutes quickly and accurately while maintaining effective communication. Conventional technologies have challenges in accurately transcribing spoken content, identifying speakers, and standardizing language, leading to decreased meeting efficiency.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes an information processing device for receiving audio information, an audio analysis device for converting the received audio information into symbolic information, a language conversion function device for converting the converted symbolic information into a reference language, a biometric information recognition device for identifying the speaker of the symbolic information, a record generation device for generating meeting minutes for each identified speaker, and an output function device for outputting the generated meeting minutes. This enables accurate and efficient creation of meeting minutes even among participants with diverse linguistic backgrounds.

[0063] An "information processing device" is a device used to receive audio information and process the data.

[0064] "Speech analysis means" refers to a technology or function for converting speech information into symbolic information.

[0065] "Language conversion function means" refers to a technology or function for converting converted symbolic information into a reference language.

[0066] "Biometric information recognition means" refers to a technology or function that analyzes biometric information, such as voiceprints, used to identify the speaker of symbolic information.

[0067] "Record generation means" refers to a technology or function for generating meeting minutes for each identified speaker.

[0068] "Output function means" refers to a technology or function for outputting the generated meeting minutes.

[0069] The system of this invention is designed for use in meetings and business negotiations, enabling smooth communication among participants with diverse backgrounds. For implementation, a mechanism is needed to efficiently handle the process from audio information collection to meeting minute generation.

[0070] The terminal acts as hardware for collecting participants' voices in the meeting room. It is equipped with microphones and voice capture software to acquire high-quality audio from multiple participants. This terminal converts the audio into digital data in real time and transmits it to a server via the network.

[0071] The server receives audio information and converts the audio data into symbolic information using speech analysis tools. In this process, existing speech recognition APIs such as Google® Speech-to-Text and IBM Watson® are utilized as speech recognition engines. The server further uses language conversion tools to accurately convert dialects and different languages ​​into a standard language. Here, dialect dictionaries and language conversion algorithms are used to ensure consistency in spoken content.

[0072] The server analyzes the characteristics of the voice data using biometric recognition means and identifies the speaker based on pre-registered voiceprints. This allows for accurate organization of statements made during the meeting, and a record generation means is used to generate meeting minutes for each identified speaker.

[0073] The generated meeting minutes can be viewed, edited, and shared in real time by users using the output function. Users can easily manage the generated meeting minutes through a dedicated interface.

[0074] As a concrete example, consider a project team meeting. If a participant from a regional area says, "What should we do today?", the server translates this to "What shall we do today?". Furthermore, it quickly identifies which participant made a particular statement and reflects this in the meeting minutes. Using this system, everyone can have a shared understanding while efficiently creating meeting minutes.

[0075] An example of a prompt using a generative AI model is: "Convert the audio data spoken during the meeting into text in real time, identify the speaker, and automatically generate meeting minutes."

[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0077] Step 1:

[0078] The device uses a microphone to collect participants' voices during meetings and business negotiations, converting analog audio into digital audio data. The input is the raw voice of each participant, and the output is prepared as digital audio data. This data is transmitted in real time to a server via the network.

[0079] Step 2:

[0080] The server converts digital audio data received from the terminal into symbolic information using speech analysis tools. Digital audio data is passed to the server as input, and speech recognition APIs such as Google Speech-to-Text are used to convert the audio into text data. The output is text data in string format.

[0081] Step 3:

[0082] The server uses language conversion functionality to convert the converted text data back to the base language. Here, text data is used as input, and specific dialects and expressions are replaced with those in the base language. Standardized text data is generated as output.

[0083] Step 4:

[0084] The server identifies the speaker from the audio data using biometric recognition means. In this process, audio feature data is used as input, and the speaker is identified by matching it with a pre-registered voiceprint database. The output is speaker identification information accompanying the text data.

[0085] Step 5:

[0086] The server uses a recording generation mechanism to generate meeting minutes for each identified speaker. Standardized text data and speaker information are used as input data, which are combined to generate organized meeting minutes. The output is a completed meeting minute.

[0087] Step 6:

[0088] Users can view meeting minutes provided by the server in real time through an output function. The input is meeting minutes data from the server, which users can edit and share with other participants via a web application or dedicated software. The output is the meeting minutes and related information displayed in a user-friendly format.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] In commercial environments where multiple languages ​​or regional dialects are spoken, there is a need to facilitate communication between customers and staff. Traditionally, insufficient communication due to different languages ​​or dialects has led to problems with decreased customer satisfaction.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes terminal means for receiving acoustic data, speech recognition means for converting the received acoustic data into text data, and display means for visually outputting the translated text data to a display device. This enables real-time assistance for voice communication with customers and facilitates smooth mutual understanding in a multilingual environment.

[0094] "Terminal means for receiving acoustic data" refers to a device that acquires sound as digital information.

[0095] "Speech recognition means" refers to technology that analyzes acoustic data and converts it into text information.

[0096] "Language conversion means" refers to technologies for converting text information from another language or dialect into a standard language.

[0097] "Voiceprint recognition means" refers to technology for identifying individual speakers from their voices.

[0098] A "recording means" is a function that stores the converted data and saves it as a record.

[0099] "Output means" refers to a device that presents the created record in physical or digital format.

[0100] "Display means" refers to a device that visually displays translated text information.

[0101] "Communication methods" refer to technologies for exchanging data between devices.

[0102] The system implementing this invention is equipped with technology for collecting audio data and processing the information in real time. The terminal primarily receives audio data and transmits it to the server. The server converts the received audio data into text data using speech recognition technology. The converted text data is then translated into a standard language by a language conversion means. Advanced natural language processing technology is used for this translation.

[0103] The server uses voiceprint recognition technology to identify the speaker. Specifically, it organizes each utterance by speaker by comparing it with a pre-registered voiceprint database. This generates a record categorized by speaker. The generated record is provided to the user through an output device.

[0104] This system also includes a display mechanism, allowing translated text data to be visually output. In a real-world example, when using smart glasses to serve foreign customers in a shop, the customer's speech is translated into standard Japanese in real time and displayed on the glasses' screen. This enables smooth communication between staff and customers.

[0105] For example, if a French-speaking tourist visits a souvenir shop and asks about the price of an item, the staff member's smart glasses will display the translated text, "This keychain costs 500 yen." Through this concrete example, users can obtain information in real time and receive a quick response. This system is an effective means of improving customer satisfaction.

[0106] Examples of prompts for a generative AI model include the following:

[0107] "Develop a real-time voice translation system using smart glasses for use in physical stores. The system will collect voice data via a microphone, send it to a server, perform voice recognition and translation, and display the results on the glasses' display. The goal is smooth communication with customers."

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The terminal receives audio data in a commercial environment. Specifically, it collects customer speech via microphones within a store and prepares the data in digital format. The input is an audio signal, and the output is digital audio data.

[0111] Step 2:

[0112] Digital audio data is transferred from the terminal to the server. The server inputs the received audio data into the speech recognition system. The input here is digital audio data, and the server prepares to process that data.

[0113] Step 3:

[0114] The server uses speech recognition technology to convert audio data into text data. Specifically, it uses the Google Cloud Speech-to-Text API to analyze speech and generate text information. The input is digital audio data, and the output is text data.

[0115] Step 4:

[0116] The server uses language conversion tools to translate text data into a reference language. It uses the Google Cloud Translation API to convert input text into the reference language. The input is the converted text data, and the output is the text data in the reference language.

[0117] Step 5:

[0118] The server uses voiceprint recognition technology to identify the speaker. It compares the received audio data with a pre-registered voiceprint database to identify the speaker. The input is the audio data and the voiceprint database, and the output is information about the identified speaker.

[0119] Step 6:

[0120] The server sends the translated text data to a display device via an output mechanism, and outputs it visually. Specifically, the text data is displayed on the smart glasses' screen to provide information to staff. The input is text data in a reference language, and the output is a visual display.

[0121] Step 7:

[0122] The user reviews the displayed information and responds to the customer. Based on the customer's questions, they understand the text display and provide appropriate information. Input is visual display information, and output is the customer interaction conversation.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention combines a speech recognition system that transcribes participants' speech in real time and converts it into standard Japanese during meetings and business negotiations with a newly developed emotion engine. This makes it possible to recognize the emotional state of participants and organize and record the flow of discussion and important points while considering the emotional context.

[0125] This system first has a terminal collect the voices of multiple participants through a microphone and send the converted audio data to a server. The server then uses a speech recognition engine to convert this audio data into text data. During this process, the server uses a dialect dictionary to perform standard language conversion appropriate for the text data.

[0126] Furthermore, the server uses voiceprint recognition technology to identify speakers and organizes each statement by speaker. By matching it with a voiceprint database, speakers can be identified, and this is used to organize the statements by speaker.

[0127] In addition, an emotion engine installed on the server estimates the emotional state of each speaker through text data and speech analysis. The emotion engine adds this emotional information to the meeting minutes and further prioritizes the importance of summaries and action items based on the intensity of the emotions.

[0128] Users can access meeting minutes provided by the server, which include details of each speaker's remarks along with their estimated emotional state. For example, if participant A says, "I have some concerns about this new proposal," during a meeting, the server recognizes A's emotional state from the audio and text, and this emotional information is included in the meeting minutes. This has the advantage of providing the meeting facilitator with information to plan discussions and follow-ups to address A's concerns.

[0129] Thus, the present invention generates meeting minutes that also take emotional nuances into account, improving the quality of meetings and effectively preventing misunderstandings and communication breakdowns among participants.

[0130] The following describes the processing flow.

[0131] Step 1:

[0132] The terminal captures audio from meetings and business negotiations via its microphone and converts it into digital audio data in real time. This audio data is immediately transmitted to the server.

[0133] Step 2:

[0134] The server prepares the audio data received from the terminal for input into the speech recognition engine. The data is processed in batches and converted into text by the speech recognition engine.

[0135] Step 3:

[0136] The server uses a dialect dictionary to convert the text data obtained through speech recognition into standard Japanese. By replacing dialects and expressions with standard Japanese, the server facilitates overall comprehension.

[0137] Step 4:

[0138] The server uses voiceprint recognition technology to identify the individual speaker. It compares features extracted from the audio data with a voiceprint database and adds speaker information to the text data.

[0139] Step 5:

[0140] The server uses an emotion engine to estimate the speaker's emotions from text data and voice features. The emotion engine identifies specific emotional states and associates that information with the text.

[0141] Step 6:

[0142] The server generates meeting minutes by adding sentiment information to text data organized by speaker. The minutes include detailed records of what was said, who said it, and the estimated sentiment.

[0143] Step 7:

[0144] The server extracts summaries and action items from the generated meeting minutes and prioritizes them based on sentiment information. This information is then highlighted within the meeting minutes.

[0145] Step 8:

[0146] Users can view meeting minutes provided by the server in real time. While reviewing the minutes, users can understand the flow of the discussion by referring to sentiment information and take necessary actions.

[0147] (Example 2)

[0148] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0149] In meetings and business negotiations, it is necessary not only to accurately record participants' statements in real time, but also to understand the emotional state of the speakers and organize and record the flow of the discussion and important points while taking emotional context into account. However, conventional speech recognition systems do not adequately identify emotions or convert dialects to standard Japanese, leading to problems such as misunderstandings and communication losses among participants.

[0150] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0151] In this invention, the server includes speech recognition means for converting speech information into text information, language conversion means for converting the converted text information into a common language, and emotion prediction means for analyzing the text information and speech information to predict emotional states. This makes it possible to record the content of meetings while taking into account emotional nuances, thereby facilitating smooth communication among participants.

[0152] "Audio information" refers to sound data collected from participants in meetings or business negotiations using terminal devices.

[0153] "Textual information" refers to digital text data converted from audio information by speech recognition technology.

[0154] "Common language" refers to a standardized form of language expression that does not include dialects.

[0155] "Identification means" refers to technology that identifies a speaker by analyzing the speaker's voice characteristics from textual information.

[0156] "Emotion prediction means" refers to technology that predicts and identifies the emotional state of a speaker by analyzing textual and auditory information.

[0157] "Recording means" refers to a means of organizing the content of a conversation for each identified speaker and generating a meeting record.

[0158] A "dialect dictionary" refers to a language database used to translate regionally specific linguistic expressions into standard Japanese.

[0159] One embodiment of the present invention is a system that records the statements of participants in meetings and business negotiations in real time and generates meeting minutes that take into account emotional context. This system handles everything from collecting audio data to sentiment analysis and generating meeting minutes in an integrated manner.

[0160] The device uses a microphone to collect audio and converts participants' speech into digital audio data. The audio data collected by the device is immediately transmitted to a server via the network. Noise reduction technology is employed during this process to maintain audio quality.

[0161] The server first converts the received audio data into text data using a speech recognition engine. A speech recognition API is used for this process, generating highly accurate textual information. Furthermore, the converted textual information is standardized to a common language through a language conversion function. The software used in this process incorporates a dialect dictionary.

[0162] Next, the server uses an identification mechanism to identify the speaker from the text information. This mechanism uses voiceprint data to identify the speaker and organizes the statements by speaker. Furthermore, the emotion prediction engine on the server analyzes the text data and voice characteristics to predict the speaker's emotional state.

[0163] Ultimately, users can access the meeting minutes generated by the server and review the details of each statement, including emotional nuances. For example, if participant A says, "I have some concerns about this new proposal," in order to ensure the meeting runs smoothly, the server can reflect A's concern in the minutes. This allows users to assess concerns about the proposal and use that information to plan follow-up.

[0164] A possible prompt would be, "What methods and technologies should be used to create a system that records participants' statements and emotions in real time during a meeting and generates meeting minutes based on the estimated emotions?" This system would improve communication among participants and enhance the quality of the meeting.

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The terminal uses a microphone to collect speech from meeting participants in real time. The input is an analog signal of raw sound, which is converted into digital audio data. The converted audio data is then transmitted to a server via the network. Noise reduction technology is used to improve the quality of the audio data.

[0168] Step 2:

[0169] The server converts digital audio data received from the terminal into text information using a speech recognition engine. The input is digital audio data, and the output is text information. Speech recognition automatically transcribes spoken content into text, generating highly accurate text data.

[0170] Step 3:

[0171] The server applies a language conversion function to convert the converted character information into the standard language. The input is text data, and the output is character information in the standard language. By using a pre-built dialect dictionary, it performs the task of standardizing regional expressions.

[0172] Step 4:

[0173] The server identifies the speaker using identification methods based on the text information converted to a common language. The input here is the text information in a common language, and the output is the identification information of each speaker. Voiceprint data is used to identify speakers with high accuracy and organize statements by speaker.

[0174] Step 5:

[0175] The server uses an emotion prediction engine to analyze the speaker's emotional state based on identified textual information. The input is standard language textual information with speaker information, and the output is text information including emotional state. It analyzes the features of the audio and text data to identify the emotional context.

[0176] Step 6:

[0177] Users access meeting minutes generated from the server, which include sentiment information. The generated minutes provide details of each statement, incorporating emotional nuances. This makes it easier for users to effectively manage meeting content and plan subsequent actions.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0180] In brick-and-mortar retail settings, understanding customer emotions and providing accurate information and suggestions has been difficult with conventional technologies. In particular, there was a lack of mechanisms to quickly detect and address customer anxiety or dissatisfaction. Therefore, there was a need for effective customer service to improve customer satisfaction.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes an information processing means for receiving audio data, an audio recognition means for converting the received audio data into text data, and an emotion analysis means for estimating emotional states from the audio and text data and adding emotional information to the generated meeting minutes. This makes it possible to analyze the emotions of customers in physical stores in real time and provide optimal responses and suggestions.

[0183] "Audio data" refers to a data format in which audio has been digitized.

[0184] "Information processing means" refers to devices and systems for appropriately processing received data.

[0185] "Speech recognition means" refers to a technology or device that analyzes speech data and converts it into text data.

[0186] "Text data" refers to data in sentence format that has been converted by speech recognition technology.

[0187] "Standard language" refers to the standard form of language used in a particular region or group.

[0188] "Language conversion means" refers to a technology or device for changing one language format to another.

[0189] "Means of personal identification" refers to technologies or devices that identify individuals based on their voice or other characteristics.

[0190] "Recording means" refers to a technology or device for storing organized information.

[0191] "Emotion analysis means" refers to a technology or device that estimates a speaker's emotions using text data or audio data.

[0192] "Information provision means" refers to the technology or device that outputs the generated data or information.

[0193] A "regional language dictionary" is a database that compiles the characteristics and vocabulary of a language in a specific region.

[0194] A "voice characteristics database" is a database that stores voice characteristics for the purpose of identifying individuals.

[0195] This invention provides a voice recognition system for enhancing customer interaction in a retail environment. This system includes the following components:

[0196] First, a terminal, acting as an information processing device, is installed in the store to receive customer voice data. This voice data is digitized and sent to a server. The server uses speech recognition to convert the voice data into text data. The software used here is the speech_recognition library. This converted text data is then converted back into standard Japanese. In this process, a regional language dictionary is used as a language conversion tool.

[0197] Next, using a voice characteristics database, the speaker is identified by a personal identification means, and meeting minutes are generated through a recording means. In addition, an emotion analysis means is used to estimate the customer's emotional state from the voice and text data, and based on this, emotion information is added to the meeting minutes. For this purpose, software to operate the emotion engine is required.

[0198] Ultimately, users can access meeting minutes and sentiment information generated from the server through the information provision mechanism. This embodiment makes it possible to analyze customer sentiment in real time in stores and use that information to improve customer service.

[0199] To give a concrete example, if a customer says, "I want to know more about this product," the system will sense the customer's interest and automatically initiate a process to provide relevant additional information. In this way, customer service that takes emotions into account is achieved.

[0200] An example of a prompt might be: "When a customer asks a question about a product, convert the question into text, analyze its sentiment, and generate an appropriate response based on that."

[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0202] Step 1:

[0203] The terminal collects customer voices using a microphone in the store, digitizes the audio data, and sends it to a server. The input is the customer's raw voice data, and the output is digitized audio data. When the terminal collects the audio data, it uses noise cancellation technology to suppress ambient noise.

[0204] Step 2:

[0205] The server converts received digital audio data into text data using speech recognition. The input is digital audio data, and the output is the converted text data. The server uses the speech_recognition library to analyze the audio data, perform word recognition, and convert it to text.

[0206] Step 3:

[0207] The server converts text data into standard language using a language conversion mechanism. The input is text data that may contain dialects, and the output is text data converted to standard language. Here, the server consults a regional language dictionary and replaces region-specific words and expressions with standard language equivalents.

[0208] Step 4:

[0209] The server uses personal identification methods to identify speakers by referring to a voice characteristics database. Input consists of text data and audio data converted to standard language, and output is the identification information of the identified speaker. The server analyzes the voiceprint and compares it with an existing database to identify the speaker.

[0210] Step 5:

[0211] The server uses emotion analysis tools to estimate the customer's emotional state from text and audio data. The input is the speaker's text and audio data, and the output is estimated emotion information. The emotion engine analyzes emotions based on the tone and keywords of the speech, and as a result identifies emotions such as joy and anxiety.

[0212] Step 6:

[0213] The server adds sentiment information to the meeting minutes for each identified speaker and stores it through a recording mechanism. The input is sentiment information and speaker information, and the output is the meeting minutes with the sentiment information added. The server generates the meeting minutes in a structured format and stores them in an appropriate format for later access.

[0214] Step 7:

[0215] Users access meeting minutes and sentiment information generated from the server through an information provision mechanism. The input is the user's access request, and the output is specific meeting minutes and sentiment information. Users can view the necessary information through the store's interface and use it in their interactions.

[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0219] [Second Embodiment]

[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0232] This invention provides a speech recognition system that facilitates communication and improves the efficiency of meeting minute creation in meetings and business negotiations where diverse dialects and multiple languages ​​are present. This system receives speech data, converts it to text in real time, and automatically converts it to standard Japanese as needed. Furthermore, it can identify speakers using voiceprint recognition technology, organize the text for each speaker, and generate meeting minutes.

[0233] Specifically, the terminal collects the voices of participants during meetings and business negotiations. The collected voice data is transmitted to a server via the network. The server converts the received voice data into text using a speech recognition engine. During this process, it refers to a dialect dictionary to replace any dialects in the converted text with standard Japanese.

[0234] Furthermore, the server utilizes voiceprint recognition technology to identify each speaker from the audio data. Using a pre-registered voiceprint database, it can identify individual statements. This makes it clear who said what during an audio conference and allows for the recording of meeting minutes organized by speaker.

[0235] Users can view meeting minutes provided by the server in real time. The minutes include summaries and key action items, and can be edited as needed or shared with other participants. This system reduces communication loss among remote meeting participants and enables effective minute-taking and meeting management.

[0236] As a concrete example, consider a project meeting with a team that includes members from different regions. During the meeting, the terminals simultaneously transmit the participants' voices to the server, which immediately converts the voices into text and standardizes it to a uniform language. For example, if person A, who is participating from Tohoku, says "What should we do today?", this is translated as "What shall we do today?" and displayed in the same format as the other members. In this process, person A's voiceprint is used to automatically classify the statement as A's and reflect it in the meeting minutes. As a result, everyone has a common understanding, and at the end of the meeting, they can obtain an action plan along with detailed meeting minutes.

[0237] The following describes the processing flow.

[0238] Step 1:

[0239] The device captures participants' voices from the microphone during meetings and business negotiations, and converts the audio data into a digital format. The converted digital audio data is then transmitted to the server in real time.

[0240] Step 2:

[0241] The server sequentially processes the audio data received from the terminal and prepares it for input to the speech recognition engine. The server sends the audio data to the processing unit and performs appropriate batch division.

[0242] Step 3:

[0243] The server starts the speech recognition engine and converts speech data into text in real time. Based on the acoustic model and language model, it sequentially converts speech segments into text data.

[0244] Step 4:

[0245] The server applies a dialect dictionary to the text data and automatically converts local dialect expressions to standard Japanese. It detects the relevant dialects in the text and replaces them with standard expressions.

[0246] Step 5:

[0247] The server compares the audio data with a pre-registered voiceprint database and identifies the speaker using voiceprint recognition technology. The server organizes the text data for each speaker and adds information about the speaker.

[0248] Step 6:

[0249] The server generates meeting minutes based on text data organized by speaker. The server arranges the text chronologically and constructs meeting minutes that reflect the content of the meeting.

[0250] Step 7:

[0251] The server extracts summaries and action items from the generated meeting minutes and prepares them for the user. It uses natural language processing to extract and list important information.

[0252] Step 8:

[0253] Users can view meeting minutes provided by the server in real time and make corrections or edits as needed. Users can then share the final meeting minutes with relevant parties.

[0254] (Example 1)

[0255] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0256] In meetings and business negotiations involving diverse dialects and multiple languages, there is a need to create meeting minutes quickly and accurately while maintaining effective communication. Conventional technologies have challenges in accurately transcribing spoken content, identifying speakers, and standardizing language, leading to decreased meeting efficiency.

[0257] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0258] In this invention, the server includes an information processing device for receiving audio information, an audio analysis device for converting the received audio information into symbolic information, a language conversion function device for converting the converted symbolic information into a reference language, a biometric information recognition device for identifying the speaker of the symbolic information, a record generation device for generating meeting minutes for each identified speaker, and an output function device for outputting the generated meeting minutes. This enables accurate and efficient creation of meeting minutes even among participants with diverse linguistic backgrounds.

[0259] An "information processing device" is a device used to receive audio information and process the data.

[0260] "Speech analysis means" refers to a technology or function for converting speech information into symbolic information.

[0261] "Language conversion function means" refers to a technology or function for converting converted symbolic information into a reference language.

[0262] "Biometric information recognition means" refers to a technology or function that analyzes biometric information, such as voiceprints, used to identify the speaker of symbolic information.

[0263] "Record generation means" refers to a technology or function for generating meeting minutes for each identified speaker.

[0264] "Output function means" refers to a technology or function for outputting the generated meeting minutes.

[0265] The system of this invention is designed for use in meetings and business negotiations, enabling smooth communication among participants with diverse backgrounds. For implementation, a mechanism is needed to efficiently handle the process from audio information collection to meeting minute generation.

[0266] The terminal acts as hardware for collecting participants' voices in the meeting room. It is equipped with microphones and voice capture software to acquire high-quality audio from multiple participants. This terminal converts the audio into digital data in real time and transmits it to a server via the network.

[0267] The server receives audio information and converts the audio data into symbolic information using speech analysis tools. In this process, it utilizes existing speech recognition APIs such as Google Speech-to-Text and IBM Watson as its speech recognition engine. Furthermore, the server accurately converts dialects and different languages ​​to a standard language using language conversion tools. Here, dialect dictionaries and language conversion algorithms are used to ensure consistency in the content of speech.

[0268] The server analyzes the characteristics of the voice data using biometric recognition means and identifies the speaker based on pre-registered voiceprints. This allows for accurate organization of statements made during the meeting, and a record generation means is used to generate meeting minutes for each identified speaker.

[0269] The generated meeting minutes can be viewed, edited, and shared in real time by users using the output function. Users can easily manage the generated meeting minutes through a dedicated interface.

[0270] As a concrete example, consider a project team meeting. If a participant from a regional area says, "What should we do today?", the server translates this to "What shall we do today?". Furthermore, it quickly identifies which participant made a particular statement and reflects this in the meeting minutes. Using this system, everyone can have a shared understanding while efficiently creating meeting minutes.

[0271] An example of a prompt using a generative AI model is: "Convert the audio data spoken during the meeting into text in real time, identify the speaker, and automatically generate meeting minutes."

[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0273] Step 1:

[0274] The device uses a microphone to collect participants' voices during meetings and business negotiations, converting analog audio into digital audio data. The input is the raw voice of each participant, and the output is prepared as digital audio data. This data is transmitted in real time to a server via the network.

[0275] Step 2:

[0276] The server converts digital audio data received from the terminal into symbolic information using speech analysis tools. Digital audio data is passed to the server as input, and speech recognition APIs such as Google Speech-to-Text are used to convert the audio into text data. The output is text data in string format.

[0277] Step 3:

[0278] The server uses language conversion functionality to convert the converted text data back to the base language. Here, text data is used as input, and specific dialects and expressions are replaced with those in the base language. Standardized text data is generated as output.

[0279] Step 4:

[0280] The server identifies the speaker from the audio data using biometric recognition means. In this process, audio feature data is used as input, and the speaker is identified by matching it with a pre-registered voiceprint database. The output is speaker identification information accompanying the text data.

[0281] Step 5:

[0282] The server uses a recording generation mechanism to generate meeting minutes for each identified speaker. Standardized text data and speaker information are used as input data, which are combined to generate organized meeting minutes. The output is a completed meeting minute.

[0283] Step 6:

[0284] Users can view meeting minutes provided by the server in real time through an output function. The input is meeting minutes data from the server, which users can edit and share with other participants via a web application or dedicated software. The output is the meeting minutes and related information displayed in a user-friendly format.

[0285] (Application Example 1)

[0286] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart glasses 214 are referred to as a "terminal".

[0287] In a business environment where multiple languages or region-specific dialects are rampant, smooth communication between customers and staff is required. Conventionally, there has been a problem that communication due to different languages and dialects is insufficient, resulting in a decline in customer satisfaction.

[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0289] In this invention, the server includes terminal means for receiving acoustic data, speech recognition means for converting the received acoustic data into text data, and display means for visually outputting the translated text data to a display device. Thereby, voice communication with customers is assisted in real time, enabling smooth mutual understanding in a multilingual environment.

[0290] The "terminal means for receiving acoustic data" is a device that acquires voice as digital information.

[0291] The "speech recognition means" is a technology for analyzing acoustic data and converting it into text information.

[0292] The "language conversion means" is a technology for converting text information from other languages or dialects into a standard language.

[0293] The "voiceprint recognition means" is a technology for identifying individual speakers from voice.

[0294] The "recording means" is a function for accumulating and storing the converted data as a record.

[0295] "Output means" refers to a device that presents the created record in physical or digital format.

[0296] "Display means" refers to a device that visually displays translated text information.

[0297] "Communication methods" refer to technologies for exchanging data between devices.

[0298] The system implementing this invention is equipped with technology for collecting audio data and processing the information in real time. The terminal primarily receives audio data and transmits it to the server. The server converts the received audio data into text data using speech recognition technology. The converted text data is then translated into a standard language by a language conversion means. Advanced natural language processing technology is used for this translation.

[0299] The server uses voiceprint recognition technology to identify the speaker. Specifically, it organizes each utterance by speaker by comparing it with a pre-registered voiceprint database. This generates a record categorized by speaker. The generated record is provided to the user through an output device.

[0300] This system also includes a display mechanism, allowing translated text data to be visually output. In a real-world example, when using smart glasses to serve foreign customers in a shop, the customer's speech is translated into standard Japanese in real time and displayed on the glasses' screen. This enables smooth communication between staff and customers.

[0301] For example, if a French-speaking tourist visits a souvenir shop and asks about the price of an item, the staff member's smart glasses will display the translated text, "This keychain costs 500 yen." Through this concrete example, users can obtain information in real time and receive a quick response. This system is an effective means of improving customer satisfaction.

[0302] Examples of prompt sentences for a generative AI model include the following.

[0303] "Please develop a real-time voice translation system using smart glasses in a physical store. Collect voice data with a microphone and send it to the server. It is a program that performs voice recognition and translation and displays the results on the glasses display. Aims for smooth communication with customers."

[0304] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0305] Step 1:

[0306] The terminal receives acoustic data in a commercial environment. Specifically, it collects the customer's speech in the store via a microphone and prepares the data in digital form. The input is an audio signal, and the output is digital acoustic data.

[0307] Step 2:

[0308] Transfer the digital acoustic data from the terminal to the server. The server inputs the received acoustic data into the voice recognition means. The input here is digital acoustic data, and the server prepares to process the data.

[0309] Step 3:

[0310] The server converts the acoustic data into text data using voice recognition technology. As a specific operation, use the Google Cloud Speech-to-Text API to analyze the voice and generate text information. The input is digital acoustic data, and the output is text data.

[0311] Step 4:

[0312] The server uses language conversion tools to translate text data into a reference language. It uses the Google Cloud Translation API to convert input text into the reference language. The input is the converted text data, and the output is the text data in the reference language.

[0313] Step 5:

[0314] The server uses voiceprint recognition technology to identify the speaker. It compares the received audio data with a pre-registered voiceprint database to identify the speaker. The input is the audio data and the voiceprint database, and the output is information about the identified speaker.

[0315] Step 6:

[0316] The server sends the translated text data to a display device via an output mechanism, and outputs it visually. Specifically, the text data is displayed on the smart glasses' screen to provide information to staff. The input is text data in a reference language, and the output is a visual display.

[0317] Step 7:

[0318] The user reviews the displayed information and responds to the customer. Based on the customer's questions, they understand the text display and provide appropriate information. Input is visual display information, and output is the customer interaction conversation.

[0319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0320] This invention combines a speech recognition system that transcribes participants' speech in real time and converts it into standard Japanese during meetings and business negotiations with a newly developed emotion engine. This makes it possible to recognize the emotional state of participants and organize and record the flow of discussion and important points while considering the emotional context.

[0321] This system first has a terminal collect the voices of multiple participants through a microphone and send the converted audio data to a server. The server then uses a speech recognition engine to convert this audio data into text data. During this process, the server uses a dialect dictionary to perform standard language conversion appropriate for the text data.

[0322] Furthermore, the server uses voiceprint recognition technology to identify speakers and organizes each statement by speaker. By matching it with a voiceprint database, speakers can be identified, and this is used to organize the statements by speaker.

[0323] In addition, an emotion engine installed on the server estimates the emotional state of each speaker through text data and speech analysis. The emotion engine adds this emotional information to the meeting minutes and further prioritizes the importance of summaries and action items based on the intensity of the emotions.

[0324] Users can access meeting minutes provided by the server, which include details of each speaker's remarks along with their estimated emotional state. For example, if participant A says, "I have some concerns about this new proposal," during a meeting, the server recognizes A's emotional state from the audio and text, and this emotional information is included in the meeting minutes. This has the advantage of providing the meeting facilitator with information to plan discussions and follow-ups to address A's concerns.

[0325] Thus, the present invention generates meeting minutes that also take emotional nuances into account, improving the quality of meetings and effectively preventing misunderstandings and communication breakdowns among participants.

[0326] The following describes the processing flow.

[0327] Step 1:

[0328] The terminal captures audio from meetings and business negotiations via its microphone and converts it into digital audio data in real time. This audio data is immediately transmitted to the server.

[0329] Step 2:

[0330] The server prepares the audio data received from the terminal for input into the speech recognition engine. The data is processed in batches and converted into text by the speech recognition engine.

[0331] Step 3:

[0332] The server uses a dialect dictionary to convert the text data obtained through speech recognition into standard Japanese. By replacing dialects and expressions with standard Japanese, the server facilitates overall comprehension.

[0333] Step 4:

[0334] The server uses voiceprint recognition technology to identify the individual speaker. It compares features extracted from the audio data with a voiceprint database and adds speaker information to the text data.

[0335] Step 5:

[0336] The server uses an emotion engine to estimate the speaker's emotions from text data and voice features. The emotion engine identifies specific emotional states and associates that information with the text.

[0337] Step 6:

[0338] The server generates meeting minutes by adding sentiment information to text data organized by speaker. The minutes include detailed records of what was said, who said it, and the estimated sentiment.

[0339] Step 7:

[0340] The server extracts summaries and action items from the generated meeting minutes and prioritizes them based on sentiment information. This information is then highlighted within the meeting minutes.

[0341] Step 8:

[0342] Users can view meeting minutes provided by the server in real time. While reviewing the minutes, users can understand the flow of the discussion by referring to sentiment information and take necessary actions.

[0343] (Example 2)

[0344] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0345] In meetings and business negotiations, it is necessary not only to accurately record participants' statements in real time, but also to understand the emotional state of the speakers and organize and record the flow of the discussion and important points while taking emotional context into account. However, conventional speech recognition systems do not adequately identify emotions or convert dialects to standard Japanese, leading to problems such as misunderstandings and communication losses among participants.

[0346] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0347] In this invention, the server includes speech recognition means for converting speech information into text information, language conversion means for converting the converted text information into a common language, and emotion prediction means for analyzing the text information and speech information to predict emotional states. This makes it possible to record the content of meetings while taking into account emotional nuances, thereby facilitating smooth communication among participants.

[0348] "Audio information" refers to sound data collected from participants in meetings or business negotiations using terminal devices.

[0349] "Textual information" refers to digital text data converted from audio information by speech recognition technology.

[0350] "Common language" refers to a standardized form of language expression that does not include dialects.

[0351] "Identification means" refers to technology that identifies a speaker by analyzing the speaker's voice characteristics from textual information.

[0352] "Emotion prediction means" refers to technology that predicts and identifies the emotional state of a speaker by analyzing textual and auditory information.

[0353] "Recording means" refers to a means of organizing the content of a conversation for each identified speaker and generating a meeting record.

[0354] A "dialect dictionary" refers to a language database used to translate regionally specific linguistic expressions into standard Japanese.

[0355] One embodiment of the present invention is a system that records the statements of participants in meetings and business negotiations in real time and generates meeting minutes that take into account emotional context. This system handles everything from collecting audio data to sentiment analysis and generating meeting minutes in an integrated manner.

[0356] The device uses a microphone to collect audio and converts participants' speech into digital audio data. The audio data collected by the device is immediately transmitted to a server via the network. Noise reduction technology is employed during this process to maintain audio quality.

[0357] The server first converts the received audio data into text data using a speech recognition engine. A speech recognition API is used for this process, generating highly accurate textual information. Furthermore, the converted textual information is standardized to a common language through a language conversion function. The software used in this process incorporates a dialect dictionary.

[0358] Next, the server uses an identification mechanism to identify the speaker from the text information. This mechanism uses voiceprint data to identify the speaker and organizes the statements by speaker. Furthermore, the emotion prediction engine on the server analyzes the text data and voice characteristics to predict the speaker's emotional state.

[0359] Ultimately, users can access the meeting minutes generated by the server and review the details of each statement, including emotional nuances. For example, if participant A says, "I have some concerns about this new proposal," in order to ensure the meeting runs smoothly, the server can reflect A's concern in the minutes. This allows users to assess concerns about the proposal and use that information to plan follow-up.

[0360] A possible prompt would be, "What methods and technologies should be used to create a system that records participants' statements and emotions in real time during a meeting and generates meeting minutes based on the estimated emotions?" This system would improve communication among participants and enhance the quality of the meeting.

[0361] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0362] Step 1:

[0363] The terminal uses a microphone to collect speech from meeting participants in real time. The input is an analog signal of raw sound, which is converted into digital audio data. The converted audio data is then transmitted to a server via the network. Noise reduction technology is used to improve the quality of the audio data.

[0364] Step 2:

[0365] The server converts digital audio data received from the terminal into text information using a speech recognition engine. The input is digital audio data, and the output is text information. Speech recognition automatically transcribes spoken content into text, generating highly accurate text data.

[0366] Step 3:

[0367] The server applies a language conversion function to convert the converted character information into the standard language. The input is text data, and the output is character information in the standard language. By using a pre-built dialect dictionary, it performs the task of standardizing regional expressions.

[0368] Step 4:

[0369] The server identifies the speaker using identification methods based on the text information converted to a common language. The input here is the text information in a common language, and the output is the identification information of each speaker. Voiceprint data is used to identify speakers with high accuracy and organize statements by speaker.

[0370] Step 5:

[0371] The server uses an emotion prediction engine to analyze the speaker's emotional state based on identified textual information. The input is standard language textual information with speaker information, and the output is text information including emotional state. It analyzes the features of the audio and text data to identify the emotional context.

[0372] Step 6:

[0373] Users access meeting minutes generated from the server, which include sentiment information. The generated minutes provide details of each statement, incorporating emotional nuances. This makes it easier for users to effectively manage meeting content and plan subsequent actions.

[0374] (Application Example 2)

[0375] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0376] In brick-and-mortar retail settings, understanding customer emotions and providing accurate information and suggestions has been difficult with conventional technologies. In particular, there was a lack of mechanisms to quickly detect and address customer anxiety or dissatisfaction. Therefore, there was a need for effective customer service to improve customer satisfaction.

[0377] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0378] In this invention, the server includes an information processing means for receiving audio data, an audio recognition means for converting the received audio data into text data, and an emotion analysis means for estimating emotional states from the audio and text data and adding emotional information to the generated meeting minutes. This makes it possible to analyze the emotions of customers in physical stores in real time and provide optimal responses and suggestions.

[0379] "Audio data" refers to a data format in which audio has been digitized.

[0380] "Information processing means" refers to devices and systems for appropriately processing received data.

[0381] "Speech recognition means" refers to a technology or device that analyzes speech data and converts it into text data.

[0382] "Text data" refers to data in sentence format that has been converted by speech recognition technology.

[0383] "Standard language" refers to the standard form of language used in a particular region or group.

[0384] "Language conversion means" refers to a technology or device for changing one language format to another.

[0385] "Means of personal identification" refers to technologies or devices that identify individuals based on their voice or other characteristics.

[0386] "Recording means" refers to a technology or device for storing organized information.

[0387] "Emotion analysis means" refers to a technology or device that estimates a speaker's emotions using text data or audio data.

[0388] "Information provision means" refers to the technology or device that outputs the generated data or information.

[0389] A "regional language dictionary" is a database that compiles the characteristics and vocabulary of a language in a specific region.

[0390] A "voice characteristics database" is a database that stores voice characteristics for the purpose of identifying individuals.

[0391] This invention provides a voice recognition system for enhancing customer interaction in a retail environment. This system includes the following components:

[0392] First, a terminal, acting as an information processing device, is installed in the store to receive customer voice data. This voice data is digitized and sent to a server. The server uses speech recognition to convert the voice data into text data. The software used here is the speech_recognition library. This converted text data is then converted back into standard Japanese. In this process, a regional language dictionary is used as a language conversion tool.

[0393] Next, using a voice characteristics database, the speaker is identified by a personal identification means, and meeting minutes are generated through a recording means. In addition, an emotion analysis means is used to estimate the customer's emotional state from the voice and text data, and based on this, emotion information is added to the meeting minutes. For this purpose, software to operate the emotion engine is required.

[0394] Ultimately, users can access meeting minutes and sentiment information generated from the server through the information provision mechanism. This embodiment makes it possible to analyze customer sentiment in real time in stores and use that information to improve customer service.

[0395] To give a concrete example, if a customer says, "I want to know more about this product," the system will sense the customer's interest and automatically initiate a process to provide relevant additional information. In this way, customer service that takes emotions into account is achieved.

[0396] An example of a prompt might be: "When a customer asks a question about a product, convert the question into text, analyze its sentiment, and generate an appropriate response based on that."

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] The terminal collects customer voices using a microphone in the store, digitizes the audio data, and sends it to a server. The input is the customer's raw voice data, and the output is digitized audio data. When the terminal collects the audio data, it uses noise cancellation technology to suppress ambient noise.

[0400] Step 2:

[0401] The server converts received digital audio data into text data using speech recognition. The input is digital audio data, and the output is the converted text data. The server uses the speech_recognition library to analyze the audio data, perform word recognition, and convert it to text.

[0402] Step 3:

[0403] The server converts text data into standard language using a language conversion mechanism. The input is text data that may contain dialects, and the output is text data converted to standard language. Here, the server consults a regional language dictionary and replaces region-specific words and expressions with standard language equivalents.

[0404] Step 4:

[0405] The server uses personal identification methods to identify speakers by referring to a voice characteristics database. Input consists of text data and audio data converted to standard language, and output is the identification information of the identified speaker. The server analyzes the voiceprint and compares it with an existing database to identify the speaker.

[0406] Step 5:

[0407] The server uses emotion analysis tools to estimate the customer's emotional state from text and audio data. The input is the speaker's text and audio data, and the output is estimated emotion information. The emotion engine analyzes emotions based on the tone and keywords of the speech, and as a result identifies emotions such as joy and anxiety.

[0408] Step 6:

[0409] The server adds sentiment information to the meeting minutes for each identified speaker and stores it through a recording mechanism. The input is sentiment information and speaker information, and the output is the meeting minutes with the sentiment information added. The server generates the meeting minutes in a structured format and stores them in an appropriate format for later access.

[0410] Step 7:

[0411] Users access meeting minutes and sentiment information generated from the server through an information provision mechanism. The input is the user's access request, and the output is specific meeting minutes and sentiment information. Users can view the necessary information through the store's interface and use it in their interactions.

[0412] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0413] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0414] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0415] [Third Embodiment]

[0416] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0417] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0418] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0419] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0420] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0422] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0423] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0424] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0425] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0426] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0427] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0428] This invention provides a speech recognition system that facilitates communication and improves the efficiency of meeting minute creation in meetings and business negotiations where diverse dialects and multiple languages ​​are present. This system receives speech data, converts it to text in real time, and automatically converts it to standard Japanese as needed. Furthermore, it can identify speakers using voiceprint recognition technology, organize the text for each speaker, and generate meeting minutes.

[0429] Specifically, the terminal collects the voices of participants during meetings and business negotiations. The collected voice data is transmitted to a server via the network. The server converts the received voice data into text using a speech recognition engine. During this process, it refers to a dialect dictionary to replace any dialects in the converted text with standard Japanese.

[0430] Furthermore, the server utilizes voiceprint recognition technology to identify each speaker from the audio data. Using a pre-registered voiceprint database, it can identify individual statements. This makes it clear who said what during an audio conference and allows for the recording of meeting minutes organized by speaker.

[0431] Users can view meeting minutes provided by the server in real time. The minutes include summaries and key action items, and can be edited as needed or shared with other participants. This system reduces communication loss among remote meeting participants and enables effective minute-taking and meeting management.

[0432] As a concrete example, consider a project meeting with a team that includes members from different regions. During the meeting, the terminals simultaneously transmit the participants' voices to the server, which immediately converts the voices into text and standardizes it to a uniform language. For example, if person A, who is participating from Tohoku, says "What should we do today?", this is translated as "What shall we do today?" and displayed in the same format as the other members. In this process, person A's voiceprint is used to automatically classify the statement as A's and reflect it in the meeting minutes. As a result, everyone has a common understanding, and at the end of the meeting, they can obtain an action plan along with detailed meeting minutes.

[0433] The following describes the processing flow.

[0434] Step 1:

[0435] The device captures participants' voices from the microphone during meetings and business negotiations, and converts the audio data into a digital format. The converted digital audio data is then transmitted to the server in real time.

[0436] Step 2:

[0437] The server sequentially processes the audio data received from the terminal and prepares it for input to the speech recognition engine. The server sends the audio data to the processing unit and performs appropriate batch division.

[0438] Step 3:

[0439] The server starts the speech recognition engine and converts speech data into text in real time. Based on the acoustic model and language model, it sequentially converts speech segments into text data.

[0440] Step 4:

[0441] The server applies a dialect dictionary to the text data and automatically converts local dialect expressions to standard Japanese. It detects the relevant dialects in the text and replaces them with standard expressions.

[0442] Step 5:

[0443] The server compares the audio data with a pre-registered voiceprint database and identifies the speaker using voiceprint recognition technology. The server organizes the text data for each speaker and adds information about the speaker.

[0444] Step 6:

[0445] The server generates meeting minutes based on text data organized by speaker. The server arranges the text chronologically and constructs meeting minutes that reflect the content of the meeting.

[0446] Step 7:

[0447] The server extracts summaries and action items from the generated meeting minutes and prepares them for the user. It uses natural language processing to extract and list important information.

[0448] Step 8:

[0449] Users can view meeting minutes provided by the server in real time and make corrections or edits as needed. Users can then share the final meeting minutes with relevant parties.

[0450] (Example 1)

[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0452] In meetings and business negotiations involving diverse dialects and multiple languages, there is a need to create meeting minutes quickly and accurately while maintaining effective communication. Conventional technologies have challenges in accurately transcribing spoken content, identifying speakers, and standardizing language, leading to decreased meeting efficiency.

[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0454] In this invention, the server includes an information processing device for receiving audio information, an audio analysis device for converting the received audio information into symbolic information, a language conversion function device for converting the converted symbolic information into a reference language, a biometric information recognition device for identifying the speaker of the symbolic information, a record generation device for generating meeting minutes for each identified speaker, and an output function device for outputting the generated meeting minutes. This enables accurate and efficient creation of meeting minutes even among participants with diverse linguistic backgrounds.

[0455] An "information processing device" is a device used to receive audio information and process the data.

[0456] "Speech analysis means" refers to a technology or function for converting speech information into symbolic information.

[0457] "Language conversion function means" refers to a technology or function for converting converted symbolic information into a reference language.

[0458] "Biometric information recognition means" refers to a technology or function that analyzes biometric information, such as voiceprints, used to identify the speaker of symbolic information.

[0459] "Record generation means" refers to a technology or function for generating meeting minutes for each identified speaker.

[0460] "Output function means" refers to a technology or function for outputting the generated meeting minutes.

[0461] The system of this invention is designed for use in meetings and business negotiations, enabling smooth communication among participants with diverse backgrounds. For implementation, a mechanism is needed to efficiently handle the process from audio information collection to meeting minute generation.

[0462] The terminal acts as hardware for collecting participants' voices in the meeting room. It is equipped with microphones and voice capture software to acquire high-quality audio from multiple participants. This terminal converts the audio into digital data in real time and transmits it to a server via the network.

[0463] The server receives audio information and converts the audio data into symbolic information using speech analysis tools. In this process, it utilizes existing speech recognition APIs such as Google Speech-to-Text and IBM Watson as its speech recognition engine. Furthermore, the server accurately converts dialects and different languages ​​to a standard language using language conversion tools. Here, dialect dictionaries and language conversion algorithms are used to ensure consistency in the content of speech.

[0464] The server analyzes the characteristics of the voice data using biometric recognition means and identifies the speaker based on pre-registered voiceprints. This allows for accurate organization of statements made during the meeting, and a record generation means is used to generate meeting minutes for each identified speaker.

[0465] The generated meeting minutes can be viewed, edited, and shared in real time by users using the output function. Users can easily manage the generated meeting minutes through a dedicated interface.

[0466] As a concrete example, consider a project team meeting. If a participant from a regional area says, "What should we do today?", the server translates this to "What shall we do today?". Furthermore, it quickly identifies which participant made a particular statement and reflects this in the meeting minutes. Using this system, everyone can have a shared understanding while efficiently creating meeting minutes.

[0467] An example of a prompt using a generative AI model is: "Convert the audio data spoken during the meeting into text in real time, identify the speaker, and automatically generate meeting minutes."

[0468] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0469] Step 1:

[0470] The device uses a microphone to collect participants' voices during meetings and business negotiations, converting analog audio into digital audio data. The input is the raw voice of each participant, and the output is prepared as digital audio data. This data is transmitted in real time to a server via the network.

[0471] Step 2:

[0472] The server converts digital audio data received from the terminal into symbolic information using speech analysis tools. Digital audio data is passed to the server as input, and speech recognition APIs such as Google Speech-to-Text are used to convert the audio into text data. The output is text data in string format.

[0473] Step 3:

[0474] The server uses language conversion functionality to convert the converted text data back to the base language. Here, text data is used as input, and specific dialects and expressions are replaced with those in the base language. Standardized text data is generated as output.

[0475] Step 4:

[0476] The server identifies the speaker from the audio data using biometric recognition means. In this process, audio feature data is used as input, and the speaker is identified by matching it with a pre-registered voiceprint database. The output is speaker identification information accompanying the text data.

[0477] Step 5:

[0478] The server uses a recording generation mechanism to generate meeting minutes for each identified speaker. Standardized text data and speaker information are used as input data, which are combined to generate organized meeting minutes. The output is a completed meeting minute.

[0479] Step 6:

[0480] Users can view meeting minutes provided by the server in real time through an output function. The input is meeting minutes data from the server, which users can edit and share with other participants via a web application or dedicated software. The output is the meeting minutes and related information displayed in a user-friendly format.

[0481] (Application Example 1)

[0482] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0483] In commercial environments where multiple languages ​​or regional dialects are spoken, there is a need to facilitate communication between customers and staff. Traditionally, insufficient communication due to different languages ​​or dialects has led to problems with decreased customer satisfaction.

[0484] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0485] In this invention, the server includes terminal means for receiving acoustic data, speech recognition means for converting the received acoustic data into text data, and display means for visually outputting the translated text data to a display device. This enables real-time assistance for voice communication with customers and facilitates smooth mutual understanding in a multilingual environment.

[0486] "Terminal means for receiving acoustic data" refers to a device that acquires sound as digital information.

[0487] "Speech recognition means" refers to technology that analyzes acoustic data and converts it into text information.

[0488] "Language conversion means" refers to technologies for converting text information from another language or dialect into a standard language.

[0489] "Voiceprint recognition means" refers to technology for identifying individual speakers from their voices.

[0490] A "recording means" is a function that stores the converted data and saves it as a record.

[0491] "Output means" refers to a device that presents the created record in physical or digital format.

[0492] "Display means" refers to a device that visually displays translated text information.

[0493] "Communication methods" refer to technologies for exchanging data between devices.

[0494] The system implementing this invention is equipped with technology for collecting audio data and processing the information in real time. The terminal primarily receives audio data and transmits it to the server. The server converts the received audio data into text data using speech recognition technology. The converted text data is then translated into a standard language by a language conversion means. Advanced natural language processing technology is used for this translation.

[0495] The server uses voiceprint recognition technology to identify the speaker. Specifically, it organizes each utterance by speaker by comparing it with a pre-registered voiceprint database. This generates a record categorized by speaker. The generated record is provided to the user through an output device.

[0496] This system also includes a display mechanism, allowing translated text data to be visually output. In a real-world example, when using smart glasses to serve foreign customers in a shop, the customer's speech is translated into standard Japanese in real time and displayed on the glasses' screen. This enables smooth communication between staff and customers.

[0497] For example, if a French-speaking tourist visits a souvenir shop and asks about the price of an item, the staff member's smart glasses will display the translated text, "This keychain costs 500 yen." Through this concrete example, users can obtain information in real time and receive a quick response. This system is an effective means of improving customer satisfaction.

[0498] Examples of prompts for a generative AI model include the following:

[0499] "Develop a real-time voice translation system using smart glasses for use in physical stores. The system will collect voice data via a microphone, send it to a server, perform voice recognition and translation, and display the results on the glasses' display. The goal is smooth communication with customers."

[0500] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0501] Step 1:

[0502] The terminal receives audio data in a commercial environment. Specifically, it collects customer speech via microphones within a store and prepares the data in digital format. The input is an audio signal, and the output is digital audio data.

[0503] Step 2:

[0504] Digital audio data is transferred from the terminal to the server. The server inputs the received audio data into the speech recognition system. The input here is digital audio data, and the server prepares to process that data.

[0505] Step 3:

[0506] The server uses speech recognition technology to convert audio data into text data. Specifically, it uses the Google Cloud Speech-to-Text API to analyze speech and generate text information. The input is digital audio data, and the output is text data.

[0507] Step 4:

[0508] The server uses language conversion tools to translate text data into a reference language. It uses the Google Cloud Translation API to convert input text into the reference language. The input is the converted text data, and the output is the text data in the reference language.

[0509] Step 5:

[0510] The server uses voiceprint recognition technology to identify the speaker. It compares the received audio data with a pre-registered voiceprint database to identify the speaker. The input is the audio data and the voiceprint database, and the output is information about the identified speaker.

[0511] Step 6:

[0512] The server sends the translated text data to a display device via an output mechanism, and outputs it visually. Specifically, the text data is displayed on the smart glasses' screen to provide information to staff. The input is text data in a reference language, and the output is a visual display.

[0513] Step 7:

[0514] The user reviews the displayed information and responds to the customer. Based on the customer's questions, they understand the text display and provide appropriate information. Input is visual display information, and output is the customer interaction conversation.

[0515] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0516] This invention combines a speech recognition system that transcribes participants' speech in real time and converts it into standard Japanese during meetings and business negotiations with a newly developed emotion engine. This makes it possible to recognize the emotional state of participants and organize and record the flow of discussion and important points while considering the emotional context.

[0517] This system first has a terminal collect the voices of multiple participants through a microphone and send the converted audio data to a server. The server then uses a speech recognition engine to convert this audio data into text data. During this process, the server uses a dialect dictionary to perform standard language conversion appropriate for the text data.

[0518] Furthermore, the server uses voiceprint recognition technology to identify speakers and organizes each statement by speaker. By matching it with a voiceprint database, speakers can be identified, and this is used to organize the statements by speaker.

[0519] In addition, an emotion engine installed on the server estimates the emotional state of each speaker through text data and speech analysis. The emotion engine adds this emotional information to the meeting minutes and further prioritizes the importance of summaries and action items based on the intensity of the emotions.

[0520] Users can access meeting minutes provided by the server, which include details of each speaker's remarks along with their estimated emotional state. For example, if participant A says, "I have some concerns about this new proposal," during a meeting, the server recognizes A's emotional state from the audio and text, and this emotional information is included in the meeting minutes. This has the advantage of providing the meeting facilitator with information to plan discussions and follow-ups to address A's concerns.

[0521] Thus, the present invention generates meeting minutes that also take emotional nuances into account, improving the quality of meetings and effectively preventing misunderstandings and communication breakdowns among participants.

[0522] The following describes the processing flow.

[0523] Step 1:

[0524] The terminal captures audio from meetings and business negotiations via its microphone and converts it into digital audio data in real time. This audio data is immediately transmitted to the server.

[0525] Step 2:

[0526] The server prepares the audio data received from the terminal for input into the speech recognition engine. The data is processed in batches and converted into text by the speech recognition engine.

[0527] Step 3:

[0528] The server uses a dialect dictionary to convert the text data obtained through speech recognition into standard Japanese. By replacing dialects and expressions with standard Japanese, the server facilitates overall comprehension.

[0529] Step 4:

[0530] The server uses voiceprint recognition technology to identify the individual speaker. It compares features extracted from the audio data with a voiceprint database and adds speaker information to the text data.

[0531] Step 5:

[0532] The server uses an emotion engine to estimate the speaker's emotions from text data and voice features. The emotion engine identifies specific emotional states and associates that information with the text.

[0533] Step 6:

[0534] The server generates meeting minutes by adding sentiment information to text data organized by speaker. The minutes include detailed records of what was said, who said it, and the estimated sentiment.

[0535] Step 7:

[0536] The server extracts summaries and action items from the generated meeting minutes and prioritizes them based on sentiment information. This information is then highlighted within the meeting minutes.

[0537] Step 8:

[0538] Users can view meeting minutes provided by the server in real time. While reviewing the minutes, users can understand the flow of the discussion by referring to sentiment information and take necessary actions.

[0539] (Example 2)

[0540] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0541] In meetings and business negotiations, it is necessary not only to accurately record participants' statements in real time, but also to understand the emotional state of the speakers and organize and record the flow of the discussion and important points while taking emotional context into account. However, conventional speech recognition systems do not adequately identify emotions or convert dialects to standard Japanese, leading to problems such as misunderstandings and communication losses among participants.

[0542] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0543] In this invention, the server includes speech recognition means for converting speech information into text information, language conversion means for converting the converted text information into a common language, and emotion prediction means for analyzing the text information and speech information to predict emotional states. This makes it possible to record the content of meetings while taking into account emotional nuances, thereby facilitating smooth communication among participants.

[0544] "Audio information" refers to sound data collected from participants in meetings or business negotiations using terminal devices.

[0545] "Textual information" refers to digital text data converted from audio information by speech recognition technology.

[0546] "Common language" refers to a standardized form of language expression that does not include dialects.

[0547] "Identification means" refers to technology that identifies a speaker by analyzing the speaker's voice characteristics from textual information.

[0548] "Emotion prediction means" refers to technology that predicts and identifies the emotional state of a speaker by analyzing textual and auditory information.

[0549] "Recording means" refers to a means of organizing the content of a conversation for each identified speaker and generating a meeting record.

[0550] A "dialect dictionary" refers to a language database used to translate regionally specific linguistic expressions into standard Japanese.

[0551] One embodiment of the present invention is a system that records the statements of participants in meetings and business negotiations in real time and generates meeting minutes that take into account emotional context. This system handles everything from collecting audio data to sentiment analysis and generating meeting minutes in an integrated manner.

[0552] The device uses a microphone to collect audio and converts participants' speech into digital audio data. The audio data collected by the device is immediately transmitted to a server via the network. Noise reduction technology is employed during this process to maintain audio quality.

[0553] The server first converts the received audio data into text data using a speech recognition engine. A speech recognition API is used for this process, generating highly accurate textual information. Furthermore, the converted textual information is standardized to a common language through a language conversion function. The software used in this process incorporates a dialect dictionary.

[0554] Next, the server uses an identification mechanism to identify the speaker from the text information. This mechanism uses voiceprint data to identify the speaker and organizes the statements by speaker. Furthermore, the emotion prediction engine on the server analyzes the text data and voice characteristics to predict the speaker's emotional state.

[0555] Ultimately, users can access the meeting minutes generated by the server and review the details of each statement, including emotional nuances. For example, if participant A says, "I have some concerns about this new proposal," in order to ensure the meeting runs smoothly, the server can reflect A's concern in the minutes. This allows users to assess concerns about the proposal and use that information to plan follow-up.

[0556] A possible prompt would be, "What methods and technologies should be used to create a system that records participants' statements and emotions in real time during a meeting and generates meeting minutes based on the estimated emotions?" This system would improve communication among participants and enhance the quality of the meeting.

[0557] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0558] Step 1:

[0559] The terminal uses a microphone to collect speech from meeting participants in real time. The input is an analog signal of raw sound, which is converted into digital audio data. The converted audio data is then transmitted to a server via the network. Noise reduction technology is used to improve the quality of the audio data.

[0560] Step 2:

[0561] The server converts digital audio data received from the terminal into text information using a speech recognition engine. The input is digital audio data, and the output is text information. Speech recognition automatically transcribes spoken content into text, generating highly accurate text data.

[0562] Step 3:

[0563] The server applies a language conversion function to convert the converted character information into the standard language. The input is text data, and the output is character information in the standard language. By using a pre-built dialect dictionary, it performs the task of standardizing regional expressions.

[0564] Step 4:

[0565] The server identifies the speaker using identification methods based on the text information converted to a common language. The input here is the text information in a common language, and the output is the identification information of each speaker. Voiceprint data is used to identify speakers with high accuracy and organize statements by speaker.

[0566] Step 5:

[0567] The server uses an emotion prediction engine to analyze the speaker's emotional state based on identified textual information. The input is standard language textual information with speaker information, and the output is text information including emotional state. It analyzes the features of the audio and text data to identify the emotional context.

[0568] Step 6:

[0569] Users access meeting minutes generated from the server, which include sentiment information. The generated minutes provide details of each statement, incorporating emotional nuances. This makes it easier for users to effectively manage meeting content and plan subsequent actions.

[0570] (Application Example 2)

[0571] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] In brick-and-mortar retail settings, understanding customer emotions and providing accurate information and suggestions has been difficult with conventional technologies. In particular, there was a lack of mechanisms to quickly detect and address customer anxiety or dissatisfaction. Therefore, there was a need for effective customer service to improve customer satisfaction.

[0573] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0574] In this invention, the server includes an information processing means for receiving audio data, an audio recognition means for converting the received audio data into text data, and an emotion analysis means for estimating emotional states from the audio and text data and adding emotional information to the generated meeting minutes. This makes it possible to analyze the emotions of customers in physical stores in real time and provide optimal responses and suggestions.

[0575] "Audio data" refers to a data format in which audio has been digitized.

[0576] "Information processing means" refers to devices and systems for appropriately processing received data.

[0577] "Speech recognition means" refers to a technology or device that analyzes speech data and converts it into text data.

[0578] "Text data" refers to data in sentence format that has been converted by speech recognition technology.

[0579] "Standard language" refers to the standard form of language used in a particular region or group.

[0580] "Language conversion means" refers to a technology or device for changing one language format to another.

[0581] "Means of personal identification" refers to technologies or devices that identify individuals based on their voice or other characteristics.

[0582] "Recording means" refers to a technology or device for storing organized information.

[0583] "Emotion analysis means" refers to a technology or device that estimates a speaker's emotions using text data or audio data.

[0584] "Information provision means" refers to the technology or device that outputs the generated data or information.

[0585] A "regional language dictionary" is a database that compiles the characteristics and vocabulary of a language in a specific region.

[0586] A "voice characteristics database" is a database that stores voice characteristics for the purpose of identifying individuals.

[0587] This invention provides a voice recognition system for enhancing customer interaction in a retail environment. This system includes the following components:

[0588] First, a terminal, acting as an information processing device, is installed in the store to receive customer voice data. This voice data is digitized and sent to a server. The server uses speech recognition to convert the voice data into text data. The software used here is the speech_recognition library. This converted text data is then converted back into standard Japanese. In this process, a regional language dictionary is used as a language conversion tool.

[0589] Next, using a voice characteristics database, the speaker is identified by a personal identification means, and meeting minutes are generated through a recording means. In addition, an emotion analysis means is used to estimate the customer's emotional state from the voice and text data, and based on this, emotion information is added to the meeting minutes. For this purpose, software to operate the emotion engine is required.

[0590] Ultimately, users can access meeting minutes and sentiment information generated from the server through the information provision mechanism. This embodiment makes it possible to analyze customer sentiment in real time in stores and use that information to improve customer service.

[0591] To give a concrete example, if a customer says, "I want to know more about this product," the system will sense the customer's interest and automatically initiate a process to provide relevant additional information. In this way, customer service that takes emotions into account is achieved.

[0592] An example of a prompt might be: "When a customer asks a question about a product, convert the question into text, analyze its sentiment, and generate an appropriate response based on that."

[0593] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0594] Step 1:

[0595] The terminal collects customer voices using a microphone in the store, digitizes the audio data, and sends it to a server. The input is the customer's raw voice data, and the output is digitized audio data. When the terminal collects the audio data, it uses noise cancellation technology to suppress ambient noise.

[0596] Step 2:

[0597] The server converts received digital audio data into text data using speech recognition. The input is digital audio data, and the output is the converted text data. The server uses the speech_recognition library to analyze the audio data, perform word recognition, and convert it to text.

[0598] Step 3:

[0599] The server converts text data into standard language using a language conversion mechanism. The input is text data that may contain dialects, and the output is text data converted into standard language. Here, the server consults a regional language dictionary and replaces region-specific words and expressions with standard language equivalents.

[0600] Step 4:

[0601] The server uses personal identification methods to identify speakers by referring to a voice characteristics database. Input consists of text data and audio data converted to standard language, and output is the identification information of the identified speaker. The server analyzes the voiceprint and compares it with an existing database to identify the speaker.

[0602] Step 5:

[0603] The server uses emotion analysis tools to estimate the customer's emotional state from text and audio data. The input is the speaker's text and audio data, and the output is estimated emotion information. The emotion engine analyzes emotions based on the tone and keywords of the speech, and as a result identifies emotions such as joy and anxiety.

[0604] Step 6:

[0605] The server adds sentiment information to the meeting minutes for each identified speaker and stores it through a recording mechanism. The input is sentiment information and speaker information, and the output is the meeting minutes with the sentiment information added. The server generates the meeting minutes in a structured format and stores them in an appropriate format for later access.

[0606] Step 7:

[0607] Users access meeting minutes and sentiment information generated from the server through an information provision mechanism. Input is the user's access request, and output is specific meeting minutes and sentiment information. Users can view the necessary information through the store's interface and use it in their interactions.

[0608] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0609] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0610] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0611] [Fourth Embodiment]

[0612] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0613] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0614] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0615] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0616] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0617] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0618] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0619] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0620] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0621] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0622] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0623] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0624] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0625] This invention provides a speech recognition system that facilitates communication and improves the efficiency of meeting minute creation in meetings and business negotiations where diverse dialects and multiple languages ​​are present. This system receives speech data, converts it to text in real time, and automatically converts it to standard Japanese as needed. Furthermore, it can identify speakers using voiceprint recognition technology, organize the text for each speaker, and generate meeting minutes.

[0626] Specifically, the terminal collects the voices of participants during meetings and business negotiations. The collected voice data is transmitted to a server via the network. The server converts the received voice data into text using a speech recognition engine. During this process, it refers to a dialect dictionary to replace any dialects in the converted text with standard Japanese.

[0627] Furthermore, the server utilizes voiceprint recognition technology to identify each speaker from the audio data. Using a pre-registered voiceprint database, it can identify individual statements. This makes it clear who said what during an audio conference and allows for the recording of meeting minutes organized by speaker.

[0628] Users can view meeting minutes provided by the server in real time. The minutes include summaries and key action items, and can be edited as needed or shared with other participants. This system reduces communication loss among remote meeting participants and enables effective minute-taking and meeting management.

[0629] As a concrete example, consider a project meeting with a team that includes members from different regions. During the meeting, the terminals simultaneously transmit the participants' voices to the server, which immediately converts the voices into text and standardizes it to a uniform language. For example, if person A, who is participating from Tohoku, says "What should we do today?", this is translated as "What shall we do today?" and displayed in the same format as the other members. In this process, person A's voiceprint is used to automatically classify the statement as A's and reflect it in the meeting minutes. As a result, everyone has a common understanding, and at the end of the meeting, they can obtain an action plan along with detailed meeting minutes.

[0630] The following describes the processing flow.

[0631] Step 1:

[0632] The device captures participants' voices from the microphone during meetings and business negotiations, and converts the audio data into a digital format. The converted digital audio data is then transmitted to the server in real time.

[0633] Step 2:

[0634] The server sequentially processes the audio data received from the terminal and prepares it for input to the speech recognition engine. The server sends the audio data to the processing unit and performs appropriate batch division.

[0635] Step 3:

[0636] The server starts the speech recognition engine and converts speech data into text in real time. Based on the acoustic model and language model, it sequentially converts speech segments into text data.

[0637] Step 4:

[0638] The server applies a dialect dictionary to the text data and automatically converts local dialect expressions to standard Japanese. It detects the relevant dialects in the text and replaces them with standard expressions.

[0639] Step 5:

[0640] The server compares the audio data with a pre-registered voiceprint database and identifies the speaker using voiceprint recognition technology. The server organizes the text data for each speaker and adds information about the speaker.

[0641] Step 6:

[0642] The server generates meeting minutes based on text data organized by speaker. The server arranges the text chronologically and constructs meeting minutes that reflect the content of the meeting.

[0643] Step 7:

[0644] The server extracts summaries and action items from the generated meeting minutes and prepares them for the user. It uses natural language processing to extract and list important information.

[0645] Step 8:

[0646] Users can view meeting minutes provided by the server in real time and make corrections or edits as needed. Users can then share the final meeting minutes with relevant parties.

[0647] (Example 1)

[0648] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0649] In meetings and business negotiations involving diverse dialects and multiple languages, there is a need to create meeting minutes quickly and accurately while maintaining effective communication. Conventional technologies have challenges in accurately transcribing spoken content, identifying speakers, and standardizing language, leading to decreased meeting efficiency.

[0650] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0651] In this invention, the server includes an information processing device for receiving audio information, an audio analysis device for converting the received audio information into symbolic information, a language conversion function device for converting the converted symbolic information into a reference language, a biometric information recognition device for identifying the speaker of the symbolic information, a record generation device for generating meeting minutes for each identified speaker, and an output function device for outputting the generated meeting minutes. This enables accurate and efficient creation of meeting minutes even among participants with diverse linguistic backgrounds.

[0652] An "information processing device" is a device used to receive audio information and process the data.

[0653] "Speech analysis means" refers to a technology or function for converting speech information into symbolic information.

[0654] "Language conversion function means" refers to a technology or function for converting converted symbolic information into a reference language.

[0655] "Biometric information recognition means" refers to a technology or function that analyzes biometric information, such as voiceprints, used to identify the speaker of symbolic information.

[0656] "Record generation means" refers to a technology or function for generating meeting minutes for each identified speaker.

[0657] "Output function means" refers to a technology or function for outputting the generated meeting minutes.

[0658] The system of this invention is designed for use in meetings and business negotiations, enabling smooth communication among participants with diverse backgrounds. For implementation, a mechanism is needed to efficiently handle the process from audio information collection to meeting minute generation.

[0659] The terminal acts as hardware for collecting participants' voices in the meeting room. It is equipped with microphones and voice capture software to acquire high-quality audio from multiple participants. This terminal converts the audio into digital data in real time and transmits it to a server via the network.

[0660] The server receives audio information and converts the audio data into symbolic information using speech analysis tools. In this process, it utilizes existing speech recognition APIs such as Google Speech-to-Text and IBM Watson as its speech recognition engine. Furthermore, the server accurately converts dialects and different languages ​​to a standard language using language conversion tools. Here, dialect dictionaries and language conversion algorithms are used to ensure consistency in the content of speech.

[0661] The server analyzes the characteristics of the voice data using biometric recognition means and identifies the speaker based on pre-registered voiceprints. This allows for accurate organization of statements made during the meeting, and a record generation means is used to generate meeting minutes for each identified speaker.

[0662] The generated meeting minutes can be viewed, edited, and shared in real time by users using the output function. Users can easily manage the generated meeting minutes through a dedicated interface.

[0663] As a concrete example, consider a project team meeting. If a participant from a regional area says, "What should we do today?", the server translates this to "What shall we do today?". Furthermore, it quickly identifies which participant made a particular statement and reflects this in the meeting minutes. Using this system, everyone can have a shared understanding while efficiently creating meeting minutes.

[0664] An example of a prompt using a generative AI model is: "Convert the audio data spoken during the meeting into text in real time, identify the speaker, and automatically generate meeting minutes."

[0665] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0666] Step 1:

[0667] The device uses a microphone to collect participants' voices during meetings and business negotiations, converting analog audio into digital audio data. The input is the raw voice of each participant, and the output is prepared as digital audio data. This data is transmitted in real time to a server via the network.

[0668] Step 2:

[0669] The server converts digital audio data received from the terminal into symbolic information using speech analysis tools. Digital audio data is passed to the server as input, and speech recognition APIs such as Google Speech-to-Text are used to convert the audio into text data. The output is text data in string format.

[0670] Step 3:

[0671] The server uses language conversion functionality to convert the converted text data back to the base language. Here, text data is used as input, and specific dialects and expressions are replaced with those in the base language. Standardized text data is generated as output.

[0672] Step 4:

[0673] The server identifies the speaker from the audio data using biometric recognition means. In this process, audio feature data is used as input, and the speaker is identified by matching it with a pre-registered voiceprint database. The output is speaker identification information accompanying the text data.

[0674] Step 5:

[0675] The server uses a recording generation mechanism to generate meeting minutes for each identified speaker. Standardized text data and speaker information are used as input data, which are combined to generate organized meeting minutes. The output is a completed meeting minute.

[0676] Step 6:

[0677] Users can view meeting minutes provided by the server in real time through an output function. The input is meeting minutes data from the server, which users can edit and share with other participants via a web application or dedicated software. The output is the meeting minutes and related information displayed in a user-friendly format.

[0678] (Application Example 1)

[0679] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] In commercial environments where multiple languages ​​or regional dialects are spoken, there is a need to facilitate communication between customers and staff. Traditionally, insufficient communication due to different languages ​​or dialects has led to problems with decreased customer satisfaction.

[0681] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0682] In this invention, the server includes terminal means for receiving acoustic data, speech recognition means for converting the received acoustic data into text data, and display means for visually outputting the translated text data to a display device. This enables real-time assistance for voice communication with customers and facilitates smooth mutual understanding in a multilingual environment.

[0683] "Terminal means for receiving acoustic data" refers to a device that acquires sound as digital information.

[0684] "Speech recognition means" refers to technology that analyzes acoustic data and converts it into text information.

[0685] "Language conversion means" refers to technologies for converting text information from another language or dialect into a standard language.

[0686] "Voiceprint recognition means" refers to technology for identifying individual speakers from their voices.

[0687] A "recording means" is a function that stores the converted data and saves it as a record.

[0688] "Output means" refers to a device that presents the created record in physical or digital format.

[0689] "Display means" refers to a device that visually displays translated text information.

[0690] "Communication methods" refer to technologies for exchanging data between devices.

[0691] The system implementing this invention is equipped with technology for collecting audio data and processing the information in real time. The terminal primarily receives audio data and transmits it to the server. The server converts the received audio data into text data using speech recognition technology. The converted text data is then translated into a standard language by a language conversion means. Advanced natural language processing technology is used for this translation.

[0692] The server uses voiceprint recognition technology to identify the speaker. Specifically, it organizes each utterance by speaker by comparing it with a pre-registered voiceprint database. This generates a record categorized by speaker. The generated record is provided to the user through an output device.

[0693] This system also includes a display mechanism, allowing translated text data to be visually output. In a real-world example, when using smart glasses to serve foreign customers in a shop, the customer's speech is translated into standard Japanese in real time and displayed on the glasses' screen. This enables smooth communication between staff and customers.

[0694] For example, if a French-speaking tourist visits a souvenir shop and asks about the price of an item, the staff member's smart glasses will display the translated text, "This keychain costs 500 yen." Through this concrete example, users can obtain information in real time and receive a quick response. This system is an effective means of improving customer satisfaction.

[0695] Examples of prompts for a generative AI model include the following:

[0696] "Develop a real-time voice translation system using smart glasses for use in physical stores. The system will collect voice data via a microphone, send it to a server, perform voice recognition and translation, and display the results on the glasses' display. The goal is smooth communication with customers."

[0697] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0698] Step 1:

[0699] The terminal receives audio data in a commercial environment. Specifically, it collects customer speech via microphones within a store and prepares the data in digital format. The input is an audio signal, and the output is digital audio data.

[0700] Step 2:

[0701] Digital audio data is transferred from the terminal to the server. The server inputs the received audio data into the speech recognition system. The input here is digital audio data, and the server prepares to process that data.

[0702] Step 3:

[0703] The server uses speech recognition technology to convert audio data into text data. Specifically, it uses the Google Cloud Speech-to-Text API to analyze speech and generate text information. The input is digital audio data, and the output is text data.

[0704] Step 4:

[0705] The server uses language conversion tools to translate text data into a reference language. It uses the Google Cloud Translation API to convert input text into the reference language. The input is the converted text data, and the output is the text data in the reference language.

[0706] Step 5:

[0707] The server uses voiceprint recognition technology to identify the speaker. It compares the received audio data with a pre-registered voiceprint database to identify the speaker. The input is the audio data and the voiceprint database, and the output is information about the identified speaker.

[0708] Step 6:

[0709] The server sends the translated text data to a display device via an output mechanism, and outputs it visually. Specifically, the text data is displayed on the smart glasses' screen to provide information to staff. The input is text data in a reference language, and the output is a visual display.

[0710] Step 7:

[0711] The user reviews the displayed information and responds to the customer. Based on the customer's questions, they understand the text display and provide appropriate information. Input is visual display information, and output is the customer interaction conversation.

[0712] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0713] This invention combines a speech recognition system that transcribes participants' speech in real time and converts it into standard Japanese during meetings and business negotiations with a newly developed emotion engine. This makes it possible to recognize the emotional state of participants and organize and record the flow of discussion and important points while considering the emotional context.

[0714] This system first has a terminal collect the voices of multiple participants through a microphone and send the converted audio data to a server. The server then uses a speech recognition engine to convert this audio data into text data. During this process, the server uses a dialect dictionary to perform standard language conversion appropriate for the text data.

[0715] Furthermore, the server uses voiceprint recognition technology to identify speakers and organizes each statement by speaker. By matching it with a voiceprint database, speakers can be identified, and this is used to organize the statements by speaker.

[0716] In addition, an emotion engine installed on the server estimates the emotional state of each speaker through text data and speech analysis. The emotion engine adds this emotional information to the meeting minutes and further prioritizes the importance of summaries and action items based on the intensity of the emotions.

[0717] Users can access meeting minutes provided by the server, which include details of each speaker's remarks along with their estimated emotional state. For example, if participant A says, "I have some concerns about this new proposal," during a meeting, the server recognizes A's emotional state from the audio and text, and this emotional information is included in the meeting minutes. This has the advantage of providing the meeting facilitator with information to plan discussions and follow-ups to address A's concerns.

[0718] Thus, the present invention generates meeting minutes that also take emotional nuances into account, improving the quality of meetings and effectively preventing misunderstandings and communication breakdowns among participants.

[0719] The following describes the processing flow.

[0720] Step 1:

[0721] The terminal captures audio from meetings and business negotiations via its microphone and converts it into digital audio data in real time. This audio data is immediately transmitted to the server.

[0722] Step 2:

[0723] The server prepares the audio data received from the terminal for input into the speech recognition engine. The data is processed in batches and converted into text by the speech recognition engine.

[0724] Step 3:

[0725] The server uses a dialect dictionary to convert the text data obtained through speech recognition into standard Japanese. By replacing dialects and expressions with standard Japanese, the server facilitates overall comprehension.

[0726] Step 4:

[0727] The server uses voiceprint recognition technology to identify the individual speaker. It compares features extracted from the audio data with a voiceprint database and adds speaker information to the text data.

[0728] Step 5:

[0729] The server uses an emotion engine to estimate the speaker's emotions from text data and voice features. The emotion engine identifies specific emotional states and associates that information with the text.

[0730] Step 6:

[0731] The server generates meeting minutes by adding sentiment information to text data organized by speaker. The minutes include detailed records of what was said, who said it, and the estimated sentiment.

[0732] Step 7:

[0733] The server extracts summaries and action items from the generated meeting minutes and prioritizes them based on sentiment information. This information is then highlighted within the meeting minutes.

[0734] Step 8:

[0735] Users can view meeting minutes provided by the server in real time. While reviewing the minutes, users can understand the flow of the discussion by referring to sentiment information and take necessary actions.

[0736] (Example 2)

[0737] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0738] In meetings and business negotiations, it is necessary not only to accurately record participants' statements in real time, but also to understand the emotional state of the speakers and organize and record the flow of the discussion and important points while taking emotional context into account. However, conventional speech recognition systems do not adequately identify emotions or convert dialects to standard Japanese, leading to problems such as misunderstandings and communication losses among participants.

[0739] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0740] In this invention, the server includes speech recognition means for converting speech information into text information, language conversion means for converting the converted text information into a common language, and emotion prediction means for analyzing the text information and speech information to predict emotional states. This makes it possible to record the content of meetings while taking into account emotional nuances, thereby facilitating smooth communication among participants.

[0741] "Audio information" refers to sound data collected from participants in meetings or business negotiations using terminal devices.

[0742] "Textual information" refers to digital text data converted from audio information by speech recognition technology.

[0743] "Common language" refers to a standardized form of language expression that does not include dialects.

[0744] "Identification means" refers to technology that identifies a speaker by analyzing the speaker's voice characteristics from textual information.

[0745] "Emotion prediction means" refers to technology that predicts and identifies the emotional state of a speaker by analyzing textual and auditory information.

[0746] "Recording means" refers to a means of organizing the content of a conversation for each identified speaker and generating a meeting record.

[0747] A "dialect dictionary" refers to a language database used to translate regionally specific linguistic expressions into standard Japanese.

[0748] One embodiment of the present invention is a system that records the statements of participants in meetings and business negotiations in real time and generates meeting minutes that take into account emotional context. This system handles everything from collecting audio data to sentiment analysis and generating meeting minutes in an integrated manner.

[0749] The device uses a microphone to collect audio and converts participants' speech into digital audio data. The audio data collected by the device is immediately transmitted to a server via the network. Noise reduction technology is employed during this process to maintain audio quality.

[0750] The server first converts the received audio data into text data using a speech recognition engine. A speech recognition API is used for this process, generating highly accurate textual information. Furthermore, the converted textual information is standardized to a common language through a language conversion function. The software used in this process incorporates a dialect dictionary.

[0751] Next, the server uses an identification mechanism to identify the speaker from the text information. This mechanism uses voiceprint data to identify the speaker and organizes the statements by speaker. Furthermore, the emotion prediction engine on the server analyzes the text data and voice characteristics to predict the speaker's emotional state.

[0752] Ultimately, users can access the meeting minutes generated by the server and review the details of each statement, including emotional nuances. For example, if participant A says, "I have some concerns about this new proposal," in order to ensure the meeting runs smoothly, the server can reflect A's concern in the minutes. This allows users to assess concerns about the proposal and use that information to plan follow-up.

[0753] A possible prompt would be, "What methods and technologies should be used to create a system that records participants' statements and emotions in real time during a meeting and generates meeting minutes based on the estimated emotions?" This system would improve communication among participants and enhance the quality of the meeting.

[0754] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0755] Step 1:

[0756] The terminal uses a microphone to collect speech from meeting participants in real time. The input is an analog signal of raw sound, which is converted into digital audio data. The converted audio data is then transmitted to a server via the network. Noise reduction technology is used to improve the quality of the audio data.

[0757] Step 2:

[0758] The server converts digital audio data received from the terminal into text information using a speech recognition engine. The input is digital audio data, and the output is text information. Speech recognition automatically transcribes spoken content into text, generating highly accurate text data.

[0759] Step 3:

[0760] The server applies a language conversion function to convert the converted character information into the standard language. The input is text data, and the output is character information in the standard language. By using a pre-built dialect dictionary, it performs the task of standardizing regional expressions.

[0761] Step 4:

[0762] The server identifies the speaker using identification methods based on the text information converted to a common language. The input here is the text information in a common language, and the output is the identification information of each speaker. Voiceprint data is used to identify speakers with high accuracy and organize statements by speaker.

[0763] Step 5:

[0764] The server uses an emotion prediction engine to analyze the speaker's emotional state based on identified textual information. The input is standard language textual information with speaker information, and the output is text information including emotional state. It analyzes the features of the audio and text data to identify the emotional context.

[0765] Step 6:

[0766] Users access meeting minutes generated from the server, which include sentiment information. The generated minutes provide details of each statement, incorporating emotional nuances. This makes it easier for users to effectively manage meeting content and plan subsequent actions.

[0767] (Application Example 2)

[0768] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0769] In brick-and-mortar retail settings, understanding customer emotions and providing accurate information and suggestions has been difficult with conventional technologies. In particular, there was a lack of mechanisms to quickly detect and address customer anxiety or dissatisfaction. Therefore, there was a need for effective customer service to improve customer satisfaction.

[0770] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0771] In this invention, the server includes an information processing means for receiving audio data, an audio recognition means for converting the received audio data into text data, and an emotion analysis means for estimating emotional states from the audio and text data and adding emotional information to the generated meeting minutes. This makes it possible to analyze the emotions of customers in physical stores in real time and provide optimal responses and suggestions.

[0772] "Audio data" refers to a data format in which audio has been digitized.

[0773] "Information processing means" refers to devices and systems for appropriately processing received data.

[0774] "Speech recognition means" refers to a technology or device that analyzes speech data and converts it into text data.

[0775] "Text data" refers to data in sentence format that has been converted by speech recognition technology.

[0776] "Standard language" refers to the standard form of language used in a particular region or group.

[0777] "Language conversion means" refers to a technology or device for changing one language format to another.

[0778] "Means of personal identification" refers to technologies or devices that identify individuals based on their voice or other characteristics.

[0779] "Recording means" refers to a technology or device for storing organized information.

[0780] "Emotion analysis means" refers to a technology or device that estimates a speaker's emotions using text data or audio data.

[0781] "Information provision means" refers to the technology or device that outputs the generated data or information.

[0782] A "regional language dictionary" is a database that compiles the characteristics and vocabulary of a language in a specific region.

[0783] A "voice characteristics database" is a database that stores voice characteristics for the purpose of identifying individuals.

[0784] This invention provides a voice recognition system for enhancing customer interaction in a retail environment. This system includes the following components:

[0785] First, a terminal, acting as an information processing device, is installed in the store to receive customer voice data. This voice data is digitized and sent to a server. The server uses speech recognition to convert the voice data into text data. The software used here is the speech_recognition library. This converted text data is then converted back into standard Japanese. In this process, a regional language dictionary is used as a language conversion tool.

[0786] Next, using a voice characteristics database, the speaker is identified by a personal identification means, and meeting minutes are generated through a recording means. In addition, an emotion analysis means is used to estimate the customer's emotional state from the voice and text data, and based on this, emotion information is added to the meeting minutes. For this purpose, software to operate the emotion engine is required.

[0787] Ultimately, users can access meeting minutes and sentiment information generated from the server through the information provision mechanism. This embodiment makes it possible to analyze customer sentiment in real time in stores and use that information to improve customer service.

[0788] To give a concrete example, if a customer says, "I want to know more about this product," the system will sense the customer's interest and automatically initiate a process to provide relevant additional information. In this way, customer service that takes emotions into account is achieved.

[0789] An example of a prompt might be: "When a customer asks a question about a product, convert the question into text, analyze its sentiment, and generate an appropriate response based on that."

[0790] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0791] Step 1:

[0792] The terminal collects customer voices using a microphone in the store, digitizes the audio data, and sends it to a server. The input is the customer's raw voice data, and the output is digitized audio data. When the terminal collects the audio data, it uses noise cancellation technology to suppress ambient noise.

[0793] Step 2:

[0794] The server converts received digital audio data into text data using speech recognition. The input is digital audio data, and the output is the converted text data. The server uses the speech_recognition library to analyze the audio data, perform word recognition, and convert it to text.

[0795] Step 3:

[0796] The server converts text data into standard language using a language conversion mechanism. The input is text data that may contain dialects, and the output is text data converted into standard language. Here, the server consults a regional language dictionary and replaces region-specific words and expressions with standard language equivalents.

[0797] Step 4:

[0798] The server uses personal identification methods to identify speakers by referring to a voice characteristics database. Input consists of text data and audio data converted to standard language, and output is the identification information of the identified speaker. The server analyzes the voiceprint and compares it with an existing database to identify the speaker.

[0799] Step 5:

[0800] The server uses emotion analysis tools to estimate the customer's emotional state from text and audio data. The input is the speaker's text and audio data, and the output is estimated emotion information. The emotion engine analyzes emotions based on the tone and keywords of the speech, and as a result identifies emotions such as joy and anxiety.

[0801] Step 6:

[0802] The server adds sentiment information to the meeting minutes for each identified speaker and stores it through a recording mechanism. The input is sentiment information and speaker information, and the output is the meeting minutes with the sentiment information added. The server generates the meeting minutes in a structured format and stores them in an appropriate format for later access.

[0803] Step 7:

[0804] Users access meeting minutes and sentiment information generated from the server through an information provision mechanism. Input is the user's access request, and output is specific meeting minutes and sentiment information. Users can view the necessary information through the store's interface and use it in their interactions.

[0805] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0806] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0807] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0808] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0809] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0810] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0811] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0812] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0813] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0814] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0815] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0816] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0817] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0818] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0819] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0820] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0821] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0822] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0823] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0824] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0825] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0826] The following is further disclosed regarding the embodiments described above.

[0827] (Claim 1)

[0828] A terminal means for receiving audio data,

[0829] A speech recognition means that converts received audio data into text data,

[0830] A language conversion means for converting converted text data into a standard language,

[0831] A voiceprint recognition means for identifying the speaker of text data,

[0832] A recording means for generating minutes for each identified speaker,

[0833] An output method for outputting the generated meeting minutes,

[0834] A system that includes this.

[0835] (Claim 2)

[0836] The system according to claim 1, wherein the speech recognition means recognizes a dialect from speech data using a dialect dictionary.

[0837] (Claim 3)

[0838] The system according to claim 1, which uses voiceprint recognition means to compare the characteristics of an identified speaker with a voiceprint database.

[0839] "Example 1"

[0840] (Claim 1)

[0841] Information processing device means for receiving audio information,

[0842] A speech analysis means that converts received speech information into symbolic information,

[0843] A language conversion function means for converting converted symbolic information into a reference language,

[0844] A biometric information recognition means for identifying the speaker of symbolic information,

[0845] A recording generation means that generates meeting minutes for each identified speaker,

[0846] An output function means for outputting the generated meeting minutes,

[0847] A system that includes this.

[0848] (Claim 2)

[0849] The system according to claim 1, wherein the speech analysis means recognizes a regional language from speech information using a language dictionary.

[0850] (Claim 3)

[0851] The system according to claim 1, which uses biometric information recognition means to compare the characteristics of an identified speaker with a biometric information database.

[0852] "Application Example 1"

[0853] (Claim 1)

[0854] A terminal means for receiving audio data,

[0855] A speech recognition means that converts received audio data into text data,

[0856] A language conversion means for converting converted text data into a base language,

[0857] A voiceprint recognition means for identifying the speaker of text data,

[0858] A recording means that generates a record for each identified speaker,

[0859] An output means for outputting the generated record,

[0860] A display means for visually outputting translated text data to a display device,

[0861] A communication means for assisting voice communication through visual displays,

[0862] A system that includes this.

[0863] (Claim 2)

[0864] The system according to claim 1, wherein the speech recognition means recognizes a dialect from acoustic data using a dialect dictionary.

[0865] (Claim 3)

[0866] The system according to claim 1, which uses voiceprint recognition means to compare the characteristics of an identified speaker with a voiceprint database.

[0867] "Example 2 of combining an emotion engine"

[0868] (Claim 1)

[0869] A terminal means for receiving voice information,

[0870] A speech recognition means that converts received audio information into text information,

[0871] A language conversion means for converting converted text information into a common language,

[0872] An identification means for identifying the speaker of text information,

[0873] A recording means for generating meeting records for each identified speaker,

[0874] An emotion prediction means that analyzes textual and audio information to predict emotional states,

[0875] An output means for outputting the generated meeting record,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, wherein the speech recognition means recognizes a regional expression from speech information using a dialect dictionary.

[0879] (Claim 3)

[0880] The system according to claim 1, which uses identification means to match the characteristics of an identified speaker with a voice database.

[0881] "Application example 2 of combining emotional engines"

[0882] (Claim 1)

[0883] Information processing means for receiving audio data,

[0884] A speech recognition means that converts received audio data into text data,

[0885] A language conversion means for converting converted text data into a standard language,

[0886] A means of identifying the speaker of text data,

[0887] A recording means for generating minutes for each identified speaker,

[0888] An emotion analysis means that estimates emotional states from audio and text data and adds emotional information to the generated meeting minutes,

[0889] A means of providing information that outputs the generated meeting minutes,

[0890] A system that includes this.

[0891] (Claim 2)

[0892] The system according to claim 1, wherein the speech recognition means recognizes a regional language from speech data using a regional language dictionary.

[0893] (Claim 3)

[0894] The system according to claim 1, which uses personal identification means to match the characteristics of an identified speaker with a voice characteristics database. [Explanation of Symbols]

[0895] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal means for receiving audio data, A speech recognition means that converts received audio data into text data, A language conversion means for converting converted text data into a standard language, A voiceprint recognition means for identifying the speaker of text data, A recording means for generating minutes for each identified speaker, An output method for outputting the generated meeting minutes, A system that includes this.

2. The system according to claim 1, wherein the speech recognition means recognizes a dialect from speech data using a dialect dictionary.

3. The system according to claim 1, which uses voiceprint recognition means to compare the characteristics of an identified speaker with a voiceprint database.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A