system

The system automates the transcription, formatting, and verification of recorded audio to streamline management and ensure accuracy and consistency, addressing the inefficiencies of manual methods.

JP2026062212APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

The management of recorded voices for customer service and administrative service relies on manual transcription and content verification, which is time-consuming and labor-intensive, and lacks accuracy and consistency.

Method used

A system that automates the process by receiving recorded audio, transcribing it into text, formatting the content, verifying its accuracy, and saving it as a history, using speech recognition APIs and user interfaces for verification.

Benefits of technology

This system reduces time and labor, ensures accurate and consistent management of recorded audio, and enables efficient recording of its history.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062212000001_ABST
    Figure 2026062212000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving recorded audio, A means for transcribing the aforementioned recorded audio, A means for formatting the transcribed content into a specific format, A means for verifying the formatted content, Means for saving the confirmed contents, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventionally, the management of recorded voices for customer service and administrative service has relied on manual transcription and content verification, which has the problem of being time-consuming and labor-intensive. In addition, since the accuracy and consistency of transcription and content verification are not guaranteed, there is a possibility of errors and omissions in the content.

Means for Solving the Problems

[0005] The present invention solves the above problems by a system including means for receiving a recorded voice, means for transcribing the recorded voice into text, means for formatting the transcribed content into a specific format, means for verifying the formatted content, and means for storing the verified content. Thereby, the management of recorded voices is automated, it becomes possible to reduce time and labor, and it is also possible to ensure accuracy and consistency.

[0006] "Recorded audio" refers to audio recorded in digital or analog format.

[0007] "Reception" refers to the act of taking in audio data within a system.

[0008] "Transcription" is the process of converting audio data into text data.

[0009] "Formatting" is the process of shaping data into a specific format.

[0010] "Formatting" refers to the act of arranging data according to specific rules or regulations.

[0011] "Verification" is the act of checking whether the entered data is accurate.

[0012] "Storage" is the act of recording data in order to retain it for a long period of time.

[0013] "Means" refers to a device or method for achieving a specific function.

[0014] A "system" is an overall configuration in which multiple devices or methods work together to achieve a specific function. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a headset type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0019] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. This system streamlines the management of recorded audio and enables accurate recording of its history.

[0037] System Overview

[0038] The system consists of the following elements:

[0039] 1. Means for receiving recorded audio

[0040] 2. Methods for transcribing recorded audio into text.

[0041] 3. Means for formatting the transcribed content into a specific format.

[0042] 4. Means of verifying formatted content

[0043] 5. Means for saving the confirmed information

[0044] Means for receiving recorded audio

[0045] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0046] Methods for transcribing recorded audio into text

[0047] The server uses a speech recognition API (for example, Google's Speech Recognition API) to transcribe the recorded audio. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is stored in a temporary variable.

[0048] A method for formatting transcribed content into a specific format.

[0049] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0050] Means of verifying formatted content

[0051] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0052] Means for saving confirmed information

[0053] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt).

[0054] Specific example

[0055] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[0056] ---

[0057] Date and Time: 2023-10-15 14:35:00

[0058] Type: Customer support

[0059] Contents:

[0060] Hello, this is Customer Service. How can I help you today?

[0061] ---

[0062] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[0063] ---

[0064] Date and Time: 2023-10-15 14:35:00

[0065] Type: Customer support

[0066] Contents:

[0067] Hello, this is Customer Service. How can I help you today?

[0068] ---

[0069] This automates the management of recorded audio, saving time and effort, while also ensuring accuracy and consistency.

[0070] The following describes the processing flow.

[0071] Processing flow

[0072] Step 1:

[0073] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[0074] Step 2:

[0075] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0076] Step 3:

[0077] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0078] Step 4:

[0079] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0080] Step 5:

[0081] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0082] (Example 1)

[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] Managing recorded audio requires manual transcription and formatting, which is time-consuming, labor-intensive, and can lead to inaccuracies and inconsistencies. Furthermore, verifying and correcting transcribed content is cumbersome, highlighting the need for an efficient management system. Additionally, a centralized method for managing and preserving the recording history is essential.

[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0086] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio using a speech recognition API, means for formatting the transcribed content into a specific format, means for confirming the formatted content through a user interface, and means for saving the confirmed content to a history file. This streamlines the automatic transcription and formatting of recorded audio, enabling accurate and consistent management.

[0087] "Recorded audio" refers to a file containing audio data recorded in a specific manner.

[0088] "Means of receiving" refers to the function that allows the server to receive recorded audio files from users.

[0089] A "speech recognition API" is a program that provides an interface for converting speech data into text data.

[0090] "Transcribing methods" refer to the function of converting recorded audio into text data using a speech recognition API.

[0091] "Methods for formatting" refer to functions for organizing transcribed text data into a specific format.

[0092] A "user interface" is software that provides screens and input methods for users to interact with a system.

[0093] A "history file" is a text file used to save the confirmed information.

[0094] Modes for carrying out the invention

[0095] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. By using this system, the management of recorded audio can be streamlined, and accurate history recording can be achieved.

[0096] System Overview

[0097] The system consists of the following elements:

[0098] 1. Means for receiving recorded audio

[0099] 2. Methods for transcription

[0100] 3. Means for formatting the transcribed content into a specific format.

[0101] 4. Means of verifying formatted content

[0102] 5. Means for saving the confirmed information to a history file.

[0103] Hardware and software configuration

[0104] server

[0105] The server receives HTTP requests from users, receives and saves recorded audio files. The Flask framework in Python is used for this receiving process.

[0106] Specific receiving operation

[0107] When a user uploads the audio file customer_call.wav using a browser, the server receives it and saves it to a specific directory. This directory is / uploads / .

[0108] server

[0109] The server uses Google's speech recognition API to convert recorded audio into text data. Here, the Python pydub library is used to read the audio file and send a request to the API.

[0110] Specific transcription actions

[0111] The server sends the audio file `customer_call.wav` to the Google Speech Recognition API and receives the text data. The received text data is stored in a temporary variable `transcription_text`. For example, text such as "Hello, this is Customer Service. How can I help you today?" is generated.

[0112] terminal

[0113] The terminal receives text data sent from the server and formats it into a specific format. This format includes the recording date and time, the type of recording, and the transcribed content.

[0114] Specific formatting behavior

[0115] The terminal generates the following format based on the character data received from the server.

[0116] ---

[0117] Date and Time: 2023-10-15 14:35:00

[0118] Type: Customer support

[0119] Contents:

[0120] Hello, this is Customer Service. How can I help you today?

[0121] ---

[0122] The formatted data is stored in the temporary variable `formatted_text`.

[0123] User

[0124] The user reviews the formatted content on their device screen. They verify the content is accurate, make corrections if necessary, and finally approve it. This review process utilizes a web browser and a dedicated web application.

[0125] Specific verification actions

[0126] The user views the formatted content in their browser and clicks the confirmation button to check the content. If corrections are needed, they use the input fields to make the necessary changes and then click the "Approve" button.

[0127] server

[0128] The server receives the approved formatted content and saves it to the history file. The saved history is appended to a text file (e.g., transcription_history.txt).

[0129] Specific saving actions

[0130] The server receives the approved format content and appends it to the history file as follows:

[0131] ---

[0132] Date and Time: 2023-10-15 14:35:00

[0133] Type: Customer support

[0134] Contents:

[0135] Hello, this is Customer Service. How can I help you today?

[0136] ---

[0137] Example of a prompt

[0138] Examples of prompt messages in this system are as follows:

[0139] Please transcribe the audio file customer_call.wav using a speech recognition API and format it into the following format:

[0140] ---

[0141] Date and Time: YYYY-MM-DD HH:MM:SS

[0142] Type: Customer support

[0143] Contents:

[0144] <Transcripted text>

[0145] ---

[0146] In this way, the automatic transcription and formatting of recorded audio is streamlined, and it becomes possible to manage it with accuracy and consistency.

[0147] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0148] Step 1: Receiving the recorded audio

[0149] The server receives the recorded audio file sent as an HTTP request from the user.

[0150] Input: A recorded audio file uploaded by the user (e.g., customer_call.wav).

[0151] Data processing: The server temporarily stores the received file in memory and then moves the file to the specified directory.

[0152] Output: Audio files saved in the specified directory on the server (e.g., / uploads / ).

[0153] Specific operation: The user uploads an audio file through a browser form, the server receives it, and saves it to the / uploads / directory.

[0154] Step 2: Transcribe the recorded audio.

[0155] The server sends the stored audio file to a speech recognition API, where it is converted into text data.

[0156] Input: Audio file stored on the server (e.g., / uploads / customer_call.wav).

[0157] Data processing: The server uses the Python pydub library to read the audio file and send a request to the speech recognition API.

[0158] Output: Text data returned from the speech recognition API (e.g., "Hello, this is customer service. How can I help you today?").

[0159] Specific operation: The server sends the audio file to the Google Speech Recognition API and saves the returned text data to a temporary variable called `transcription_text`.

[0160] Step 3: Format of the transcript

[0161] The device receives text data and formats it into a specific format.

[0162] Input: Text data sent from the server (e.g., "Hello, this is Customer Service. How can I help you today?").

[0163] Data processing: Apply a format to the text data that includes the recording date and time, recording type, and transcript content.

[0164] Output: Formatted text data (e.g., in the format shown below).

[0165] ---

[0166] Date and Time: 2023-10-15 14:35:00

[0167] Type: Customer support

[0168] Contents:

[0169] Hello, this is Customer Service. How can I help you today?

[0170] ---

[0171] Specific operation: The terminal receives text data from the server, applies a format to it, shapes it, and saves it to a temporary variable called `formatted_text`.

[0172] Step 4: Check the format content

[0173] The user confirms the formatted content on the device screen.

[0174] Input: Formatted text data displayed on the terminal.

[0175] Data processing: The user reviews the content and makes corrections if necessary. The corrected data is then formatted again.

[0176] Output: The final text data approved or modified by the user.

[0177] Specific action: The user reviews the formatted content using a browser, makes corrections as needed, and clicks the "Approve" button.

[0178] Step 5: Save the confirmation details

[0179] The server saves the format content approved by the user to a history file.

[0180] Input: User-approved, formatted text data.

[0181] Data processing: Approved text data is appended to the history file (e.g., transcription_history.txt).

[0182] Output: Formatted text saved in the history file.

[0183] Specific operation: The server receives the approved text data and appends it to the history file in the following format.

[0184] ---

[0185] Date and Time: 2023-10-15 14:35:00

[0186] Type: Customer support

[0187] Contents:

[0188] Hello, this is Customer Service. How can I help you today?

[0189] ---

[0190] (Application Example 1)

[0191] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0192] As autonomous vehicles become more widespread, there is a growing need to efficiently manage, record, and store the voice interactions between passengers and vehicles as a history. However, existing systems rely on manual management of recorded voices, which is time-consuming and labor-intensive, and makes it difficult to maintain accuracy and consistency. Furthermore, there is a lack of adequate systems for effectively utilizing voice interactions within vehicles after the fact.

[0193] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0194] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for verifying the formatted content, means for saving the verified content, and, if the recorded audio is voice dialogue in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors, and means for displaying the transcribed content on the vehicle's display device or a portable information terminal. This automates the recording, transcription, formatting, verification, and saving of voice dialogue in autonomous vehicles, thereby improving the quality of operation management and customer service.

[0195] An "autonomous vehicle" is a vehicle that can operate autonomously without driver intervention.

[0196] An "acoustic sensor" is a device used to detect and record sound, such as a microphone installed in a vehicle.

[0197] A "portable information terminal" refers to a portable electronic device, such as a smartphone or tablet.

[0198] "Transcription" refers to the process of converting recorded audio into text data.

[0199] A "speech recognition API" is a programmatic interface for converting speech data into text data, and includes cloud-based and local-based services.

[0200] "Formatting to a specific format" refers to organizing and arranging transcribed text data according to a predetermined format and layout.

[0201] "Vehicle display devices" refer to display devices such as monitors and touch panels installed inside a vehicle.

[0202] "User interface" refers to the means by which a device or system and a user exchange information, and includes displays and input devices.

[0203] "Means of verification" refers to processes or devices that allow users to check formatted text content and make corrections or approvals as needed.

[0204] "Means of preservation" refers to the processes and systems used to record and retain verified text data in databases or files.

[0205] This invention relates to a system for recording, transcribing, and formatting voice dialogue within an autonomous vehicle as a history. This system streamlines the management of recorded audio and ensures accurate history recording.

[0206] System Overview

[0207] The system consists of the following elements:

[0208] 1. Means for receiving recorded audio

[0209] 2. Methods for transcribing recorded audio into text.

[0210] 3. Means for formatting the transcribed content into a specific format.

[0211] 4. Means of verifying formatted content

[0212] 5. Means for saving the confirmed information

[0213] 6. If the recorded audio is for voice interaction in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors.

[0214] 7. Means for displaying the transcribed content on a vehicle display device or portable information terminal.

[0215] Means for receiving recorded audio

[0216] Acoustic sensors (such as in-vehicle microphones) installed in autonomous vehicles record passenger voices. The recorded voice data is transmitted to the onboard computer and temporarily stored. This prepares the system for processing voice interactions with passengers.

[0217] Methods for transcribing recorded audio into text

[0218] The server uses a speech recognition API (e.g., a speech recognition cloud service) to transcribe the recorded audio. Specifically, it reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is temporarily stored.

[0219] A method for formatting transcribed content into a specific format.

[0220] The device formats the transcribed text data into a specific format. Specifically, it includes the recording date and time, the type of recording (customer interaction or passenger interaction), and the transcribed text content. The formatted text is temporarily stored.

[0221] Means of verifying formatted content

[0222] The user reviews the formatted text content via the vehicle's display device or a mobile device. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0223] Means for saving confirmed information

[0224] The server saves the formatted text approved by the user to a history file. Specifically, it appends the text to the history file (for example, transcription_history.txt).

[0225] Specific example

[0226] For example, suppose a passenger's voice interaction inside an autonomous vehicle is recorded and saved as a file named customer_interaction.wav. Let's say this audio file was uploaded on October 15, 2023 at 14:35:00. The system processes this audio file and generates text formatted as follows:

[0227] ---

[0228] Date and Time: 2023-10-15 14:35:00

[0229] Contents:

[0230] Hello, this is an autonomous vehicle. Where are you going today?

[0231] ---

[0232] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[0233] ---

[0234] Date and Time: 2023-10-15 14:35:00

[0235] Contents:

[0236] Hello, this is an autonomous vehicle. Where are you going today?

[0237] ---

[0238] This automates the management of voice interaction within autonomous vehicles, saving time and effort while ensuring accuracy and consistency.

[0239] Example prompts for generative AI models

[0240] "Create a sample Python program that records voice conversations with customers, transcribes them using a voice recognition cloud service, and saves them in a specific format."

[0241] As described above, this system automates the recording, transcription, formatting, verification, and saving of voice conversations within autonomous vehicles, thereby improving the quality of operational management and customer service.

[0242] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0243] Step 1:

[0244] Receiving recorded audio

[0245] The server receives recorded audio data through acoustic sensors installed in the autonomous vehicle. Specifically, a microphone captures the passenger's voice and sends that audio data to the server. The input is audio data, and the output is a temporarily stored audio file.

[0246] Step 2:

[0247] Transcription of recorded audio

[0248] The server uses a speech recognition API to convert the received audio data into text data. Specifically, a Python script is used to read the audio file and send a request to the speech recognition API. The input is an audio file, and the output is the transcribed text data.

[0249] Step 3:

[0250] Formatting of the transcribed content

[0251] The terminal formats the transcribed text data into a specific format. Specifically, it adds the recording date and time, and the type of recording (e.g., customer service, passenger interaction) to the text data and compiles it into a predetermined format. The input is the text data and the current date and time, and the output is the formatted text.

[0252] Step 4:

[0253] Checking the formatted content

[0254] The user reviews the formatted content via the vehicle's display device or a mobile device. Specifically, the formatted text is displayed on the screen of a display or smartphone, allowing the user to review the content and make corrections as needed. The input is the formatted text, and the output is the reviewed and corrected text.

[0255] Step 5:

[0256] Saving of confirmed information

[0257] The server saves the formatted text confirmed by the user to a history file. Specifically, it appends text data to the history file (e.g., transcription_history.txt). The input is the confirmed text data, and the output is the updated history file.

[0258] Step 6:

[0259] Managing history files

[0260] The server periodically manages history files, creates backups as needed, and archives old data. Specifically, it uses scripts to create file backups and manage file sizes. The input is history files, and the output is organized history data.

[0261] The above outlines the specific processing steps of the program, detailing the input, data processing, and output in each step. This process efficiently records, transcribes, formats, verifies, and saves voice conversations within autonomous vehicles.

[0262] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0263] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[0264] System Overview

[0265] The system consists of the following elements:

[0266] 1. Means for receiving recorded audio

[0267] 2. Methods for transcribing recorded audio into text.

[0268] 3. Means for formatting the transcribed content into a specific format.

[0269] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[0270] 5. Means of verifying formatted content

[0271] 6. Means for saving the confirmed information

[0272] Means for receiving recorded audio

[0273] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0274] Methods for transcribing recorded audio into text

[0275] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0276] A method for formatting transcribed content into a specific format.

[0277] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0278] An emotion engine that recognizes user emotions from transcribed content.

[0279] The emotion engine installed on the server analyzes the text data and voice data that have been transcribed and recognizes the user's emotion. The recognized emotion information is included in the text data and added to the format.

[0280] Means for checking the formatted content

[0281] The user checks the text content formatted on the terminal screen. The user checks that the formatted content is accurate, makes corrections if necessary, and then approves it.

[0282] Means for saving the confirmed content

[0283] The server saves the formatted text approved by the user in the history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server checks the saved content and ensures that it is properly saved.

[0284] Specific example

[0285] For example, assume that the recorded voice file customer_call.wav for customer service is uploaded at 2023-10-15 14:35:00. This recorded voice file is processed by the system, and the following formatted text is generated:

[0286] ---

[0287] Date and time: 2023-10-15 14:35:00

[0288] Type: Customer service

[0289] Content:

[0290] Hello, this is customer service. What can I do for you today?

[0291] Emotion: Joy

[0292] ---

[0293] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[0294] ---

[0295] Date and Time: 2023-10-15 14:35:00

[0296] Type: Customer support

[0297] Contents:

[0298] Hello, this is Customer Service. How can I help you today?

[0299] Emotion: Joy

[0300] ---

[0301] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0302] The following describes the processing flow.

[0303] Processing flow

[0304] Step 1:

[0305] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[0306] Step 2:

[0307] The server uses a speech recognition API (e.g., Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is saved in a temporary variable.

[0308] Step 3:

[0309] The terminal formats the transcribed text data into a specific format. Specifically, it formats it into a format that includes the recording date and time, the type of recording (customer response or administrative response), and the transcribed text content. The formatted text is saved in a temporary variable for subsequent processing.

[0310] Step 4:

[0311] The emotion engine installed on the server analyzes the transcribed text data and audio data to recognize the user's emotion. The emotion engine uses natural language processing technology to extract emotions (e.g., joy, anger, sadness, etc.) from the text data. The recognized emotion information is added to the text data.

[0312] Step 5:

[0313] The user checks on the terminal screen the content combined with the formatted text data and the recognized emotion information. The user confirms that the content is accurate, makes corrections if necessary, and then approves it.

[0314] Step 6:

[0315] The server saves the formatted text and emotion information approved by the user in a history file. Specifically, the server appends the formatted text to a history file (e.g., transcription_history.txt). The server checks the saved content and ensures that it is properly saved.

[0316] A concrete example of the process is as follows:

[0317] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 at 14:35:00. This recording file is processed by the system, generating text and sentiment information in the following format:

[0318] ---

[0319] Date and Time: 2023-10-15 14:35:00

[0320] Type: Customer support

[0321] Contents:

[0322] Hello, this is Customer Service. How can I help you today?

[0323] Emotion: Joy

[0324] ---

[0325] After the user reviews and approves this formatted content and sentiment information, the server saves it to a history file. The saved content will look like this:

[0326] ---

[0327] Date and Time: 2023-10-15 14:35:00

[0328] Type: Customer support

[0329] Contents:

[0330] Hello, this is Customer Service. How can I help you today?

[0331] Emotion: Joy

[0332] ---

[0333] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0334] (Example 2)

[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0336] Conventional audio recording management systems often involve manual transcription, formatting, and emotional analysis of recordings, which is time-consuming, labor-intensive, and can lack accuracy. Furthermore, the lack of a function to analyze and record changes in emotion makes detailed conversation history management difficult.

[0337] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing it, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, and means for saving the confirmed content. This enables automatic transcription of recorded audio, formatting, emotion recognition, and history saving.

[0338] "Recorded audio" refers to audio data uploaded by users.

[0339] "Means of receiving" refers to the function by which a server receives audio data from a user and temporarily stores it.

[0340] "Transcribing" refers to a function that analyzes recorded audio and converts it into text data.

[0341] A "speech recognition API" is an interface for external services that convert speech data into text data.

[0342] "Methods for formatting" refer to functions that arrange text data into a specific format.

[0343] "Means of recognizing emotions" refers to a function that analyzes transcribed text data and audio data to identify the user's emotions.

[0344] A "sentiment analysis engine" is a software component used to analyze emotions from text and audio data.

[0345] A "means of verification" refers to a function that allows users to view formatted content and make corrections or approvals.

[0346] The "means of saving" refer to the function of saving the verified text data to a history file.

[0347] A "user interface" refers to the screens and control devices that allow a user to interact with and operate a system.

[0348] Defining it in this way clarifies the role of each function in the system.

[0349] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[0350] System Overview

[0351] The system consists of the following elements:

[0352] 1. Means for receiving recorded audio

[0353] 2. Methods for transcribing recorded audio into text.

[0354] 3. Means for formatting the transcribed content into a specific format.

[0355] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[0356] 5. Means of verifying formatted content

[0357] 6. Means for saving the confirmed information

[0358] Means for receiving recorded audio

[0359] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0360] Methods for transcribing recorded audio into text

[0361] The server uses a speech recognition API (for example, a common speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0362] A method for formatting transcribed content into a specific format.

[0363] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (e.g., customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0364] An emotion engine that recognizes user emotions from transcribed content.

[0365] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[0366] Means of verifying formatted content

[0367] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0368] Means for saving confirmed information

[0369] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0370] Specific example

[0371] For example, suppose a recorded audio file of a customer interaction, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recorded audio file is processed by the system, and text formatted as follows is generated:

[0372] ---

[0373] Date and Time: 2023-10-15 14:35:00

[0374] Type: Customer support

[0375] Contents:

[0376] Hello, this is Customer Service. How can I help you today?

[0377] Emotion: Joy

[0378] ---

[0379] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[0380] ---

[0381] Date and Time: 2023-10-15 14:35:00

[0382] Type: Customer support

[0383] Contents:

[0384] Hello, this is Customer Service. How can I help you today?

[0385] Emotion: Joy

[0386] ---

[0387] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0388] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0389] Step 1:

[0390] The server receives the recorded audio file and stores it temporarily.

[0391] Input: Audio file sent by the user.

[0392] Specific operation: The user uploads an audio file from their device, and the server receives and temporarily stores the file.

[0393] Output: Temporarily saved audio file.

[0394] Step 2:

[0395] The server transcribes the audio file into text.

[0396] Input: Temporarily saved audio file.

[0397] Specific operation: The server sends an audio file to a speech recognition API (a common speech recognition service) and retrieves text data.

[0398] Output: Transcripted text data.

[0399] Step 3:

[0400] The device formats the transcribed text data into a specific format.

[0401] Input: Transcribed text data.

[0402] Specific operation: The device acquires text data and formats it into a format that includes the recording date and time, recording type, and transcript content.

[0403] Output: Formatted text data.

[0404] Step 4:

[0405] The server uses an emotion engine to recognize the user's emotions from the transcribed text data.

[0406] Input: Formatted text data.

[0407] Specific operation: The server uses a sentiment analysis engine to analyze text data and recognize sentiment information. The recognized sentiment information is added to the format.

[0408] Output: Formatted text data with added sentiment information.

[0409] Step 5:

[0410] The user checks the formatted text content on the device screen.

[0411] Input: Formatted text data with added sentiment information.

[0412] Specific operation: The device displays text data to the user, who then reviews the content. The user makes corrections as needed and then approves it.

[0413] Output: Text data reviewed and modified by the user.

[0414] Step 6:

[0415] The server saves the formatted text approved by the user to a history file.

[0416] Input: Text data reviewed and modified by the user.

[0417] Specific action: The server appends the verified text data to the history file. For example, the data is saved to transcription_history.txt.

[0418] Output: Text data saved in the history file.

[0419] Through the above processing, a series of tasks such as transcribing recorded audio, formatting, emotion recognition, and history saving are performed efficiently.

[0420] (Application Example 2)

[0421] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0422] Conventional audio recording management systems often require manual transcription and sentiment analysis, which is time-consuming and labor-intensive. Furthermore, they lack mechanisms for accurately understanding user emotions, resulting in limited quality in reviews and feedback.

[0423] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, means for saving the confirmed content, and means for managing the saved content as history. This not only automates the management of recorded audio, but also records changes in the user's emotions in detail, enabling richer and more accurate reviews and feedback.

[0424] "Means for receiving recorded audio" refers to a mechanism for incorporating audio data provided by the user into the system.

[0425] "Methods for transcribing recorded audio" refer to the technologies and processes used to convert recorded audio data into text data.

[0426] "Methods for formatting transcribed content into a specific format" refers to methods of organizing and arranging transcribed text data according to a predetermined format.

[0427] "Methods for recognizing user emotions from transcribed content" refer to algorithms and technologies that analyze text and audio data to identify the user's emotional state.

[0428] "Means of verifying formatted content" refers to interfaces or methods that allow users to verify formatted text data and make corrections as needed.

[0429] "Means for saving confirmed content" refers to a mechanism for recording and saving the format content confirmed by the user in a database or file system.

[0430] "Means for managing saved content as history" refers to systems and processes for organizing saved data in chronological order so that it can be accessed and referenced later.

[0431] Modes for carrying out the invention

[0432] The invention is a system that automatically transcribes recorded audio, formats the content, saves it as a history, and further recognizes the user's emotions using an emotion engine, incorporating them into the format. This system is designed to allow users to easily manage their recorded audio and is particularly useful in content distribution services.

[0433] System Overview

[0434] The system consists of the following elements:

[0435] 1. Means for receiving recorded audio

[0436] The server has an interface for users to provide audio data, receives recorded audio files, and temporarily stores them. Users record audio reviews using a smartphone application and send them to the server.

[0437] 2. Methods for transcribing recorded audio into text.

[0438] The server transcribes the recorded audio file using the Google Speech-to-Text API. This API provides highly accurate speech recognition technology and converts the audio data into text data. The transcribed text data is temporarily stored on the server.

[0439] 3. Means for formatting the transcribed content into a specific format.

[0440] The server formats the transcribed text data into a specific format. Specifically, it formats it to include information such as the review date and time, content name, and username. The formatted text data is then temporarily stored in a database.

[0441] 4. Means of recognizing user emotions from transcribed content

[0442] The server uses Microsoft® Azure® Text Analytics API to recognize user emotions from transcribed text data. This API analyzes emotions (joy, sadness, surprise, etc.) within the text with high accuracy. The recognized emotion information is added to the format.

[0443] 5. Means of verifying formatted content

[0444] Users can review the formatted content through a smartphone application. They can then review the content, make corrections as needed, and finally approve it.

[0445] 6. Means for saving the confirmed information

[0446] The server saves the format content approved by the user to a history file. Specifically, it appends the confirmed data to the database and manages the saved content as history.

[0447] Specific example

[0448] Specific examples are given below.

[0449] A user records a movie review using the "Sentiment Analysis Content Guide" app and sends it to the server. Assume the recorded audio review contains the following:

[0450] "The movie I saw yesterday was very moving. The last scene was especially wonderful."

[0451] This audio review is processed by the server and a formatted text like the following is generated:

[0452] ---

[0453] Date and Time: 2023-10-15 17:45:00

[0454] Content name: Movie A

[0455] Username: user123

[0456] Contents:

[0457] The movie I saw yesterday was very moving. The last scene was especially wonderful.

[0458] Emotion: Joy

[0459] ---

[0460] The user reviews this formatted content in the application and saves it to the database after approval.

[0461] Example of a prompt

[0462] Please enter your thoughts on the movie you saw yesterday using voice input. Example: "The movie I saw yesterday was very moving! There were so many great scenes."

[0463] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0464] Step 1:

[0465] The user records an audio review within the application using their smartphone. The input is the user's voice data (e.g., "I was very moved by the movie I saw yesterday."), which is saved as digital data using the smartphone's built-in microphone. The output is the recorded audio file (e.g., "review_20231015.wav").

[0466] Step 2:

[0467] The server receives the recorded audio file. The audio file sent from the smartphone is uploaded to the server and temporarily stored. The input is the audio file sent by the user from their smartphone. The output is an audio file that is stored on the server.

[0468] Step 3:

[0469] The server uses the Google Speech-to-Text API to transcribe the received audio file. The input is an audio file stored on the server, which is sent to the API. This API converts the audio data into text data. The output is the transcribed text data (e.g., "The movie I saw yesterday was very moving.").

[0470] Step 4:

[0471] The server formats the transcribed text data into a specific format. Metadata such as review date and time, content name, and username are also added at this time. The input is the transcribed text data and associated metadata. The output is the formatted text data (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A Username: user123 Content: I was very moved by the movie I saw yesterday. ---").

[0472] Step 5:

[0473] The server uses Microsoft Azure's Text Analytics API to recognize the user's emotions from formatted text data. The input is formatted text data, which is sent to the API. This API analyzes the emotions within the text (e.g., "joy"). The output is formatted data with emotion information added (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A User Name: user123 Content: I was very moved by the movie I saw yesterday. Emotion: Joy ---").

[0474] Step 6:

[0475] Users review formatted text content through a smartphone application. The input is formatted data with added sentiment information, displayed on the application. Users review the content, make corrections as needed, and finally approve it. The output is the text data reviewed and corrected by the user.

[0476] Step 7:

[0477] The server saves the verified text data to the database. The input consists of user-approved formatted content, which is then recorded in the database. The output generates saved data that is managed as history.

[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0479] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0480] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0481] [Second Embodiment]

[0482] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0483] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0484] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0486] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0488] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0489] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0490] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0491] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0492] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0493] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0494] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. This system streamlines the management of recorded audio and enables accurate recording of its history.

[0495] System Overview

[0496] The system consists of the following elements:

[0497] 1. Means for receiving recorded audio

[0498] 2. Methods for transcribing recorded audio into text.

[0499] 3. Means for formatting the transcribed content into a specific format.

[0500] 4. Means of verifying formatted content

[0501] 5. Means for saving the confirmed information

[0502] Means for receiving recorded audio

[0503] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0504] Methods for transcribing recorded audio into text

[0505] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the recorded audio. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is stored in a temporary variable.

[0506] A method for formatting transcribed content into a specific format.

[0507] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0508] Means of verifying formatted content

[0509] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0510] Means for saving confirmed information

[0511] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt).

[0512] Specific example

[0513] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[0514] ---

[0515] Date and Time: 2023-10-15 14:35:00

[0516] Type: Customer support

[0517] Contents:

[0518] Hello, this is Customer Service. How can I help you today?

[0519] ---

[0520] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[0521] ---

[0522] Date and Time: 2023-10-15 14:35:00

[0523] Type: Customer support

[0524] Contents:

[0525] Hello, this is Customer Service. How can I help you today?

[0526] ---

[0527] This automates the management of recorded audio, saving time and effort, while also ensuring accuracy and consistency.

[0528] The following describes the processing flow.

[0529] Processing flow

[0530] Step 1:

[0531] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[0532] Step 2:

[0533] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0534] Step 3:

[0535] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0536] Step 4:

[0537] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0538] Step 5:

[0539] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0540] (Example 1)

[0541] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0542] Managing recorded audio requires manual transcription and formatting, which is time-consuming, labor-intensive, and can lead to inaccuracies and inconsistencies. Furthermore, verifying and correcting transcribed content is cumbersome, highlighting the need for an efficient management system. Additionally, a centralized method for managing and preserving the recording history is essential.

[0543] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0544] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio using a speech recognition API, means for formatting the transcribed content into a specific format, means for confirming the formatted content through a user interface, and means for saving the confirmed content to a history file. This streamlines the automatic transcription and formatting of recorded audio, enabling accurate and consistent management.

[0545] "Recorded audio" refers to a file containing audio data recorded in a specific manner.

[0546] "Means of receiving" refers to the function that allows the server to receive recorded audio files from users.

[0547] A "speech recognition API" is a program that provides an interface for converting speech data into text data.

[0548] "Transcribing methods" refer to the function of converting recorded audio into text data using a speech recognition API.

[0549] "Methods for formatting" refer to functions for organizing transcribed text data into a specific format.

[0550] A "user interface" is software that provides screens and input methods for users to interact with a system.

[0551] A "history file" is a text file used to save the confirmed information.

[0552] Modes for carrying out the invention

[0553] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. By using this system, the management of recorded audio can be streamlined, and accurate history recording can be achieved.

[0554] System Overview

[0555] The system consists of the following elements:

[0556] 1. Means for receiving recorded audio

[0557] 2. Methods for transcription

[0558] 3. Means for formatting the transcribed content into a specific format.

[0559] 4. Means of verifying formatted content

[0560] 5. Means for saving the confirmed information to a history file.

[0561] Hardware and software configuration

[0562] server

[0563] The server receives HTTP requests from users, receives and saves recorded audio files. The Flask framework in Python is used for this receiving process.

[0564] Specific receiving operation

[0565] When a user uploads the audio file customer_call.wav using a browser, the server receives it and saves it to a specific directory. This directory is / uploads / .

[0566] server

[0567] The server uses Google's speech recognition API to convert recorded audio into text data. Here, the Python pydub library is used to read the audio file and send a request to the API.

[0568] Specific transcription actions

[0569] The server sends the audio file `customer_call.wav` to the Google Speech Recognition API and receives the text data. The received text data is stored in a temporary variable `transcription_text`. For example, text such as "Hello, this is Customer Service. How can I help you today?" is generated.

[0570] terminal

[0571] The terminal receives text data sent from the server and formats it into a specific format. This format includes the recording date and time, the type of recording, and the transcribed content.

[0572] Specific formatting behavior

[0573] The terminal generates the following format based on the character data received from the server.

[0574] ---

[0575] Date and Time: 2023-10-15 14:35:00

[0576] Type: Customer support

[0577] Contents:

[0578] Hello, this is Customer Service. How can I help you today?

[0579] ---

[0580] The formatted data is stored in the temporary variable `formatted_text`.

[0581] User

[0582] The user reviews the formatted content on their device screen. They verify the content is accurate, make corrections if necessary, and finally approve it. This review process utilizes a web browser and a dedicated web application.

[0583] Specific verification actions

[0584] The user views the formatted content in their browser and clicks the confirmation button to check the content. If corrections are needed, they use the input fields to make the necessary changes and then click the "Approve" button.

[0585] server

[0586] The server receives the approved formatted content and saves it to the history file. The saved history is appended to a text file (e.g., transcription_history.txt).

[0587] Specific saving actions

[0588] The server receives the approved format content and appends it to the history file as follows:

[0589] ---

[0590] Date and Time: 2023-10-15 14:35:00

[0591] Type: Customer support

[0592] Contents:

[0593] Hello, this is Customer Service. How can I help you today?

[0594] ---

[0595] Example of a prompt

[0596] Examples of prompt messages in this system are as follows:

[0597] Please transcribe the audio file customer_call.wav using a speech recognition API and format it into the following format:

[0598] ---

[0599] Date and Time: YYYY-MM-DD HH:MM:SS

[0600] Type: Customer support

[0601] Contents:

[0602] <Transcripted text>

[0603] ---

[0604] In this way, the automatic transcription and formatting of recorded audio is streamlined, and it becomes possible to manage it with accuracy and consistency.

[0605] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0606] Step 1: Receiving the recorded audio

[0607] The server receives the recorded audio file sent as an HTTP request from the user.

[0608] Input: A recorded audio file uploaded by the user (e.g., customer_call.wav).

[0609] Data processing: The server temporarily stores the received file in memory and then moves the file to the specified directory.

[0610] Output: Audio files saved in the specified directory on the server (e.g., / uploads / ).

[0611] Specific operation: The user uploads an audio file through a browser form, the server receives it, and saves it to the / uploads / directory.

[0612] Step 2: Transcribe the recorded audio.

[0613] The server sends the stored audio file to a speech recognition API, where it is converted into text data.

[0614] Input: Audio file stored on the server (e.g., / uploads / customer_call.wav).

[0615] Data processing: The server uses the Python pydub library to read the audio file and send a request to the speech recognition API.

[0616] Output: Text data returned from the speech recognition API (e.g., "Hello, this is customer service. How can I help you today?").

[0617] Specific operation: The server sends the audio file to the Google Speech Recognition API and saves the returned text data to a temporary variable called `transcription_text`.

[0618] Step 3: Format of the transcript

[0619] The device receives text data and formats it into a specific format.

[0620] Input: Text data sent from the server (e.g., "Hello, this is Customer Service. How can I help you today?").

[0621] Data processing: Apply a format to the text data that includes the recording date and time, recording type, and transcript content.

[0622] Output: Formatted text data (e.g., in the format shown below).

[0623] ---

[0624] Date and Time: 2023-10-15 14:35:00

[0625] Type: Customer support

[0626] Contents:

[0627] Hello, this is Customer Service. How can I help you today?

[0628] ---

[0629] Specific operation: The terminal receives text data from the server, applies a format to it, shapes it, and saves it to a temporary variable called `formatted_text`.

[0630] Step 4: Check the format content

[0631] The user confirms the formatted content on the device screen.

[0632] Input: Formatted text data displayed on the terminal.

[0633] Data processing: The user reviews the content and makes corrections if necessary. The corrected data is then formatted again.

[0634] Output: The final text data approved or modified by the user.

[0635] Specific action: The user reviews the formatted content using a browser, makes corrections as needed, and clicks the "Approve" button.

[0636] Step 5: Save the confirmation details

[0637] The server saves the format content approved by the user to a history file.

[0638] Input: User-approved, formatted text data.

[0639] Data processing: Approved text data is appended to the history file (e.g., transcription_history.txt).

[0640] Output: Formatted text saved in the history file.

[0641] Specific operation: The server receives the approved text data and appends it to the history file in the following format.

[0642] ---

[0643] Date and Time: 2023-10-15 14:35:00

[0644] Type: Customer support

[0645] Contents:

[0646] Hello, this is Customer Service. How can I help you today?

[0647] ---

[0648] (Application Example 1)

[0649] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0650] As autonomous vehicles become more widespread, there is a growing need to efficiently manage, record, and store the voice interactions between passengers and vehicles as a history. However, existing systems rely on manual management of recorded voices, which is time-consuming and labor-intensive, and makes it difficult to maintain accuracy and consistency. Furthermore, there is a lack of adequate systems for effectively utilizing voice interactions within vehicles after the fact.

[0651] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0652] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for verifying the formatted content, means for saving the verified content, and, if the recorded audio is voice dialogue in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors, and means for displaying the transcribed content on the vehicle's display device or a portable information terminal. This automates the recording, transcription, formatting, verification, and saving of voice dialogue in autonomous vehicles, thereby improving the quality of operation management and customer service.

[0653] An "autonomous vehicle" is a vehicle that can operate autonomously without driver intervention.

[0654] An "acoustic sensor" is a device used to detect and record sound, such as a microphone installed in a vehicle.

[0655] A "portable information terminal" refers to a portable electronic device, such as a smartphone or tablet.

[0656] "Transcription" refers to the process of converting recorded audio into text data.

[0657] A "speech recognition API" is a programmatic interface for converting speech data into text data, and includes cloud-based and local-based services.

[0658] "Formatting to a specific format" refers to organizing and arranging transcribed text data according to a predetermined format and layout.

[0659] "Vehicle display devices" refer to display devices such as monitors and touch panels installed inside a vehicle.

[0660] "User interface" refers to the means by which a device or system and a user exchange information, and includes displays and input devices.

[0661] "Means of verification" refers to processes or devices that allow users to check formatted text content and make corrections or approvals as needed.

[0662] "Means of preservation" refers to the processes and systems used to record and retain verified text data in databases or files.

[0663] This invention relates to a system for recording, transcribing, and formatting voice dialogue within an autonomous vehicle as a history. This system streamlines the management of recorded audio and ensures accurate history recording.

[0664] System Overview

[0665] The system consists of the following elements:

[0666] 1. Means for receiving recorded audio

[0667] 2. Methods for transcribing recorded audio into text.

[0668] 3. Means for formatting the transcribed content into a specific format.

[0669] 4. Means of verifying formatted content

[0670] 5. Means for saving the confirmed information

[0671] 6. If the recorded audio is for voice interaction in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors.

[0672] 7. Means for displaying the transcribed content on a vehicle display device or portable information terminal.

[0673] Means for receiving recorded audio

[0674] Acoustic sensors (such as in-vehicle microphones) installed in autonomous vehicles record passenger voices. The recorded voice data is transmitted to the onboard computer and temporarily stored. This prepares the system for processing voice interactions with passengers.

[0675] Methods for transcribing recorded audio into text

[0676] The server uses a speech recognition API (e.g., a speech recognition cloud service) to transcribe the recorded audio. Specifically, it reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is temporarily stored.

[0677] A method for formatting transcribed content into a specific format.

[0678] The device formats the transcribed text data into a specific format. Specifically, it includes the recording date and time, the type of recording (customer interaction or passenger interaction), and the transcribed text content. The formatted text is temporarily stored.

[0679] Means of verifying formatted content

[0680] The user reviews the formatted text content via the vehicle's display device or a mobile device. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0681] Means for saving confirmed information

[0682] The server saves the formatted text approved by the user to a history file. Specifically, it appends the text to the history file (for example, transcription_history.txt).

[0683] Specific example

[0684] For example, suppose a passenger's voice interaction inside an autonomous vehicle is recorded and saved as a file named customer_interaction.wav. Let's say this audio file was uploaded on October 15, 2023 at 14:35:00. The system processes this audio file and generates text formatted as follows:

[0685] ---

[0686] Date and Time: 2023-10-15 14:35:00

[0687] Contents:

[0688] Hello, this is an autonomous vehicle. Where are you going today?

[0689] ---

[0690] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[0691] ---

[0692] Date and Time: 2023-10-15 14:35:00

[0693] Contents:

[0694] Hello, this is an autonomous vehicle. Where are you going today?

[0695] ---

[0696] This automates the management of voice interaction within autonomous vehicles, saving time and effort while ensuring accuracy and consistency.

[0697] Example prompts for generative AI models

[0698] "Create a sample Python program that records voice conversations with customers, transcribes them using a voice recognition cloud service, and saves them in a specific format."

[0699] As described above, this system automates the recording, transcription, formatting, verification, and saving of voice conversations within autonomous vehicles, thereby improving the quality of operational management and customer service.

[0700] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0701] Step 1:

[0702] Receiving recorded audio

[0703] The server receives recorded audio data through acoustic sensors installed in the autonomous vehicle. Specifically, a microphone captures the passenger's voice and sends that audio data to the server. The input is audio data, and the output is a temporarily stored audio file.

[0704] Step 2:

[0705] Transcription of recorded audio

[0706] The server uses a speech recognition API to convert the received audio data into text data. Specifically, a Python script is used to read the audio file and send a request to the speech recognition API. The input is an audio file, and the output is the transcribed text data.

[0707] Step 3:

[0708] Formatting of the transcribed content

[0709] The terminal formats the transcribed text data into a specific format. Specifically, it adds the recording date and time, and the type of recording (e.g., customer service, passenger interaction) to the text data and compiles it into a predetermined format. The input is the text data and the current date and time, and the output is the formatted text.

[0710] Step 4:

[0711] Checking the formatted content

[0712] The user reviews the formatted content via the vehicle's display device or a mobile device. Specifically, the formatted text is displayed on the screen of a display or smartphone, allowing the user to review the content and make corrections as needed. The input is the formatted text, and the output is the reviewed and corrected text.

[0713] Step 5:

[0714] Saving of confirmed information

[0715] The server saves the formatted text confirmed by the user to a history file. Specifically, it appends text data to the history file (e.g., transcription_history.txt). The input is the confirmed text data, and the output is the updated history file.

[0716] Step 6:

[0717] Managing history files

[0718] The server periodically manages history files, creates backups as needed, and archives old data. Specifically, it uses scripts to create file backups and manage file sizes. The input is history files, and the output is organized history data.

[0719] The above outlines the specific processing steps of the program, detailing the input, data processing, and output in each step. This process efficiently records, transcribes, formats, verifies, and saves voice conversations within autonomous vehicles.

[0720] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0721] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[0722] System Overview

[0723] The system consists of the following elements:

[0724] 1. Means for receiving recorded audio

[0725] 2. Methods for transcribing recorded audio into text.

[0726] 3. Means for formatting the transcribed content into a specific format.

[0727] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[0728] 5. Means of verifying formatted content

[0729] 6. Means for saving the confirmed information

[0730] Means for receiving recorded audio

[0731] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0732] Methods for transcribing recorded audio into text

[0733] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0734] A method for formatting transcribed content into a specific format.

[0735] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0736] An emotion engine that recognizes user emotions from transcribed content.

[0737] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[0738] Means of verifying formatted content

[0739] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0740] Means for saving confirmed information

[0741] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0742] Specific example

[0743] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[0744] ---

[0745] Date and Time: 2023-10-15 14:35:00

[0746] Type: Customer support

[0747] Contents:

[0748] Hello, this is Customer Service. How can I help you today?

[0749] Emotion: Joy

[0750] ---

[0751] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[0752] ---

[0753] Date and Time: 2023-10-15 14:35:00

[0754] Type: Customer support

[0755] Contents:

[0756] Hello, this is Customer Service. How can I help you today?

[0757] Emotion: Joy

[0758] ---

[0759] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0760] The following describes the processing flow.

[0761] Processing flow

[0762] Step 1:

[0763] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[0764] Step 2:

[0765] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0766] Step 3:

[0767] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0768] Step 4:

[0769] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The emotion engine uses natural language processing techniques to extract emotions (e.g., joy, anger, sadness) from the text data. The recognized emotion information is added to the text data.

[0770] Step 5:

[0771] The user reviews the combined formatted text data and recognized sentiment information on their device screen. The user verifies the accuracy of the content, makes corrections as needed, and then approves it.

[0772] Step 6:

[0773] The server saves the user-approved formatted text and sentiment information to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it has been saved correctly.

[0774] A concrete example of the process is as follows:

[0775] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 at 14:35:00. This recording file is processed by the system, generating text and sentiment information in the following format:

[0776] ---

[0777] Date and Time: 2023-10-15 14:35:00

[0778] Type: Customer support

[0779] Contents:

[0780] Hello, this is Customer Service. How can I help you today?

[0781] Emotion: Joy

[0782] ---

[0783] After the user reviews and approves this formatted content and sentiment information, the server saves it to a history file. The saved content will look like this:

[0784] ---

[0785] Date and Time: 2023-10-15 14:35:00

[0786] Type: Customer support

[0787] Contents:

[0788] Hello, this is Customer Service. How can I help you today?

[0789] Emotion: Joy

[0790] ---

[0791] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0792] (Example 2)

[0793] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0794] Conventional audio recording management systems often involve manual transcription, formatting, and emotional analysis of recordings, which is time-consuming, labor-intensive, and can lack accuracy. Furthermore, the lack of a function to analyze and record changes in emotion makes detailed conversation history management difficult.

[0795] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing it, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, and means for saving the confirmed content. This enables automatic transcription of recorded audio, formatting, emotion recognition, and history saving.

[0796] "Recorded audio" refers to audio data uploaded by users.

[0797] "Means of receiving" refers to the function by which a server receives audio data from a user and temporarily stores it.

[0798] "Transcribing" refers to a function that analyzes recorded audio and converts it into text data.

[0799] A "speech recognition API" is an interface for external services that convert speech data into text data.

[0800] "Methods for formatting" refer to functions that arrange text data into a specific format.

[0801] "Means of recognizing emotions" refers to a function that analyzes transcribed text data and audio data to identify the user's emotions.

[0802] A "sentiment analysis engine" is a software component used to analyze emotions from text and audio data.

[0803] A "means of verification" refers to a function that allows users to view formatted content and make corrections or approvals.

[0804] The "means of saving" refer to the function of saving the verified text data to a history file.

[0805] A "user interface" refers to the screens and control devices that allow a user to interact with and operate a system.

[0806] Defining it in this way clarifies the role of each function in the system.

[0807] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[0808] System Overview

[0809] The system consists of the following elements:

[0810] 1. Means for receiving recorded audio

[0811] 2. Methods for transcribing recorded audio into text.

[0812] 3. Means for formatting the transcribed content into a specific format.

[0813] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[0814] 5. Means of verifying formatted content

[0815] 6. Means for saving the confirmed information

[0816] Means for receiving recorded audio

[0817] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0818] Methods for transcribing recorded audio into text

[0819] The server uses a speech recognition API (for example, a common speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0820] A method for formatting transcribed content into a specific format.

[0821] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (e.g., customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0822] An emotion engine that recognizes user emotions from transcribed content.

[0823] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[0824] Means of verifying formatted content

[0825] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0826] Means for saving confirmed information

[0827] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0828] Specific example

[0829] For example, suppose a recorded audio file of a customer interaction, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recorded audio file is processed by the system, and text formatted as follows is generated:

[0830] ---

[0831] Date and Time: 2023-10-15 14:35:00

[0832] Type: Customer support

[0833] Contents:

[0834] Hello, this is Customer Service. How can I help you today?

[0835] Emotion: Joy

[0836] ---

[0837] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[0838] ---

[0839] Date and Time: 2023-10-15 14:35:00

[0840] Type: Customer support

[0841] Contents:

[0842] Hello, this is Customer Service. How can I help you today?

[0843] Emotion: Joy

[0844] ---

[0845] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[0846] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0847] Step 1:

[0848] The server receives the recorded audio file and stores it temporarily.

[0849] Input: Audio file sent by the user.

[0850] Specific operation: The user uploads an audio file from their device, and the server receives and temporarily stores the file.

[0851] Output: Temporarily saved audio file.

[0852] Step 2:

[0853] The server transcribes the audio file into text.

[0854] Input: Temporarily saved audio file.

[0855] Specific operation: The server sends an audio file to a speech recognition API (a common speech recognition service) and retrieves text data.

[0856] Output: Transcripted text data.

[0857] Step 3:

[0858] The device formats the transcribed text data into a specific format.

[0859] Input: Transcribed text data.

[0860] Specific operation: The device acquires text data and formats it into a format that includes the recording date and time, recording type, and transcript content.

[0861] Output: Formatted text data.

[0862] Step 4:

[0863] The server uses an emotion engine to recognize the user's emotions from the transcribed text data.

[0864] Input: Formatted text data.

[0865] Specific operation: The server uses a sentiment analysis engine to analyze text data and recognize sentiment information. The recognized sentiment information is added to the format.

[0866] Output: Formatted text data with added sentiment information.

[0867] Step 5:

[0868] The user checks the formatted text content on the device screen.

[0869] Input: Formatted text data with added sentiment information.

[0870] Specific operation: The device displays text data to the user, who then reviews the content. The user makes corrections as needed and then approves it.

[0871] Output: Text data reviewed and modified by the user.

[0872] Step 6:

[0873] The server saves the formatted text approved by the user to a history file.

[0874] Input: Text data reviewed and modified by the user.

[0875] Specific action: The server appends the verified text data to the history file. For example, the data is saved to transcription_history.txt.

[0876] Output: Text data saved in the history file.

[0877] Through the above processing, a series of tasks such as transcribing recorded audio, formatting, emotion recognition, and history saving are performed efficiently.

[0878] (Application Example 2)

[0879] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0880] Conventional audio recording management systems often require manual transcription and sentiment analysis, which is time-consuming and labor-intensive. Furthermore, they lack mechanisms for accurately understanding user emotions, resulting in limited quality in reviews and feedback.

[0881] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, means for saving the confirmed content, and means for managing the saved content as history. This not only automates the management of recorded audio, but also records changes in the user's emotions in detail, enabling richer and more accurate reviews and feedback.

[0882] "Means for receiving recorded audio" refers to a mechanism for incorporating audio data provided by the user into the system.

[0883] "Methods for transcribing recorded audio" refer to the technologies and processes used to convert recorded audio data into text data.

[0884] "Methods for formatting transcribed content into a specific format" refers to methods of organizing and arranging transcribed text data according to a predetermined format.

[0885] "Methods for recognizing user emotions from transcribed content" refer to algorithms and technologies that analyze text and audio data to identify the user's emotional state.

[0886] "Means of verifying formatted content" refers to interfaces or methods that allow users to verify formatted text data and make corrections as needed.

[0887] "Means for saving confirmed content" refers to a mechanism for recording and saving the format content confirmed by the user in a database or file system.

[0888] "Means for managing saved content as history" refers to systems and processes for organizing saved data in chronological order so that it can be accessed and referenced later.

[0889] Modes for carrying out the invention

[0890] The invention is a system that automatically transcribes recorded audio, formats the content, saves it as a history, and further recognizes the user's emotions using an emotion engine, incorporating them into the format. This system is designed to allow users to easily manage their recorded audio and is particularly useful in content distribution services.

[0891] System Overview

[0892] The system consists of the following elements:

[0893] 1. Means for receiving recorded audio

[0894] The server has an interface for users to provide audio data, receives recorded audio files, and temporarily stores them. Users record audio reviews using a smartphone application and send them to the server.

[0895] 2. Methods for transcribing recorded audio into text.

[0896] The server transcribes the recorded audio file using the Google Speech-to-Text API. This API provides highly accurate speech recognition technology and converts the audio data into text data. The transcribed text data is temporarily stored on the server.

[0897] 3. Means for formatting the transcribed content into a specific format.

[0898] The server formats the transcribed text data into a specific format. Specifically, it formats it to include information such as the review date and time, content name, and username. The formatted text data is then temporarily stored in a database.

[0899] 4. Means of recognizing user emotions from transcribed content

[0900] The server uses Microsoft Azure's Text Analytics API to recognize user emotions from transcribed text data. This API analyzes emotions (joy, sadness, surprise, etc.) within the text with high accuracy. The recognized emotion information is added to the format.

[0901] 5. Means of verifying formatted content

[0902] Users can review the formatted content through a smartphone application. They can then review the content, make corrections as needed, and finally approve it.

[0903] 6. Means for saving the confirmed information

[0904] The server saves the format content approved by the user to a history file. Specifically, it appends the confirmed data to the database and manages the saved content as history.

[0905] Specific example

[0906] Specific examples are given below.

[0907] A user records a movie review using the "Sentiment Analysis Content Guide" app and sends it to the server. Assume the recorded audio review contains the following:

[0908] "The movie I saw yesterday was very moving. The last scene was especially wonderful."

[0909] This audio review is processed by the server and a formatted text like the following is generated:

[0910] ---

[0911] Date and Time: 2023-10-15 17:45:00

[0912] Content name: Movie A

[0913] Username: user123

[0914] Contents:

[0915] The movie I saw yesterday was very moving. The last scene was especially wonderful.

[0916] Emotion: Joy

[0917] ---

[0918] The user reviews this formatted content in the application and saves it to the database after approval.

[0919] Example of a prompt

[0920] Please enter your thoughts on the movie you saw yesterday using voice input. Example: "The movie I saw yesterday was very moving! There were so many great scenes."

[0921] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0922] Step 1:

[0923] The user records an audio review within the application using their smartphone. The input is the user's voice data (e.g., "I was very moved by the movie I saw yesterday."), which is saved as digital data using the smartphone's built-in microphone. The output is the recorded audio file (e.g., "review_20231015.wav").

[0924] Step 2:

[0925] The server receives the recorded audio file. The audio file sent from the smartphone is uploaded to the server and temporarily stored. The input is the audio file sent by the user from their smartphone. The output is an audio file that is stored on the server.

[0926] Step 3:

[0927] The server uses the Google Speech-to-Text API to transcribe the received audio file. The input is an audio file stored on the server, which is sent to the API. This API converts the audio data into text data. The output is the transcribed text data (e.g., "The movie I saw yesterday was very moving.").

[0928] Step 4:

[0929] The server formats the transcribed text data into a specific format. Metadata such as review date and time, content name, and username are also added at this time. The input is the transcribed text data and associated metadata. The output is the formatted text data (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A Username: user123 Content: I was very moved by the movie I saw yesterday. ---").

[0930] Step 5:

[0931] The server uses Microsoft Azure's Text Analytics API to recognize the user's emotions from formatted text data. The input is formatted text data, which is sent to the API. This API analyzes the emotions within the text (e.g., "joy"). The output is formatted data with emotion information added (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A User Name: user123 Content: I was very moved by the movie I saw yesterday. Emotion: Joy ---").

[0932] Step 6:

[0933] Users review formatted text content through a smartphone application. The input is formatted data with added sentiment information, displayed on the application. Users review the content, make corrections as needed, and finally approve it. The output is the text data reviewed and corrected by the user.

[0934] Step 7:

[0935] The server saves the verified text data to the database. The input consists of user-approved formatted content, which is then recorded in the database. The output generates saved data that is managed as history.

[0936] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0937] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0938] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0939] [Third Embodiment]

[0940] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0941] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0942] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0943] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0944] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0945] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0946] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0947] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0948] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0949] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0950] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0951] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0952] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. This system streamlines the management of recorded audio and enables accurate recording of its history.

[0953] System Overview

[0954] The system consists of the following elements:

[0955] 1. Means for receiving recorded audio

[0956] 2. Methods for transcribing recorded audio into text.

[0957] 3. Means for formatting the transcribed content into a specific format.

[0958] 4. Means of verifying formatted content

[0959] 5. Means for saving the confirmed information

[0960] Means for receiving recorded audio

[0961] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[0962] Methods for transcribing recorded audio into text

[0963] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the recorded audio. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is stored in a temporary variable.

[0964] A method for formatting transcribed content into a specific format.

[0965] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0966] Means of verifying formatted content

[0967] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0968] Means for saving confirmed information

[0969] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt).

[0970] Specific example

[0971] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[0972] ---

[0973] Date and Time: 2023-10-15 14:35:00

[0974] Type: Customer support

[0975] Contents:

[0976] Hello, this is Customer Service. How can I help you today?

[0977] ---

[0978] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[0979] ---

[0980] Date and Time: 2023-10-15 14:35:00

[0981] Type: Customer support

[0982] Contents:

[0983] Hello, this is Customer Service. How can I help you today?

[0984] ---

[0985] This automates the management of recorded audio, saving time and effort, while also ensuring accuracy and consistency.

[0986] The following describes the processing flow.

[0987] Processing flow

[0988] Step 1:

[0989] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[0990] Step 2:

[0991] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[0992] Step 3:

[0993] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[0994] Step 4:

[0995] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[0996] Step 5:

[0997] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[0998] (Example 1)

[0999] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1000] Managing recorded audio requires manual transcription and formatting, which is time-consuming, labor-intensive, and can lead to inaccuracies and inconsistencies. Furthermore, verifying and correcting transcribed content is cumbersome, highlighting the need for an efficient management system. Additionally, a centralized method for managing and preserving the recording history is essential.

[1001] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1002] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio using a speech recognition API, means for formatting the transcribed content into a specific format, means for confirming the formatted content through a user interface, and means for saving the confirmed content to a history file. This streamlines the automatic transcription and formatting of recorded audio, enabling accurate and consistent management.

[1003] "Recorded audio" refers to a file containing audio data recorded in a specific manner.

[1004] "Means of receiving" refers to the function that allows the server to receive recorded audio files from users.

[1005] A "speech recognition API" is a program that provides an interface for converting speech data into text data.

[1006] "Transcribing methods" refer to the function of converting recorded audio into text data using a speech recognition API.

[1007] "Methods for formatting" refer to functions for organizing transcribed text data into a specific format.

[1008] A "user interface" is software that provides screens and input methods for users to interact with a system.

[1009] A "history file" is a text file used to save the confirmed information.

[1010] Modes for carrying out the invention

[1011] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. By using this system, the management of recorded audio can be streamlined, and accurate history recording can be achieved.

[1012] System Overview

[1013] The system consists of the following elements:

[1014] 1. Means for receiving recorded audio

[1015] 2. Methods for transcription

[1016] 3. Means for formatting the transcribed content into a specific format.

[1017] 4. Means of verifying formatted content

[1018] 5. Means for saving the confirmed information to a history file.

[1019] Hardware and software configuration

[1020] server

[1021] The server receives HTTP requests from users, receives and saves recorded audio files. The Flask framework in Python is used for this receiving process.

[1022] Specific receiving operation

[1023] When a user uploads the audio file customer_call.wav using a browser, the server receives it and saves it to a specific directory. This directory is / uploads / .

[1024] server

[1025] The server uses Google's speech recognition API to convert recorded audio into text data. Here, the Python pydub library is used to read the audio file and send a request to the API.

[1026] Specific transcription actions

[1027] The server sends the audio file `customer_call.wav` to the Google Speech Recognition API and receives the text data. The received text data is stored in a temporary variable `transcription_text`. For example, text such as "Hello, this is Customer Service. How can I help you today?" is generated.

[1028] terminal

[1029] The terminal receives text data sent from the server and formats it into a specific format. This format includes the recording date and time, the type of recording, and the transcribed content.

[1030] Specific formatting behavior

[1031] The terminal generates the following format based on the character data received from the server.

[1032] ---

[1033] Date and Time: 2023-10-15 14:35:00

[1034] Type: Customer support

[1035] Contents:

[1036] Hello, this is Customer Service. How can I help you today?

[1037] ---

[1038] The formatted data is stored in the temporary variable `formatted_text`.

[1039] User

[1040] The user reviews the formatted content on their device screen. They verify the content is accurate, make corrections if necessary, and finally approve it. This review process utilizes a web browser and a dedicated web application.

[1041] Specific verification actions

[1042] The user views the formatted content in their browser and clicks the confirmation button to check the content. If corrections are needed, they use the input fields to make the necessary changes and then click the "Approve" button.

[1043] server

[1044] The server receives the approved formatted content and saves it to the history file. The saved history is appended to a text file (e.g., transcription_history.txt).

[1045] Specific saving actions

[1046] The server receives the approved format content and appends it to the history file as follows:

[1047] ---

[1048] Date and Time: 2023-10-15 14:35:00

[1049] Type: Customer support

[1050] Contents:

[1051] Hello, this is Customer Service. How can I help you today?

[1052] ---

[1053] Example of a prompt

[1054] Examples of prompt messages in this system are as follows:

[1055] Please transcribe the audio file customer_call.wav using a speech recognition API and format it into the following format:

[1056] ---

[1057] Date and Time: YYYY-MM-DD HH:MM:SS

[1058] Type: Customer support

[1059] Contents:

[1060] <Transcripted text>

[1061] ---

[1062] In this way, the automatic transcription and formatting of recorded audio is streamlined, and it becomes possible to manage it with accuracy and consistency.

[1063] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1064] Step 1: Receiving the recorded audio

[1065] The server receives the recorded audio file sent as an HTTP request from the user.

[1066] Input: A recorded audio file uploaded by the user (e.g., customer_call.wav).

[1067] Data processing: The server temporarily stores the received file in memory and then moves the file to the specified directory.

[1068] Output: Audio files saved in the specified directory on the server (e.g., / uploads / ).

[1069] Specific operation: The user uploads an audio file through a browser form, the server receives it, and saves it to the / uploads / directory.

[1070] Step 2: Transcribe the recorded audio.

[1071] The server sends the stored audio file to a speech recognition API, where it is converted into text data.

[1072] Input: Audio file stored on the server (e.g., / uploads / customer_call.wav).

[1073] Data processing: The server uses the Python pydub library to read the audio file and send a request to the speech recognition API.

[1074] Output: Text data returned from the speech recognition API (e.g., "Hello, this is customer service. How can I help you today?").

[1075] Specific operation: The server sends the audio file to the Google Speech Recognition API and saves the returned text data to a temporary variable called `transcription_text`.

[1076] Step 3: Format of the transcript

[1077] The device receives text data and formats it into a specific format.

[1078] Input: Text data sent from the server (e.g., "Hello, this is Customer Service. How can I help you today?").

[1079] Data processing: Apply a format to the text data that includes the recording date and time, recording type, and transcript content.

[1080] Output: Formatted text data (e.g., in the format shown below).

[1081] ---

[1082] Date and Time: 2023-10-15 14:35:00

[1083] Type: Customer support

[1084] Contents:

[1085] Hello, this is Customer Service. How can I help you today?

[1086] ---

[1087] Specific operation: The terminal receives text data from the server, applies a format to it, shapes it, and saves it to a temporary variable called `formatted_text`.

[1088] Step 4: Check the format content

[1089] The user confirms the formatted content on the device screen.

[1090] Input: Formatted text data displayed on the terminal.

[1091] Data processing: The user reviews the content and makes corrections if necessary. The corrected data is then formatted again.

[1092] Output: The final text data approved or modified by the user.

[1093] Specific action: The user reviews the formatted content using a browser, makes corrections as needed, and clicks the "Approve" button.

[1094] Step 5: Save the confirmation details

[1095] The server saves the format content approved by the user to a history file.

[1096] Input: User-approved, formatted text data.

[1097] Data processing: Approved text data is appended to the history file (e.g., transcription_history.txt).

[1098] Output: Formatted text saved in the history file.

[1099] Specific operation: The server receives the approved text data and appends it to the history file in the following format.

[1100] ---

[1101] Date and Time: 2023-10-15 14:35:00

[1102] Type: Customer support

[1103] Contents:

[1104] Hello, this is Customer Service. How can I help you today?

[1105] ---

[1106] (Application Example 1)

[1107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1108] As autonomous vehicles become more widespread, there is a growing need to efficiently manage, record, and store the voice interactions between passengers and vehicles as a history. However, existing systems rely on manual management of recorded voices, which is time-consuming and labor-intensive, and makes it difficult to maintain accuracy and consistency. Furthermore, there is a lack of adequate systems for effectively utilizing voice interactions within vehicles after the fact.

[1109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1110] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for verifying the formatted content, means for saving the verified content, and, if the recorded audio is voice dialogue in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors, and means for displaying the transcribed content on the vehicle's display device or a portable information terminal. This automates the recording, transcription, formatting, verification, and saving of voice dialogue in autonomous vehicles, thereby improving the quality of operation management and customer service.

[1111] An "autonomous vehicle" is a vehicle that can operate autonomously without driver intervention.

[1112] An "acoustic sensor" is a device used to detect and record sound, such as a microphone installed in a vehicle.

[1113] A "portable information terminal" refers to a portable electronic device, such as a smartphone or tablet.

[1114] "Transcription" refers to the process of converting recorded audio into text data.

[1115] A "speech recognition API" is a programmatic interface for converting speech data into text data, and includes cloud-based and local-based services.

[1116] "Formatting to a specific format" refers to organizing and arranging transcribed text data according to a predetermined format and layout.

[1117] "Vehicle display devices" refer to display devices such as monitors and touch panels installed inside a vehicle.

[1118] "User interface" refers to the means by which a device or system and a user exchange information, and includes displays and input devices.

[1119] "Means of verification" refers to processes or devices that allow users to check formatted text content and make corrections or approvals as needed.

[1120] "Means of preservation" refers to the processes and systems used to record and retain verified text data in databases or files.

[1121] This invention relates to a system for recording, transcribing, and formatting voice dialogue within an autonomous vehicle as a history. This system streamlines the management of recorded audio and ensures accurate history recording.

[1122] System Overview

[1123] The system consists of the following elements:

[1124] 1. Means for receiving recorded audio

[1125] 2. Methods for transcribing recorded audio into text.

[1126] 3. Means for formatting the transcribed content into a specific format.

[1127] 4. Means of verifying formatted content

[1128] 5. Means for saving the confirmed information

[1129] 6. If the recorded audio is for voice interaction in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors.

[1130] 7. Means for displaying the transcribed content on a vehicle display device or portable information terminal.

[1131] Means for receiving recorded audio

[1132] Acoustic sensors (such as in-vehicle microphones) installed in autonomous vehicles record passenger voices. The recorded voice data is transmitted to the onboard computer and temporarily stored. This prepares the system for processing voice interactions with passengers.

[1133] Methods for transcribing recorded audio into text

[1134] The server uses a speech recognition API (e.g., a speech recognition cloud service) to transcribe the recorded audio. Specifically, it reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is temporarily stored.

[1135] A method for formatting transcribed content into a specific format.

[1136] The device formats the transcribed text data into a specific format. Specifically, it includes the recording date and time, the type of recording (customer interaction or passenger interaction), and the transcribed text content. The formatted text is temporarily stored.

[1137] Means of verifying formatted content

[1138] The user reviews the formatted text content via the vehicle's display device or a mobile device. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1139] Means for saving confirmed information

[1140] The server saves the formatted text approved by the user to a history file. Specifically, it appends the text to the history file (for example, transcription_history.txt).

[1141] Specific example

[1142] For example, suppose a passenger's voice interaction inside an autonomous vehicle is recorded and saved as a file named customer_interaction.wav. Let's say this audio file was uploaded on October 15, 2023 at 14:35:00. The system processes this audio file and generates text formatted as follows:

[1143] ---

[1144] Date and Time: 2023-10-15 14:35:00

[1145] Contents:

[1146] Hello, this is an autonomous vehicle. Where are you going today?

[1147] ---

[1148] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[1149] ---

[1150] Date and Time: 2023-10-15 14:35:00

[1151] Contents:

[1152] Hello, this is an autonomous vehicle. Where are you going today?

[1153] ---

[1154] This automates the management of voice interaction within autonomous vehicles, saving time and effort while ensuring accuracy and consistency.

[1155] Example prompts for generative AI models

[1156] "Create a sample Python program that records voice conversations with customers, transcribes them using a voice recognition cloud service, and saves them in a specific format."

[1157] As described above, this system automates the recording, transcription, formatting, verification, and saving of voice conversations within autonomous vehicles, thereby improving the quality of operational management and customer service.

[1158] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1159] Step 1:

[1160] Receiving recorded audio

[1161] The server receives recorded audio data through acoustic sensors installed in the autonomous vehicle. Specifically, a microphone captures the passenger's voice and sends that audio data to the server. The input is audio data, and the output is a temporarily stored audio file.

[1162] Step 2:

[1163] Transcription of recorded audio

[1164] The server uses a speech recognition API to convert the received audio data into text data. Specifically, a Python script is used to read the audio file and send a request to the speech recognition API. The input is an audio file, and the output is the transcribed text data.

[1165] Step 3:

[1166] Formatting of the transcribed content

[1167] The terminal formats the transcribed text data into a specific format. Specifically, it adds the recording date and time, and the type of recording (e.g., customer service, passenger interaction) to the text data and compiles it into a predetermined format. The input is the text data and the current date and time, and the output is the formatted text.

[1168] Step 4:

[1169] Checking the formatted content

[1170] The user reviews the formatted content via the vehicle's display device or a mobile device. Specifically, the formatted text is displayed on the screen of a display or smartphone, allowing the user to review the content and make corrections as needed. The input is the formatted text, and the output is the reviewed and corrected text.

[1171] Step 5:

[1172] Saving of confirmed information

[1173] The server saves the formatted text confirmed by the user to a history file. Specifically, it appends text data to the history file (e.g., transcription_history.txt). The input is the confirmed text data, and the output is the updated history file.

[1174] Step 6:

[1175] Managing history files

[1176] The server periodically manages history files, creates backups as needed, and archives old data. Specifically, it uses scripts to create file backups and manage file sizes. The input is history files, and the output is organized history data.

[1177] The above outlines the specific processing steps of the program, detailing the input, data processing, and output in each step. This process efficiently records, transcribes, formats, verifies, and saves voice conversations within autonomous vehicles.

[1178] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1179] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[1180] System Overview

[1181] The system consists of the following elements:

[1182] 1. Means for receiving recorded audio

[1183] 2. Methods for transcribing recorded audio into text.

[1184] 3. Means for formatting the transcribed content into a specific format.

[1185] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[1186] 5. Means of verifying formatted content

[1187] 6. Means for saving the confirmed information

[1188] Means for receiving recorded audio

[1189] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[1190] Methods for transcribing recorded audio into text

[1191] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1192] A method for formatting transcribed content into a specific format.

[1193] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1194] An emotion engine that recognizes user emotions from transcribed content.

[1195] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[1196] Means of verifying formatted content

[1197] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1198] Means for saving confirmed information

[1199] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[1200] Specific example

[1201] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[1202] ---

[1203] Date and Time: 2023-10-15 14:35:00

[1204] Type: Customer support

[1205] Contents:

[1206] Hello, this is Customer Service. How can I help you today?

[1207] Emotion: Joy

[1208] ---

[1209] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[1210] ---

[1211] Date and Time: 2023-10-15 14:35:00

[1212] Type: Customer support

[1213] Contents:

[1214] Hello, this is Customer Service. How can I help you today?

[1215] Emotion: Joy

[1216] ---

[1217] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1218] The following describes the processing flow.

[1219] Processing flow

[1220] Step 1:

[1221] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[1222] Step 2:

[1223] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1224] Step 3:

[1225] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1226] Step 4:

[1227] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The emotion engine uses natural language processing techniques to extract emotions (e.g., joy, anger, sadness) from the text data. The recognized emotion information is added to the text data.

[1228] Step 5:

[1229] The user reviews the combined formatted text data and recognized sentiment information on their device screen. The user verifies the accuracy of the content, makes corrections as needed, and then approves it.

[1230] Step 6:

[1231] The server saves the user-approved formatted text and sentiment information to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it has been saved correctly.

[1232] A concrete example of the process is as follows:

[1233] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 at 14:35:00. This recording file is processed by the system, generating text and sentiment information in the following format:

[1234] ---

[1235] Date and Time: 2023-10-15 14:35:00

[1236] Type: Customer support

[1237] Contents:

[1238] Hello, this is Customer Service. How can I help you today?

[1239] Emotion: Joy

[1240] ---

[1241] After the user reviews and approves this formatted content and sentiment information, the server saves it to a history file. The saved content will look like this:

[1242] ---

[1243] Date and Time: 2023-10-15 14:35:00

[1244] Type: Customer support

[1245] Contents:

[1246] Hello, this is Customer Service. How can I help you today?

[1247] Emotion: Joy

[1248] ---

[1249] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1250] (Example 2)

[1251] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1252] Conventional audio recording management systems often involve manual transcription, formatting, and emotional analysis of recordings, which is time-consuming, labor-intensive, and can lack accuracy. Furthermore, the lack of a function to analyze and record changes in emotion makes detailed conversation history management difficult.

[1253] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing it, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, and means for saving the confirmed content. This enables automatic transcription of recorded audio, formatting, emotion recognition, and history saving.

[1254] "Recorded audio" refers to audio data uploaded by users.

[1255] "Means of receiving" refers to the function by which a server receives audio data from a user and temporarily stores it.

[1256] "Transcribing" refers to a function that analyzes recorded audio and converts it into text data.

[1257] A "speech recognition API" is an interface for external services that convert speech data into text data.

[1258] "Methods for formatting" refer to functions that arrange text data into a specific format.

[1259] "Means of recognizing emotions" refers to a function that analyzes transcribed text data and audio data to identify the user's emotions.

[1260] A "sentiment analysis engine" is a software component used to analyze emotions from text and audio data.

[1261] A "means of verification" refers to a function that allows users to view formatted content and make corrections or approvals.

[1262] The "means of saving" refer to the function of saving the verified text data to a history file.

[1263] A "user interface" refers to the screens and control devices that allow a user to interact with and operate a system.

[1264] Defining it in this way clarifies the role of each function in the system.

[1265] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[1266] System Overview

[1267] The system consists of the following elements:

[1268] 1. Means for receiving recorded audio

[1269] 2. Methods for transcribing recorded audio into text.

[1270] 3. Means for formatting the transcribed content into a specific format.

[1271] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[1272] 5. Means of verifying formatted content

[1273] 6. Means for saving the confirmed information

[1274] Means for receiving recorded audio

[1275] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[1276] Methods for transcribing recorded audio into text

[1277] The server uses a speech recognition API (for example, a common speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1278] A method for formatting transcribed content into a specific format.

[1279] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (e.g., customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1280] An emotion engine that recognizes user emotions from transcribed content.

[1281] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[1282] Means of verifying formatted content

[1283] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1284] Means for saving confirmed information

[1285] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[1286] Specific example

[1287] For example, suppose a recorded audio file of a customer interaction, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recorded audio file is processed by the system, and text formatted as follows is generated:

[1288] ---

[1289] Date and Time: 2023-10-15 14:35:00

[1290] Type: Customer support

[1291] Contents:

[1292] Hello, this is Customer Service. How can I help you today?

[1293] Emotion: Joy

[1294] ---

[1295] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[1296] ---

[1297] Date and Time: 2023-10-15 14:35:00

[1298] Type: Customer support

[1299] Contents:

[1300] Hello, this is Customer Service. How can I help you today?

[1301] Emotion: Joy

[1302] ---

[1303] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1304] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1305] Step 1:

[1306] The server receives the recorded audio file and stores it temporarily.

[1307] Input: Audio file sent by the user.

[1308] Specific operation: The user uploads an audio file from their device, and the server receives and temporarily stores the file.

[1309] Output: Temporarily saved audio file.

[1310] Step 2:

[1311] The server transcribes the audio file into text.

[1312] Input: Temporarily saved audio file.

[1313] Specific operation: The server sends an audio file to a speech recognition API (a common speech recognition service) and retrieves text data.

[1314] Output: Transcripted text data.

[1315] Step 3:

[1316] The device formats the transcribed text data into a specific format.

[1317] Input: Transcribed text data.

[1318] Specific operation: The device acquires text data and formats it into a format that includes the recording date and time, recording type, and transcript content.

[1319] Output: Formatted text data.

[1320] Step 4:

[1321] The server uses an emotion engine to recognize the user's emotions from the transcribed text data.

[1322] Input: Formatted text data.

[1323] Specific operation: The server uses a sentiment analysis engine to analyze text data and recognize sentiment information. The recognized sentiment information is added to the format.

[1324] Output: Formatted text data with added sentiment information.

[1325] Step 5:

[1326] The user checks the formatted text content on the device screen.

[1327] Input: Formatted text data with added sentiment information.

[1328] Specific operation: The device displays text data to the user, who then reviews the content. The user makes corrections as needed and then approves it.

[1329] Output: Text data reviewed and modified by the user.

[1330] Step 6:

[1331] The server saves the formatted text approved by the user to a history file.

[1332] Input: Text data reviewed and modified by the user.

[1333] Specific action: The server appends the verified text data to the history file. For example, the data is saved to transcription_history.txt.

[1334] Output: Text data saved in the history file.

[1335] Through the above processing, a series of tasks such as transcribing recorded audio, formatting, emotion recognition, and history saving are performed efficiently.

[1336] (Application Example 2)

[1337] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1338] Conventional audio recording management systems often require manual transcription and sentiment analysis, which is time-consuming and labor-intensive. Furthermore, they lack mechanisms for accurately understanding user emotions, resulting in limited quality in reviews and feedback.

[1339] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, means for saving the confirmed content, and means for managing the saved content as history. This not only automates the management of recorded audio, but also records changes in the user's emotions in detail, enabling richer and more accurate reviews and feedback.

[1340] "Means for receiving recorded audio" refers to a mechanism for incorporating audio data provided by the user into the system.

[1341] "Methods for transcribing recorded audio" refer to the technologies and processes used to convert recorded audio data into text data.

[1342] "Methods for formatting transcribed content into a specific format" refers to methods of organizing and arranging transcribed text data according to a predetermined format.

[1343] "Methods for recognizing user emotions from transcribed content" refer to algorithms and technologies that analyze text and audio data to identify the user's emotional state.

[1344] "Means of verifying formatted content" refers to interfaces or methods that allow users to verify formatted text data and make corrections as needed.

[1345] "Means for saving confirmed content" refers to a mechanism for recording and saving the format content confirmed by the user in a database or file system.

[1346] "Means for managing saved content as history" refers to systems and processes for organizing saved data in chronological order so that it can be accessed and referenced later.

[1347] Modes for carrying out the invention

[1348] The invention is a system that automatically transcribes recorded audio, formats the content, saves it as a history, and further recognizes the user's emotions using an emotion engine, incorporating them into the format. This system is designed to allow users to easily manage their recorded audio and is particularly useful in content distribution services.

[1349] System Overview

[1350] The system consists of the following elements:

[1351] 1. Means for receiving recorded audio

[1352] The server has an interface for users to provide audio data, receives recorded audio files, and temporarily stores them. Users record audio reviews using a smartphone application and send them to the server.

[1353] 2. Methods for transcribing recorded audio into text.

[1354] The server transcribes the recorded audio file using the Google Speech-to-Text API. This API provides highly accurate speech recognition technology and converts the audio data into text data. The transcribed text data is temporarily stored on the server.

[1355] 3. Means for formatting the transcribed content into a specific format.

[1356] The server formats the transcribed text data into a specific format. Specifically, it formats it to include information such as the review date and time, content name, and username. The formatted text data is then temporarily stored in a database.

[1357] 4. Means of recognizing user emotions from transcribed content

[1358] The server uses Microsoft Azure's Text Analytics API to recognize user emotions from transcribed text data. This API analyzes emotions (joy, sadness, surprise, etc.) within the text with high accuracy. The recognized emotion information is added to the format.

[1359] 5. Means of verifying formatted content

[1360] Users can review the formatted content through a smartphone application. They can then review the content, make corrections as needed, and finally approve it.

[1361] 6. Means for saving the confirmed information

[1362] The server saves the format content approved by the user to a history file. Specifically, it appends the confirmed data to the database and manages the saved content as history.

[1363] Specific example

[1364] Specific examples are given below.

[1365] A user records a movie review using the "Sentiment Analysis Content Guide" app and sends it to the server. Assume the recorded audio review contains the following:

[1366] "The movie I saw yesterday was very moving. The last scene was especially wonderful."

[1367] This audio review is processed by the server and a formatted text like the following is generated:

[1368] ---

[1369] Date and Time: 2023-10-15 17:45:00

[1370] Content name: Movie A

[1371] Username: user123

[1372] Contents:

[1373] The movie I saw yesterday was very moving. The last scene was especially wonderful.

[1374] Emotion: Joy

[1375] ---

[1376] The user reviews this formatted content in the application and saves it to the database after approval.

[1377] Example of a prompt

[1378] Please enter your thoughts on the movie you saw yesterday using voice input. Example: "The movie I saw yesterday was very moving! There were so many great scenes."

[1379] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1380] Step 1:

[1381] The user records an audio review within the application using their smartphone. The input is the user's voice data (e.g., "I was very moved by the movie I saw yesterday."), which is saved as digital data using the smartphone's built-in microphone. The output is the recorded audio file (e.g., "review_20231015.wav").

[1382] Step 2:

[1383] The server receives the recorded audio file. The audio file sent from the smartphone is uploaded to the server and temporarily stored. The input is the audio file sent by the user from their smartphone. The output is an audio file that is stored on the server.

[1384] Step 3:

[1385] The server uses the Google Speech-to-Text API to transcribe the received audio file. The input is an audio file stored on the server, which is sent to the API. This API converts the audio data into text data. The output is the transcribed text data (e.g., "The movie I saw yesterday was very moving.").

[1386] Step 4:

[1387] The server formats the transcribed text data into a specific format. Metadata such as review date and time, content name, and username are also added at this time. The input is the transcribed text data and associated metadata. The output is the formatted text data (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A Username: user123 Content: I was very moved by the movie I saw yesterday. ---").

[1388] Step 5:

[1389] The server uses Microsoft Azure's Text Analytics API to recognize the user's emotions from formatted text data. The input is formatted text data, which is sent to the API. This API analyzes the emotions within the text (e.g., "joy"). The output is formatted data with emotion information added (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A User Name: user123 Content: I was very moved by the movie I saw yesterday. Emotion: Joy ---").

[1390] Step 6:

[1391] Users review formatted text content through a smartphone application. The input is formatted data with added sentiment information, displayed on the application. Users review the content, make corrections as needed, and finally approve it. The output is the text data reviewed and corrected by the user.

[1392] Step 7:

[1393] The server saves the verified text data to the database. The input consists of user-approved formatted content, which is then recorded in the database. The output generates saved data that is managed as history.

[1394] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1395] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1396] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1397] [Fourth Embodiment]

[1398] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1399] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1400] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1401] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1402] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1404] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1405] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1406] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1407] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1408] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1409] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1410] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1411] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. This system streamlines the management of recorded audio and enables accurate recording of its history.

[1412] System Overview

[1413] The system consists of the following elements:

[1414] 1. Means for receiving recorded audio

[1415] 2. Methods for transcribing recorded audio into text.

[1416] 3. Means for formatting the transcribed content into a specific format.

[1417] 4. Means of verifying formatted content

[1418] 5. Means for saving the confirmed information

[1419] Means for receiving recorded audio

[1420] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[1421] Methods for transcribing recorded audio into text

[1422] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the recorded audio. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is stored in a temporary variable.

[1423] A method for formatting transcribed content into a specific format.

[1424] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1425] Means of verifying formatted content

[1426] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1427] Means for saving confirmed information

[1428] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt).

[1429] Specific example

[1430] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[1431] ---

[1432] Date and Time: 2023-10-15 14:35:00

[1433] Type: Customer support

[1434] Contents:

[1435] Hello, this is Customer Service. How can I help you today?

[1436] ---

[1437] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[1438] ---

[1439] Date and Time: 2023-10-15 14:35:00

[1440] Type: Customer support

[1441] Contents:

[1442] Hello, this is Customer Service. How can I help you today?

[1443] ---

[1444] This automates the management of recorded audio, saving time and effort, while also ensuring accuracy and consistency.

[1445] The following describes the processing flow.

[1446] Processing flow

[1447] Step 1:

[1448] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[1449] Step 2:

[1450] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1451] Step 3:

[1452] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1453] Step 4:

[1454] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1455] Step 5:

[1456] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[1457] (Example 1)

[1458] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1459] Managing recorded audio requires manual transcription and formatting, which is time-consuming, labor-intensive, and can lead to inaccuracies and inconsistencies. Furthermore, verifying and correcting transcribed content is cumbersome, highlighting the need for an efficient management system. Additionally, a centralized method for managing and preserving the recording history is essential.

[1460] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1461] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio using a speech recognition API, means for formatting the transcribed content into a specific format, means for confirming the formatted content through a user interface, and means for saving the confirmed content to a history file. This streamlines the automatic transcription and formatting of recorded audio, enabling accurate and consistent management.

[1462] "Recorded audio" refers to a file containing audio data recorded in a specific manner.

[1463] "Means of receiving" refers to the function that allows the server to receive recorded audio files from users.

[1464] A "speech recognition API" is a program that provides an interface for converting speech data into text data.

[1465] "Transcribing methods" refer to the function of converting recorded audio into text data using a speech recognition API.

[1466] "Methods for formatting" refer to functions for organizing transcribed text data into a specific format.

[1467] A "user interface" is software that provides screens and input methods for users to interact with a system.

[1468] A "history file" is a text file used to save the confirmed information.

[1469] Modes for carrying out the invention

[1470] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history. By using this system, the management of recorded audio can be streamlined, and accurate history recording can be achieved.

[1471] System Overview

[1472] The system consists of the following elements:

[1473] 1. Means for receiving recorded audio

[1474] 2. Methods for transcription

[1475] 3. Means for formatting the transcribed content into a specific format.

[1476] 4. Means of verifying formatted content

[1477] 5. Means for saving the confirmed information to a history file.

[1478] Hardware and software configuration

[1479] server

[1480] The server receives HTTP requests from users, receives and saves recorded audio files. The Flask framework in Python is used for this receiving process.

[1481] Specific receiving operation

[1482] When a user uploads the audio file customer_call.wav using a browser, the server receives it and saves it to a specific directory. This directory is / uploads / .

[1483] server

[1484] The server uses Google's speech recognition API to convert recorded audio into text data. Here, the Python pydub library is used to read the audio file and send a request to the API.

[1485] Specific transcription actions

[1486] The server sends the audio file `customer_call.wav` to the Google Speech Recognition API and receives the text data. The received text data is stored in a temporary variable `transcription_text`. For example, text such as "Hello, this is Customer Service. How can I help you today?" is generated.

[1487] terminal

[1488] The terminal receives text data sent from the server and formats it into a specific format. This format includes the recording date and time, the type of recording, and the transcribed content.

[1489] Specific formatting behavior

[1490] The terminal generates the following format based on the character data received from the server.

[1491] ---

[1492] Date and Time: 2023-10-15 14:35:00

[1493] Type: Customer support

[1494] Contents:

[1495] Hello, this is Customer Service. How can I help you today?

[1496] ---

[1497] The formatted data is stored in the temporary variable `formatted_text`.

[1498] User

[1499] The user reviews the formatted content on their device screen. They verify the content is accurate, make corrections if necessary, and finally approve it. This review process utilizes a web browser and a dedicated web application.

[1500] Specific verification actions

[1501] The user views the formatted content in their browser and clicks the confirmation button to check the content. If corrections are needed, they use the input fields to make the necessary changes and then click the "Approve" button.

[1502] server

[1503] The server receives the approved formatted content and saves it to the history file. The saved history is appended to a text file (e.g., transcription_history.txt).

[1504] Specific saving actions

[1505] The server receives the approved format content and appends it to the history file as follows:

[1506] ---

[1507] Date and Time: 2023-10-15 14:35:00

[1508] Type: Customer support

[1509] Contents:

[1510] Hello, this is Customer Service. How can I help you today?

[1511] ---

[1512] Example of a prompt

[1513] Examples of prompt messages in this system are as follows:

[1514] Please transcribe the audio file customer_call.wav using a speech recognition API and format it into the following format:

[1515] ---

[1516] Date and Time: YYYY-MM-DD HH:MM:SS

[1517] Type: Customer support

[1518] Contents:

[1519] <Transcripted text>

[1520] ---

[1521] In this way, the automatic transcription and formatting of recorded audio is streamlined, and it becomes possible to manage it with accuracy and consistency.

[1522] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1523] Step 1: Receiving the recorded audio

[1524] The server receives the recorded audio file sent as an HTTP request from the user.

[1525] Input: A recorded audio file uploaded by the user (e.g., customer_call.wav).

[1526] Data processing: The server temporarily stores the received file in memory and then moves the file to the specified directory.

[1527] Output: Audio files saved in the specified directory on the server (e.g., / uploads / ).

[1528] Specific operation: The user uploads an audio file through a browser form, the server receives it, and saves it to the / uploads / directory.

[1529] Step 2: Transcribe the recorded audio.

[1530] The server sends the stored audio file to a speech recognition API, where it is converted into text data.

[1531] Input: Audio file stored on the server (e.g., / uploads / customer_call.wav).

[1532] Data processing: The server uses the Python pydub library to read the audio file and send a request to the speech recognition API.

[1533] Output: Text data returned from the speech recognition API (e.g., "Hello, this is customer service. How can I help you today?").

[1534] Specific operation: The server sends the audio file to the Google Speech Recognition API and saves the returned text data to a temporary variable called `transcription_text`.

[1535] Step 3: Format of the transcript

[1536] The device receives text data and formats it into a specific format.

[1537] Input: Text data sent from the server (e.g., "Hello, this is Customer Service. How can I help you today?").

[1538] Data processing: Apply a format to the text data that includes the recording date and time, recording type, and transcript content.

[1539] Output: Formatted text data (e.g., in the format shown below).

[1540] ---

[1541] Date and Time: 2023-10-15 14:35:00

[1542] Type: Customer support

[1543] Contents:

[1544] Hello, this is Customer Service. How can I help you today?

[1545] ---

[1546] Specific operation: The terminal receives text data from the server, applies a format to it, shapes it, and saves it to a temporary variable called `formatted_text`.

[1547] Step 4: Check the format content

[1548] The user confirms the formatted content on the device screen.

[1549] Input: Formatted text data displayed on the terminal.

[1550] Data processing: The user reviews the content and makes corrections if necessary. The corrected data is then formatted again.

[1551] Output: The final text data approved or modified by the user.

[1552] Specific action: The user reviews the formatted content using a browser, makes corrections as needed, and clicks the "Approve" button.

[1553] Step 5: Save the confirmation details

[1554] The server saves the format content approved by the user to a history file.

[1555] Input: User-approved, formatted text data.

[1556] Data processing: Approved text data is appended to the history file (e.g., transcription_history.txt).

[1557] Output: Formatted text saved in the history file.

[1558] Specific operation: The server receives the approved text data and appends it to the history file in the following format.

[1559] ---

[1560] Date and Time: 2023-10-15 14:35:00

[1561] Type: Customer support

[1562] Contents:

[1563] Hello, this is Customer Service. How can I help you today?

[1564] ---

[1565] (Application Example 1)

[1566] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1567] As autonomous vehicles become more widespread, there is a growing need to efficiently manage, record, and store the voice interactions between passengers and vehicles as a history. However, existing systems rely on manual management of recorded voices, which is time-consuming and labor-intensive, and makes it difficult to maintain accuracy and consistency. Furthermore, there is a lack of adequate systems for effectively utilizing voice interactions within vehicles after the fact.

[1568] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1569] In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for verifying the formatted content, means for saving the verified content, and, if the recorded audio is voice dialogue in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors, and means for displaying the transcribed content on the vehicle's display device or a portable information terminal. This automates the recording, transcription, formatting, verification, and saving of voice dialogue in autonomous vehicles, thereby improving the quality of operation management and customer service.

[1570] An "autonomous vehicle" is a vehicle that can operate autonomously without driver intervention.

[1571] An "acoustic sensor" is a device used to detect and record sound, such as a microphone installed in a vehicle.

[1572] A "portable information terminal" refers to a portable electronic device, such as a smartphone or tablet.

[1573] "Transcription" refers to the process of converting recorded audio into text data.

[1574] A "speech recognition API" is a programmatic interface for converting speech data into text data, and includes cloud-based and local-based services.

[1575] "Formatting to a specific format" refers to organizing and arranging transcribed text data according to a predetermined format and layout.

[1576] "Vehicle display devices" refer to display devices such as monitors and touch panels installed inside a vehicle.

[1577] "User interface" refers to the means by which a device or system and a user exchange information, and includes displays and input devices.

[1578] "Means of verification" refers to processes or devices that allow users to check formatted text content and make corrections or approvals as needed.

[1579] "Means of preservation" refers to the processes and systems used to record and retain verified text data in databases or files.

[1580] This invention relates to a system for recording, transcribing, and formatting voice dialogue within an autonomous vehicle as a history. This system streamlines the management of recorded audio and ensures accurate history recording.

[1581] System Overview

[1582] The system consists of the following elements:

[1583] 1. Means for receiving recorded audio

[1584] 2. Methods for transcribing recorded audio into text.

[1585] 3. Means for formatting the transcribed content into a specific format.

[1586] 4. Means of verifying formatted content

[1587] 5. Means for saving the confirmed information

[1588] 6. If the recorded audio is for voice interaction in an autonomous vehicle, means for receiving the recorded audio using the vehicle's acoustic sensors.

[1589] 7. Means for displaying the transcribed content on a vehicle display device or portable information terminal.

[1590] Means for receiving recorded audio

[1591] Acoustic sensors (such as in-vehicle microphones) installed in autonomous vehicles record passenger voices. The recorded voice data is transmitted to the onboard computer and temporarily stored. This prepares the system for processing voice interactions with passengers.

[1592] Methods for transcribing recorded audio into text

[1593] The server uses a speech recognition API (e.g., a speech recognition cloud service) to transcribe the recorded audio. Specifically, it reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text is temporarily stored.

[1594] A method for formatting transcribed content into a specific format.

[1595] The device formats the transcribed text data into a specific format. Specifically, it includes the recording date and time, the type of recording (customer interaction or passenger interaction), and the transcribed text content. The formatted text is temporarily stored.

[1596] Means of verifying formatted content

[1597] The user reviews the formatted text content via the vehicle's display device or a mobile device. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1598] Means for saving confirmed information

[1599] The server saves the formatted text approved by the user to a history file. Specifically, it appends the text to the history file (for example, transcription_history.txt).

[1600] Specific example

[1601] For example, suppose a passenger's voice interaction inside an autonomous vehicle is recorded and saved as a file named customer_interaction.wav. Let's say this audio file was uploaded on October 15, 2023 at 14:35:00. The system processes this audio file and generates text formatted as follows:

[1602] ---

[1603] Date and Time: 2023-10-15 14:35:00

[1604] Contents:

[1605] Hello, this is an autonomous vehicle. Where are you going today?

[1606] ---

[1607] After the user reviews and approves this formatted content, the server saves it to a history file. The saved content will look like this:

[1608] ---

[1609] Date and Time: 2023-10-15 14:35:00

[1610] Contents:

[1611] Hello, this is an autonomous vehicle. Where are you going today?

[1612] ---

[1613] This automates the management of voice interaction within autonomous vehicles, saving time and effort while ensuring accuracy and consistency.

[1614] Example prompts for generative AI models

[1615] "Create a sample Python program that records voice conversations with customers, transcribes them using a voice recognition cloud service, and saves them in a specific format."

[1616] As described above, this system automates the recording, transcription, formatting, verification, and saving of voice conversations within autonomous vehicles, thereby improving the quality of operational management and customer service.

[1617] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1618] Step 1:

[1619] Receiving recorded audio

[1620] The server receives recorded audio data through acoustic sensors installed in the autonomous vehicle. Specifically, a microphone captures the passenger's voice and sends that audio data to the server. The input is audio data, and the output is a temporarily stored audio file.

[1621] Step 2:

[1622] Transcription of recorded audio

[1623] The server uses a speech recognition API to convert the received audio data into text data. Specifically, a Python script is used to read the audio file and send a request to the speech recognition API. The input is an audio file, and the output is the transcribed text data.

[1624] Step 3:

[1625] Formatting of the transcribed content

[1626] The terminal formats the transcribed text data into a specific format. Specifically, it adds the recording date and time, and the type of recording (e.g., customer service, passenger interaction) to the text data and compiles it into a predetermined format. The input is the text data and the current date and time, and the output is the formatted text.

[1627] Step 4:

[1628] Checking the formatted content

[1629] The user reviews the formatted content via the vehicle's display device or a mobile device. Specifically, the formatted text is displayed on the screen of a display or smartphone, allowing the user to review the content and make corrections as needed. The input is the formatted text, and the output is the reviewed and corrected text.

[1630] Step 5:

[1631] Saving of confirmed information

[1632] The server saves the formatted text confirmed by the user to a history file. Specifically, it appends text data to the history file (e.g., transcription_history.txt). The input is the confirmed text data, and the output is the updated history file.

[1633] Step 6:

[1634] Managing history files

[1635] The server periodically manages history files, creates backups as needed, and archives old data. Specifically, it uses scripts to create file backups and manage file sizes. The input is history files, and the output is organized history data.

[1636] The above outlines the specific processing steps of the program, detailing the input, data processing, and output in each step. This process efficiently records, transcribes, formats, verifies, and saves voice conversations within autonomous vehicles.

[1637] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1638] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[1639] System Overview

[1640] The system consists of the following elements:

[1641] 1. Means for receiving recorded audio

[1642] 2. Methods for transcribing recorded audio into text.

[1643] 3. Means for formatting the transcribed content into a specific format.

[1644] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[1645] 5. Means of verifying formatted content

[1646] 6. Means for saving the confirmed information

[1647] Means for receiving recorded audio

[1648] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[1649] Methods for transcribing recorded audio into text

[1650] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1651] A method for formatting transcribed content into a specific format.

[1652] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1653] An emotion engine that recognizes user emotions from transcribed content.

[1654] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[1655] Means of verifying formatted content

[1656] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1657] Means for saving confirmed information

[1658] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[1659] Specific example

[1660] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recording file is processed by the system, and the following formatted text is generated:

[1661] ---

[1662] Date and Time: 2023-10-15 14:35:00

[1663] Type: Customer support

[1664] Contents:

[1665] Hello, this is Customer Service. How can I help you today?

[1666] Emotion: Joy

[1667] ---

[1668] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[1669] ---

[1670] Date and Time: 2023-10-15 14:35:00

[1671] Type: Customer support

[1672] Contents:

[1673] Hello, this is Customer Service. How can I help you today?

[1674] Emotion: Joy

[1675] ---

[1676] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1677] The following describes the processing flow.

[1678] Processing flow

[1679] Step 1:

[1680] The user uploads a recorded audio file (e.g., customer_call.wav) to the system. The server receives this file and stores it temporarily.

[1681] Step 2:

[1682] The server uses a speech recognition API (for example, Google's speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1683] Step 3:

[1684] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1685] Step 4:

[1686] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The emotion engine uses natural language processing techniques to extract emotions (e.g., joy, anger, sadness) from the text data. The recognized emotion information is added to the text data.

[1687] Step 5:

[1688] The user reviews the combined formatted text data and recognized sentiment information on their device screen. The user verifies the accuracy of the content, makes corrections as needed, and then approves it.

[1689] Step 6:

[1690] The server saves the user-approved formatted text and sentiment information to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it has been saved correctly.

[1691] A concrete example of the process is as follows:

[1692] For example, suppose a customer service recording file, customer_call.wav, was uploaded on 2023-10-15 at 14:35:00. This recording file is processed by the system, generating text and sentiment information in the following format:

[1693] ---

[1694] Date and Time: 2023-10-15 14:35:00

[1695] Type: Customer support

[1696] Contents:

[1697] Hello, this is Customer Service. How can I help you today?

[1698] Emotion: Joy

[1699] ---

[1700] After the user reviews and approves this formatted content and sentiment information, the server saves it to a history file. The saved content will look like this:

[1701] ---

[1702] Date and Time: 2023-10-15 14:35:00

[1703] Type: Customer support

[1704] Contents:

[1705] Hello, this is Customer Service. How can I help you today?

[1706] Emotion: Joy

[1707] ---

[1708] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1709] (Example 2)

[1710] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1711] Conventional audio recording management systems often involve manual transcription, formatting, and emotional analysis of recordings, which is time-consuming, labor-intensive, and can lack accuracy. Furthermore, the lack of a function to analyze and record changes in emotion makes detailed conversation history management difficult.

[1712] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing it, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, and means for saving the confirmed content. This enables automatic transcription of recorded audio, formatting, emotion recognition, and history saving.

[1713] "Recorded audio" refers to audio data uploaded by users.

[1714] "Means of receiving" refers to the function by which a server receives audio data from a user and temporarily stores it.

[1715] "Transcribing" refers to a function that analyzes recorded audio and converts it into text data.

[1716] A "speech recognition API" is an interface for external services that convert speech data into text data.

[1717] "Methods for formatting" refer to functions that arrange text data into a specific format.

[1718] "Means of recognizing emotions" refers to a function that analyzes transcribed text data and audio data to identify the user's emotions.

[1719] A "sentiment analysis engine" is a software component used to analyze emotions from text and audio data.

[1720] A "means of verification" refers to a function that allows users to view formatted content and make corrections or approvals.

[1721] The "means of saving" refer to the function of saving the verified text data to a history file.

[1722] A "user interface" refers to the screens and control devices that allow a user to interact with and operate a system.

[1723] Defining it in this way clarifies the role of each function in the system.

[1724] This invention relates to a system that automatically transcribes recorded audio, formats the content, and saves it as a history, as well as a system that uses an emotion engine to recognize the user's emotions and include them in the format. This system streamlines the management of recorded audio, enables accurate history recording, and can analyze and record changes in emotions during conversations.

[1725] System Overview

[1726] The system consists of the following elements:

[1727] 1. Means for receiving recorded audio

[1728] 2. Methods for transcribing recorded audio into text.

[1729] 3. Means for formatting the transcribed content into a specific format.

[1730] 4. An emotion engine that recognizes the user's emotions from the transcribed content.

[1731] 5. Means of verifying formatted content

[1732] 6. Means for saving the confirmed information

[1733] Means for receiving recorded audio

[1734] The server receives the recorded audio file and stores it temporarily. This prepares the system to process the recorded audio uploaded by the user.

[1735] Methods for transcribing recorded audio into text

[1736] The server uses a speech recognition API (for example, a common speech recognition API) to transcribe the uploaded audio file. Specifically, the server reads the audio file, sends a request to the API, and converts the audio data into text data. The transcribed text data is stored in a temporary variable.

[1737] A method for formatting transcribed content into a specific format.

[1738] The terminal formats the transcribed text data into a specific format. Specifically, it formats it to include the recording date and time, the type of recording (e.g., customer service or government service), and the transcribed text content. The formatted text is stored in a temporary variable for subsequent processing.

[1739] An emotion engine that recognizes user emotions from transcribed content.

[1740] The emotion engine installed on the server analyzes the transcribed text and audio data to recognize the user's emotions. The recognized emotion information is included in the text data and added to the format.

[1741] Means of verifying formatted content

[1742] The user reviews the formatted text content on their device screen. The user verifies that the formatted content is accurate, makes any necessary corrections, and then approves it.

[1743] Means for saving confirmed information

[1744] The server saves the formatted text approved by the user to a history file. Specifically, the server appends the formatted text to the history file (for example, transcription_history.txt). The server verifies the saved content and ensures that it was saved correctly.

[1745] Specific example

[1746] For example, suppose a recorded audio file of a customer interaction, customer_call.wav, was uploaded on 2023-10-15 14:35:00. This recorded audio file is processed by the system, and text formatted as follows is generated:

[1747] ---

[1748] Date and Time: 2023-10-15 14:35:00

[1749] Type: Customer support

[1750] Contents:

[1751] Hello, this is Customer Service. How can I help you today?

[1752] Emotion: Joy

[1753] ---

[1754] In this example, the emotion engine recognizes the emotion of "joy" from the text data and includes it in the format. After the user reviews and approves this formatted content, the server saves it to a history file. The saved content looks like this:

[1755] ---

[1756] Date and Time: 2023-10-15 14:35:00

[1757] Type: Customer support

[1758] Contents:

[1759] Hello, this is Customer Service. How can I help you today?

[1760] Emotion: Joy

[1761] ---

[1762] This not only automates the management of recorded audio, but also records changes in the user's emotions, enabling more detailed history management.

[1763] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1764] Step 1:

[1765] The server receives the recorded audio file and stores it temporarily.

[1766] Input: Audio file sent by the user.

[1767] Specific operation: The user uploads an audio file from their device, and the server receives and temporarily stores the file.

[1768] Output: Temporarily saved audio file.

[1769] Step 2:

[1770] The server transcribes the audio file into text.

[1771] Input: Temporarily saved audio file.

[1772] Specific operation: The server sends an audio file to a speech recognition API (a common speech recognition service) and retrieves text data.

[1773] Output: Transcripted text data.

[1774] Step 3:

[1775] The device formats the transcribed text data into a specific format.

[1776] Input: Transcribed text data.

[1777] Specific operation: The device acquires text data and formats it into a format that includes the recording date and time, recording type, and transcript content.

[1778] Output: Formatted text data.

[1779] Step 4:

[1780] The server uses an emotion engine to recognize the user's emotions from the transcribed text data.

[1781] Input: Formatted text data.

[1782] Specific operation: The server uses a sentiment analysis engine to analyze text data and recognize sentiment information. The recognized sentiment information is added to the format.

[1783] Output: Formatted text data with added sentiment information.

[1784] Step 5:

[1785] The user checks the formatted text content on the device screen.

[1786] Input: Formatted text data with added sentiment information.

[1787] Specific operation: The device displays text data to the user, who then reviews the content. The user makes corrections as needed and then approves it.

[1788] Output: Text data reviewed and modified by the user.

[1789] Step 6:

[1790] The server saves the formatted text approved by the user to a history file.

[1791] Input: Text data reviewed and modified by the user.

[1792] Specific action: The server appends the verified text data to the history file. For example, the data is saved to transcription_history.txt.

[1793] Output: Text data saved in the history file.

[1794] Through the above processing, a series of tasks such as transcribing recorded audio, formatting, emotion recognition, and history saving are performed efficiently.

[1795] (Application Example 2)

[1796] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1797] Conventional audio recording management systems often require manual transcription and sentiment analysis, which is time-consuming and labor-intensive. Furthermore, they lack mechanisms for accurately understanding user emotions, resulting in limited quality in reviews and feedback.

[1798] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for receiving recorded audio, means for transcribing the recorded audio, means for formatting the transcribed content into a specific format, means for recognizing the user's emotions from the transcribed content, means for confirming the formatted content, means for saving the confirmed content, and means for managing the saved content as history. This not only automates the management of recorded audio, but also records changes in the user's emotions in detail, enabling richer and more accurate reviews and feedback.

[1799] "Means for receiving recorded audio" refers to a mechanism for incorporating audio data provided by the user into the system.

[1800] "Methods for transcribing recorded audio" refer to the technologies and processes used to convert recorded audio data into text data.

[1801] "Methods for formatting transcribed content into a specific format" refers to methods of organizing and arranging transcribed text data according to a predetermined format.

[1802] "Methods for recognizing user emotions from transcribed content" refer to algorithms and technologies that analyze text and audio data to identify the user's emotional state.

[1803] "Means of verifying formatted content" refers to interfaces or methods that allow users to verify formatted text data and make corrections as needed.

[1804] "Means for saving confirmed content" refers to a mechanism for recording and saving the format content confirmed by the user in a database or file system.

[1805] "Means for managing saved content as history" refers to systems and processes for organizing saved data in chronological order so that it can be accessed and referenced later.

[1806] Modes for carrying out the invention

[1807] The invention is a system that automatically transcribes recorded audio, formats the content, saves it as a history, and further recognizes the user's emotions using an emotion engine, incorporating them into the format. This system is designed to allow users to easily manage their recorded audio and is particularly useful in content distribution services.

[1808] System Overview

[1809] The system consists of the following elements:

[1810] 1. Means for receiving recorded audio

[1811] The server has an interface for users to provide audio data, receives recorded audio files, and temporarily stores them. Users record audio reviews using a smartphone application and send them to the server.

[1812] 2. Methods for transcribing recorded audio into text.

[1813] The server transcribes the recorded audio file using the Google Speech-to-Text API. This API provides highly accurate speech recognition technology and converts the audio data into text data. The transcribed text data is temporarily stored on the server.

[1814] 3. Means for formatting the transcribed content into a specific format.

[1815] The server formats the transcribed text data into a specific format. Specifically, it formats it to include information such as the review date and time, content name, and username. The formatted text data is then temporarily stored in a database.

[1816] 4. Means of recognizing user emotions from transcribed content

[1817] The server uses Microsoft Azure's Text Analytics API to recognize user emotions from transcribed text data. This API analyzes emotions (joy, sadness, surprise, etc.) within the text with high accuracy. The recognized emotion information is added to the format.

[1818] 5. Means of verifying formatted content

[1819] Users can review the formatted content through a smartphone application. They can then review the content, make corrections as needed, and finally approve it.

[1820] 6. Means for saving the confirmed information

[1821] The server saves the format content approved by the user to a history file. Specifically, it appends the confirmed data to the database and manages the saved content as history.

[1822] Specific example

[1823] Specific examples are given below.

[1824] A user records a movie review using the "Sentiment Analysis Content Guide" app and sends it to the server. Assume the recorded audio review contains the following:

[1825] "The movie I saw yesterday was very moving. The last scene was especially wonderful."

[1826] This audio review is processed by the server and a formatted text like the following is generated:

[1827] ---

[1828] Date and Time: 2023-10-15 17:45:00

[1829] Content name: Movie A

[1830] Username: user123

[1831] Contents:

[1832] The movie I saw yesterday was very moving. The last scene was especially wonderful.

[1833] Emotion: Joy

[1834] ---

[1835] The user reviews this formatted content in the application and saves it to the database after approval.

[1836] Example of a prompt

[1837] Please enter your thoughts on the movie you saw yesterday using voice input. Example: "The movie I saw yesterday was very moving! There were so many great scenes."

[1838] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1839] Step 1:

[1840] The user records an audio review within the application using their smartphone. The input is the user's voice data (e.g., "I was very moved by the movie I saw yesterday."), which is saved as digital data using the smartphone's built-in microphone. The output is the recorded audio file (e.g., "review_20231015.wav").

[1841] Step 2:

[1842] The server receives the recorded audio file. The audio file sent from the smartphone is uploaded to the server and temporarily stored. The input is the audio file sent by the user from their smartphone. The output is an audio file that is stored on the server.

[1843] Step 3:

[1844] The server uses the Google Speech-to-Text API to transcribe the received audio file. The input is an audio file stored on the server, which is sent to the API. This API converts the audio data into text data. The output is the transcribed text data (e.g., "The movie I saw yesterday was very moving.").

[1845] Step 4:

[1846] The server formats the transcribed text data into a specific format. Metadata such as review date and time, content name, and username are also added at this time. The input is the transcribed text data and associated metadata. The output is the formatted text data (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A Username: user123 Content: I was very moved by the movie I saw yesterday. ---").

[1847] Step 5:

[1848] The server uses Microsoft Azure's Text Analytics API to recognize the user's emotions from formatted text data. The input is formatted text data, which is sent to the API. This API analyzes the emotions within the text (e.g., "joy"). The output is formatted data with emotion information added (e.g., "--- Date and Time: 2023-10-15 17:45:00 Content Name: Movie A User Name: user123 Content: I was very moved by the movie I saw yesterday. Emotion: Joy ---").

[1849] Step 6:

[1850] Users review formatted text content through a smartphone application. The input is formatted data with added sentiment information, displayed on the application. Users review the content, make corrections as needed, and finally approve it. The output is the text data reviewed and corrected by the user.

[1851] Step 7:

[1852] The server saves the verified text data to the database. The input consists of user-approved formatted content, which is then recorded in the database. The output generates saved data that is managed as history.

[1853] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1854] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1855] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1856] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1857] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1858] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1859] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1860] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1861] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1862] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1863] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1864] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1865] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1866] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1867] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1868] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1869] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1870] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1871] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1872] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1873] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1874] The following is further disclosed regarding the embodiments described above.

[1875] (Claim 1)

[1876] A means of receiving recorded audio,

[1877] A means for transcribing the aforementioned recorded audio,

[1878] A means for formatting the transcribed content into a specific format,

[1879] A means for verifying the formatted content,

[1880] Means for saving the confirmed contents,

[1881] A system that includes this.

[1882] (Claim 2)

[1883] The system according to claim 1, characterized in that the means for transcribing uses a speech recognition API.

[1884] (Claim 3)

[1885] The system according to claim 1, characterized in that the means for verification is performed through a user interface.

[1886] "Example 1"

[1887] (Claim 1)

[1888] A means of receiving recorded audio,

[1889] A means for transcribing the aforementioned recorded audio using a speech recognition API,

[1890] A means for formatting the transcribed content into a specific format,

[1891] A means for confirming the formatted content through a user interface,

[1892] Means for saving the confirmed information to a history file,

[1893] A system that includes this.

[1894] (Claim 2)

[1895] The system according to claim 1, characterized in that a specific format includes the date and time of recording, the type of recording, and the transcribed content.

[1896] (Claim 3)

[1897] The system according to claim 1, characterized in that recorded audio is received when it is sent from the user's terminal to the server as an HTTP request.

[1898] (Claim 4)

[1899] The system according to claim 1, characterized in that the history file is a text file.

[1900] "Application Example 1"

[1901] (Claim 1)

[1902] A means of receiving recorded audio,

[1903] A means for transcribing the aforementioned recorded audio,

[1904] A means for formatting the transcribed content into a specific format,

[1905] A means for verifying the formatted content,

[1906] Means for saving the confirmed contents,

[1907] If the recorded audio is a voice dialogue in an autonomous vehicle, the means for receiving the recorded audio using the vehicle's acoustic sensors,

[1908] Means for displaying the transcribed content on a vehicle display device or a portable information terminal,

[1909] A system that includes this.

[1910] (Claim 2)

[1911] The system according to claim 1, characterized in that the means for transcribing uses a speech recognition API.

[1912] (Claim 3)

[1913] The system according to claim 1, characterized in that the means for confirmation is performed through a user interface and displayed on the vehicle's display device or a portable information terminal.

[1914] "Example 2 of combining an emotion engine"

[1915] (Claim 1)

[1916] A means of receiving recorded audio,

[1917] A means for transcribing the aforementioned recorded audio,

[1918] A means for formatting the transcribed content into a specific format,

[1919] A means of recognizing the user's emotions from the transcribed content,

[1920] A means for verifying the formatted content,

[1921] Means for saving the confirmed contents,

[1922] A system that includes this.

[1923] (Claim 2)

[1924] The system according to claim 1, characterized in that the means for transcribing uses a speech recognition API.

[1925] (Claim 3)

[1926] The system according to claim 1, characterized in that the means for recognizing the emotion uses an emotion analysis engine.

[1927] (Claim 4)

[1928] The system according to claim 1, characterized in that the means for verification is performed through a user interface.

[1929] "Application example 2 when combining with an emotional engine"

[1930] (Claim 1)

[1931] A means of receiving recorded audio,

[1932] A means for transcribing the aforementioned recorded audio,

[1933] A means for formatting the transcribed content into a specific format,

[1934] A means of recognizing the user's emotions from the transcribed content,

[1935] A means for verifying the formatted content,

[1936] Means for saving the confirmed contents,

[1937] A means for managing the aforementioned saved content as a history,

[1938] A system that includes this.

[1939] (Claim 2)

[1940] The system according to claim 1, characterized in that the means for transcribing uses a speech recognition API.

[1941] (Claim 3)

[1942] The system according to claim 1, characterized in that the means for recognizing the emotion uses a text analytics API. [Explanation of Symbols]

[1943] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving recorded audio, A means for transcribing the aforementioned recorded audio, A means for formatting the transcribed content into a specific format, A means for verifying the formatted content, Means for saving the confirmed contents, A system that includes this.

2. The system according to claim 1, characterized in that the means for transcribing uses a speech recognition API.

3. The system according to claim 1, characterized in that the means for verification is performed through a user interface.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A