How to convert audio data to text

The method enhances voice transcription by identifying and highlighting unconverted dialects or accents in audio data, addressing the challenge of non-standard languages in business negotiations.

JP2026085192APending Publication Date: 2026-05-22TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2024-11-12
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies lack effective voice transcription capabilities for non-standard languages such as dialects and accents in business negotiations, necessitating improved text conversion technology for audio data.

Method used

An information processing device detects specific expressions and characteristic information related to pronunciation in audio data, and outputs text information that identifies portions unable to be converted to standard language, using methods like deep learning and conversion rules.

Benefits of technology

Enhances the understanding of audio data text conversion by clearly indicating unconverted portions, improving the accuracy and comprehensibility of voice transcription.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085192000001_ABST
    Figure 2026085192000001_ABST
Patent Text Reader

Abstract

Improve the technology for converting audio data to text. [Solution] A method for converting audio data to text by an information processing device, comprising: acquiring audio data; detecting a specific expression in the audio data and characteristic information related to the pronunciation of the specific expression; and, if a standard language conversion process cannot be performed to convert the specific expression to the corresponding standard language based on the detected specific expression and characteristic information, outputting text information related to the audio data in a manner that allows for the identification of the portion that could not be converted to standard language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method for text conversion of voice data.

Background Art

[0002] Conventionally, technologies for analyzing the content of business negotiations are known. For example, Patent Document 1 discloses a dialogue analysis system that checks whether a salesperson explains what should be explained and does not state what should not be stated in a business negotiation with a customer. Also, for example, Non-Patent Document 1 discloses a voice recognition technology for the Toyama dialect.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Although Patent Document 1 shows a technology for analyzing the content of business negotiations by machine learning, neither Patent Document 1 nor Non-Patent Document 1 mentions the technology of transcribing voices in business negotiations, etc., that is, the text conversion technology of voice data. In particular, there was room for improvement in the voice transcription technology for voice data including non-standard languages such as dialects and accents. On the other hand, for the analysis of the content of business negotiations, etc., feedback, etc., it is preferable to improve the text conversion technology of voice data. Thus, there was room for improvement in the text conversion technology of voice data in business negotiations, etc. [[ID=at41]]

[0005] In light of these circumstances, the purpose of this disclosure is to improve the technology for converting audio data to text. [Means for solving the problem]

[0006] A method for converting audio data to text according to one embodiment of this disclosure is: A method for converting audio data to text, performed by an information processing device, Acquiring audio data and To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, If a standard language conversion process cannot be performed to convert the detected specific expression into the corresponding standard language based on the detected specific expression and the feature information, the text information relating to the audio data is output in a manner that allows for the identification of the portion that could not be converted into the standard language. Includes. [Effects of the Invention]

[0007] According to one embodiment of this disclosure, the technique for converting speech data to text is improved. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing the schematic configuration of the system according to this embodiment. [Figure 2] This is a flowchart showing the operation of an information processing device. [Modes for carrying out the invention]

[0009] The embodiments of this disclosure will now be described. Referring to Figure 1, the overview and configuration of System 1 according to this embodiment will be described. System 1 according to this embodiment comprises an information processing device 10 and a terminal device 20. The information processing device 10 is a server device installed, for example, in a data center. The terminal device 20 is any device used by a user. These devices are connected to each other so as to be able to communicate via a network 30 such as the Internet. In Figure 1, one information processing device 10 and one terminal device 20 are shown, but System 1 may comprise multiple such devices.

[0010] First, an overview of the method for converting audio data to text according to this embodiment will be described, and further details will be provided later. The audio data may be, for example, audio data from a business negotiation. In this embodiment, the business negotiation is, for example, a negotiation related to the sale of a vehicle, and the offering related to the negotiation is a vehicle, but is not limited to this. For example, the business negotiation may be a meeting aimed at concluding various types of contracts, such as the buying and selling of real estate, the contract of an insurance product, or the sale of a financial product. Furthermore, the offering related to the business negotiation in this embodiment may be goods, services, digital content, licenses, data / information, financial products, real estate, intangible assets, other tradable rights, etc.

[0011] The information processing device 10 acquires audio data. The information processing device 10 also detects specific expressions in the audio data and characteristic information related to the pronunciation of those specific expressions. If the information processing device 10 is unable to perform a standard language conversion process to convert the specific expressions into corresponding standard language based on the detected specific expressions and characteristic information, it outputs text information related to the audio data in a manner that allows for the identification of the portion that could not be converted into standard language.

[0012] Thus, according to this embodiment, the information processing device 10 detects a specific expression in the audio data and characteristic information related to the pronunciation of the specific expression. If standard language conversion processing is not possible, it outputs text information related to the audio data in a manner that makes it possible to identify the portion that could not be converted to standard language. Therefore, when a specific expression cannot be converted, text information is output in a manner that makes it possible to identify that portion, thus improving the audio data text conversion technology in that it is easy to understand that conversion to standard language has not been completed.

[0013] Next, the configurations of the information processing device 10 and the terminal device 20 will be described in detail. As shown in Figure 1, the information processing device 10 comprises a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15. The control unit 11 includes at least one processor. The processor is a general-purpose processor such as a CPU, or a dedicated processor specialized for a specific process. The control unit 11 controls each part of the information processing device 10 and executes processes related to the operation of the information processing device 10. The storage unit 12 includes at least one semiconductor memory, etc. The semiconductor memory is, for example, RAM or ROM. The storage unit 12 functions, for example, as a main memory or auxiliary memory. The storage unit 12 stores data used for the operation of the information processing device 10 and data obtained by the operation of the information processing device 10. The input unit 13 includes at least one input interface. The input interface may be, for example, a physical key, a touchscreen, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts operations to input data used in the operation of the information processing device 10. The output unit 14 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as audio. The output unit 14 outputs data obtained by the operation of the information processing device 10. The communication unit 15 includes at least one external communication interface. The communication interface may be either a wired communication interface or a wireless communication interface. In the case of wired communication, the communication interface is, for example, a LAN or USB. In the case of wireless communication, the communication interface is, for example, an interface compatible with a mobile communication standard such as 5G, or an interface compatible with short-range wireless communication. The communication unit 15 receives data used in the operation of the information processing device 10 and transmits data obtained by the operation of the information processing device 10.

[0014] As shown in Figure 1, the terminal device 20 comprises a control unit 21, a storage unit 22, an input unit 23, an output unit 24, and a communication unit 25. The control unit 21 includes at least one processor. The processor is a general-purpose processor such as a CPU, or a dedicated processor specialized for specific processing. The control unit 21 controls each part of the terminal device 20 and executes processing related to the operation of the terminal device 20. The storage unit 22 includes at least one semiconductor memory, etc. The semiconductor memory is, for example, RAM or ROM. The storage unit 22 functions, for example, as a main memory or auxiliary memory. The storage unit 22 stores data used for the operation of the terminal device 20 and data obtained by the operation of the terminal device 20. The input unit 23 includes at least one input interface. The input interface may be, for example, a physical key, a touchscreen, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 23 accepts operations to input data used for the operation of the terminal device 20. The output unit 24 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as audio. The output unit 24 outputs data obtained by the operation of the terminal device 20. The communication unit 25 includes at least one external communication interface. The communication interface may be either a wired communication interface or a wireless communication interface. In the case of wired communication, the communication interface is, for example, a LAN or USB. In the case of wireless communication, the communication interface is, for example, an interface compatible with a mobile communication standard such as 5G, or an interface compatible with short-range wireless communication. The communication unit 25 receives data used in the operation of the terminal device 20 and transmits data obtained by the operation of the terminal device 20.

[0015] The functions of the information processing device 10 or terminal device 20 are realized by executing a program according to this embodiment on a processor corresponding to the control unit 11 or control unit 21. In other words, the functions of the information processing device 10 or terminal device 20 are realized by software. The program causes the computer to function as the information processing device 10 or terminal device 20 by having the computer execute the operations of the information processing device 10 or terminal device 20. In other words, the computer functions as the information processing device 10 or terminal device 20 by executing the operations of the information processing device 10 or terminal device 20 according to the program. In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-temporary computer-readable media, such as magnetic recording devices and semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending a portable recording medium such as a DVD on which the program is recorded. Alternatively, the program may be distributed by storing the program on the storage of an external server and transmitting the program from the external server to another computer. The program may also be provided as a program product. Some or all of the functions of the information processing device 10 or the terminal device 20 may be implemented by a dedicated circuit corresponding to the control unit 11 or the control unit 21. In other words, some or all of the functions of the information processing device 10 or the terminal device 20 may be implemented by hardware.

[0016] Referring to Figure 2, the operation of the information processing device 10 according to this embodiment will be described. Here, we will mainly describe an example where the business negotiation is related to the sale of a vehicle. First, the control unit 11 of the information processing device 10 acquires voice data (step S10). Any method can be used for the voice data acquisition process. For example, the control unit 11 may acquire voice data from an external device including a terminal device 20 via the communication unit 15 and the network 30. Alternatively, the control unit 11 may acquire voice data via the input unit 13.

[0017] Next, the control unit 11 detects specific expressions in the audio data acquired in step S10, and characteristic information related to the pronunciation of those specific expressions (step S20). Specific expressions include dialects, non-standard languages, slang, and other specific expressions used in actual speech. Any method can be used to detect specific expressions in the audio data. For example, a method may be employed in which keywords, phrases, etc., registered in advance as specific expressions are matched using a speech recognition engine. Alternatively, for example, a method may be employed in which speech patterns are classified using a machine learning model such as deep learning, and predetermined expressions are identified. Characteristic information related to the pronunciation of specific expressions includes, for example, tone information, pronunciation rhythm, pitch variation, and speech speed. Such characteristic information often reflects the speaker's vocal habits, emotions, and intentions, and is considered important for capturing in detail how specific expressions were pronounced. Any method can be used to detect characteristic information related to the pronunciation of specific expressions. For example, a method may be employed in which tone variation is analyzed using a pitch detection algorithm. Alternatively, methods such as spectrogram analysis may be used to extract temporal and frequency features of the audio signal. Furthermore, machine learning models such as deep learning may be utilized to comprehensively capture the characteristics of the speech, thereby detecting feature information related to speech production.

[0018] Next, the control unit 11 determines whether or not a standard language conversion process is possible to convert the specific expression into the corresponding standard language based on the detected specific expression and feature information (step S30). Any method may be used for this determination process. For example, suppose the standard language conversion process is a process that converts the specific expression into the standard language based on the detected pair of specific expression and feature information and the conversion rule between the non-standard language and the standard language. In this case, if the pair of specific expression and feature information exists in the conversion rule, the control unit 11 determines that the standard language conversion process is possible. On the other hand, if the pair of specific expression and feature information does not exist in the conversion rule, the control unit 11 determines that the standard language conversion process is not possible.

[0019] When it is determined in step S30 that the standard language conversion process is possible, the control unit 11 converts the detected specific expression into the corresponding standard language based on the detected specific expression and the feature information (step S40). Any method may be adopted for the conversion process to the standard language. For example, the control unit 11 may perform a process of converting the detected specific expression into the standard language based on the pair of the detected specific expression and the feature information and the conversion rule between the non-standard language and the standard language. The conversion rule may be stored in the storage unit 12, for example, and the control unit 11 may execute the above conversion process by referring to the conversion rule in the storage unit 12. In this case, the conversion rule may be a conversion table between the pair of the specific expression and the feature information and the corresponding standard language. The corresponding standard language is the standard language to be converted for each specific expression. In other words, the corresponding standard language is the language expression obtained as a result of the conversion process. Alternatively, the control unit 11 may execute the conversion process to the standard language by using a machine learning model that executes the standard language conversion process based on the detected specific expression and the feature information.

[0020] Subsequent to step S40, the control unit 11 outputs text information related to the voice data (step S50). The text information is information obtained by converting the voice data into text, and among the voice data, it is information in which the above-mentioned specific expression has been converted into standard language. In other words, the output text information is a sentence in standard language obtained as a result of voice recognition, and is text after the detection of the specific expression and the conversion process to standard language. For example, for voice data including the Aichi dialect "この道はとーめんこ", with respect to the specific expression "とーめんこ" and using the conversion rule based on the tone, the control unit 11 outputs the text information "この道は通行止め". Thus, even if the specific expression is a dialect, turn of phrase, etc. used in different regions and cultures, it can be converted into standard language by using an appropriate conversion rule, and the text information is provided in a form understandable to a wider range of users. Any method can be adopted for the output process of the text information. For example, the control unit 11 may transmit data to the terminal device 20 via the communication unit 15, and the output unit 24 of the terminal device 20 may output the determination result through a user interface for display output. Alternatively, the control unit 11 may output the determination result through a user interface that the output unit 14 outputs for display output.

[0021] If it is determined in step S30 that the standard language conversion process is not possible, the control unit 11 outputs text information related to the voice data in a manner that the portion where the standard language conversion process could not be performed can be discriminated (step S60). Any method can be adopted for the output process of the text information. For example, the control unit 11 may transmit data to the terminal device 20 via the communication unit 15, and the output unit 24 of the terminal device 20 may output the determination result through a user interface for display output. Alternatively, the control unit 11 may output the determination result through a user interface that the output unit 14 outputs for display output. The manner in which the portion where the standard language conversion process could not be performed can be discriminated includes, for example, highlighting the specific expression for which the standard language conversion process could not be performed, displaying an underline, etc.

[0022] With this configuration, the information processing device 10 detects a specific expression in the audio data and characteristic information related to the pronunciation of that specific expression. If standard language conversion processing is not possible, it outputs text information related to the audio data in a manner that allows identification of the portion that could not be converted to standard language. Therefore, when a specific expression cannot be converted, text information is output in a manner that allows identification of that portion, making it easy to understand that conversion to standard language has not been completed, thus improving the audio data text conversion technology.

[0023] While this disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art may make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of this disclosure. For example, the functions, etc., included in each component or step can be rearranged in a logically consistent manner, and multiple components or steps can be combined into one or divided into two.

[0024] For example, the control unit 11 of the information processing device 10 may perform annotation processing on specific expressions that could not be converted to standard language. Any method can be used for annotation processing. For example, the control unit 11 may accept user input and perform manual annotation processing on specific expressions. Alternatively, the control unit 11 may automatically perform annotation processing by inputting specific expressions that could not be converted to standard language into a Large Language Model (LLM) or the like to estimate the standard language for those specific expressions. As a result, the specific expressions may be corrected based on the annotations, and the text information related to the audio data may be saved and stored. The control unit 11 may also update the conversion rules based on the annotation processing and save the updated conversion rules in the storage unit 12. Alternatively, the control unit 11 may retrain the machine learning model related to the standard language conversion processing based on the annotation processing.

[0025] Furthermore, in the embodiment described above, it is also possible to have an embodiment in which the configuration and operation of the information processing device 10 are distributed among multiple computers that can communicate with each other. [Explanation of Symbols]

[0026] 10: Information processing device, 11: Control unit, 12: Storage unit, 13: Input unit, 14: Output unit, 15: Communication unit, 20: Terminal device, 21: Control unit, 22: Storage unit, 23: Input unit, 24: Output unit, 25: Communication unit, 30: Network

Claims

1. A method for converting audio data to text, performed by an information processing device, Acquiring audio data and To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, If a standard language conversion process cannot be performed to convert the detected specific expression into the corresponding standard language based on the detected specific expression and the feature information, the text information relating to the audio data is output in a manner that allows for the identification of the portion that could not be converted into the standard language. A method for converting audio data containing text into text.

2. A method for converting audio data to text according to claim 1, The method for converting speech data to text is a process in which the standard language conversion process converts a specific expression into a standard language based on a detected specific expression and a pair of feature information, and a conversion rule between a non-standard language and a standard language.

3. A method for converting audio data to text according to claim 1, The characteristic information related to the aforementioned vocalization is tone information, and the method is for converting speech data into text.

4. A method for converting audio data to text according to claim 1, The aforementioned specific expression is a method for converting audio data to text, including dialects and non-standard languages.

5. A method for converting audio data to text according to claim 1, further A method for converting speech data to text, including performing annotations on specific expressions that could not be converted to standard Japanese.