Method for converting audio data to text, information processing device, and program
The information processing device enhances voice transcription by detecting and converting non-standard language expressions into standard language, improving the analysis of business negotiations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-22
AI Technical Summary
Existing technologies lack effective methods for transcribing voice data, particularly in business negotiations, especially for non-standard languages such as dialects and accents, which hinders the analysis of negotiation content and feedback.
An information processing device acquires audio data, detects specific expressions and characteristic pronunciation information, and converts them into standard language using a conversion method based on detected expressions and feature information.
Improves voice transcription by accurately extracting and converting specific expressions into standard language, enhancing the understanding and analysis of business negotiations.
Smart Images

Figure 2026085178000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for text conversion of voice data, an information processing device, and a program.
Background Art
[0002] Conventionally, techniques for analyzing the content of business negotiations have been known. For example, Patent Document 1 discloses a dialogue analysis system that checks whether a salesperson explains what should be explained and does not mention what should not be mentioned in a business negotiation with a customer. Also, for example, Non-Patent Document 1 discloses a voice recognition technology for the Toyama dialect.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Although Patent Document 1 shows a technique for analyzing the content of business negotiations by machine learning, neither Patent Document 1 nor Non-Patent Document 1 mentions the technique of transcribing voice in business negotiations or the like, that is, the text conversion technique of voice data. In particular, there was room for improvement in the voice transcription technique for voice data including non-standard languages such as dialects and accents. On the other hand, for the analysis of the content of business negotiations or the like, feedback, etc., improvement of the text conversion technique of voice data is preferable. Thus, there was room for improvement in the text conversion technique of voice data in business negotiations or the like.
[0005] In light of these circumstances, the purpose of this disclosure is to improve the technology for converting audio data to text. [Means for solving the problem]
[0006] A method for converting audio data to text according to one embodiment of this disclosure is: A text conversion of audio data performed by an information processing device, Acquiring audio data and To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, Based on the detected specific expression and the characteristic information, the specific expression is converted into standard language. To output text information related to the aforementioned audio data, Includes. [Effects of the Invention]
[0007] According to one embodiment of this disclosure, the technique for converting speech data to text is improved. [Brief explanation of the drawing]
[0008] [Figure 1] This is a block diagram showing the schematic configuration of the system according to this embodiment. [Figure 2] This is a flowchart showing the operation of an information processing device. [Modes for carrying out the invention]
[0009] The embodiments of this disclosure will be described below.
[0010] (Summary of the embodiment) Referring to Figure 1, the overview and configuration of System 1 according to this embodiment will be described. System 1 according to this embodiment comprises an information processing device 10 and a terminal device 20. The information processing device 10 and the terminal device 20 are connected to a network 30, which includes, for example, a mobile communication network and the Internet.
[0011] The information processing device 10 is, for example, a server device installed in a data center or the like. For example, the information processing device 10 is a server belonging to a cloud computing system or other computing system. Although Figure 1 shows an example where system 1 has one information processing device 10, it is not limited to this. System 1 may have two or more information processing devices 10.
[0012] The terminal device 20 is any device used by users such as sales staff involved in vehicle sales. For example, general-purpose electronic devices such as personal computers, smartphones, tablet devices, and wearable devices, or dedicated electronic devices, can be used as the terminal device 20. Although Figure 1 shows an example where System 1 has one terminal device 20, it is not limited to this. System 1 may have two or more terminal devices 20.
[0013] First, an overview of the audio data text conversion technology according to this embodiment will be described, and further details will be provided later. The audio data may be, for example, audio data from a business negotiation. In this embodiment, the business negotiation is, for example, a negotiation related to the sale of a vehicle, and the offering related to the negotiation is a vehicle, but is not limited to this. For example, the business negotiation may be a meeting aimed at concluding various types of contracts, such as the buying and selling of real estate, the contract of an insurance product, or the sale of a financial product. Furthermore, the offering related to the business negotiation in this embodiment may be goods, services, digital content, licenses, data / information, financial products, real estate, intangible assets, other tradable rights, etc.
[0014] The information processing device 10 acquires audio data. The information processing device 10 also detects specific expressions in the audio data and characteristic information related to the pronunciation of those specific expressions. Based on the detected specific expressions and characteristic information, the information processing device 10 converts the specific expressions into standard Japanese. Finally, the information processing device 10 outputs text information related to the audio data.
[0015] Thus, according to this embodiment, the information processing apparatus 10 detects a specific expression in the voice data and feature information related to the pronunciation of the specific expression, and converts the specific expression into a standard language based on the detected specific expression and feature information. Therefore, the text conversion technology of voice data is improved in that the specific expression can be appropriately extracted and converted into a standard language.
[0016] Next, each component of the information processing apparatus 10 and the terminal apparatus 20 will be described in detail.
[0017] (Configuration of Information Processing Apparatus 10) As shown in FIG. 1, the information processing apparatus 10 includes a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15.
[0018] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (central processing unit) or a GPU (graphics processing unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (field-programmable gate array) or an ASIC (application specific integrated circuit). The control unit 11 executes processing related to the operation of the information processing apparatus 10 while controlling each part of the information processing apparatus 10.
[0019] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a RAM (random access memory) or a ROM (read only memory). The RAM is, for example, a SRAM (static random access memory) or a DRAM (dynamic random access memory). The ROM is, for example, an EEPROM (electrically erasable programmable read only memory). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data used for the operation of the information processing apparatus 10 and data obtained by the operation of the information processing apparatus 10.
[0020] The input unit 13 includes at least one input interface. The input interface is, for example, a physical key, a capacitive key, a pointing device, or a touch screen provided integrally with a display. The input interface may also be, for example, a sound sensor that receives voice input, or a camera that receives gesture input. The input unit 13 receives an operation for inputting data used for the operation of the information processing apparatus 10. Instead of being provided in the information processing apparatus 10, the input unit 13 may be connected to the information processing apparatus 10 as an external input device. As the connection method, for example, any method such as USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), or Bluetooth (registered trademark) can be used.
[0021] The output unit 14 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as sound. The display is, for example, an LCD (liquid crystal display) or an organic EL (electroluminescence) display. The output unit 14 outputs data obtained by the operation of the information processing device 10. Instead of being provided in the information processing device 10, the output unit 14 may be connected to the information processing device 10 as an external output device. Any connection method can be used, for example, USB, HDMI (registered trademark), or Bluetooth (registered trademark).
[0022] The communication unit 15 includes at least one external communication interface. The communication interface may be either a wired or wireless communication interface. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface compatible with mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface compatible with short-range wireless communication such as Bluetooth (registered trademark). The communication unit 15 receives data used for the operation of the information processing device 10 and transmits data obtained by the operation of the information processing device 10.
[0023] The functions of the information processing device 10 are realized by executing the program according to this embodiment on a processor corresponding to the control unit 11. In other words, the functions of the information processing device 10 are realized by software. The program causes the computer to perform the operations of the information processing device 10, thereby causing the computer to function as the information processing device 10. That is, the computer functions as the information processing device 10 by performing the operations of the information processing device 10 according to the program.
[0024] In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-temporary computer-readable media, such as magnetic recording devices, optical discs, magneto-optical recording media, or semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending portable recording media such as DVDs (digital versatile discs) or CD-ROMs (compact disc read-only memory) on which the program is recorded. Alternatively, the program may be distributed by storing it on the storage of an external server and transmitting it from the external server to other computers. The program may also be provided as a program product.
[0025] Some or all of the functions of the information processing device 10 may be implemented by a dedicated circuit corresponding to the control unit 11. In other words, some or all of the functions of the information processing device 10 may be implemented by hardware.
[0026] (Configuration of terminal device 20) As shown in Figure 1, the terminal device 20 comprises a control unit 21, a storage unit 22, an input unit 23, an output unit 24, and a communication unit 25.
[0027] The control unit 21 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (central processing unit) or GPU (graphics processing unit), or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). The control unit 21 controls each part of the terminal device 20 and executes processes related to the operation of the terminal device 20.
[0028] The storage unit 22 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or at least two combinations thereof. The semiconductor memory is, for example, RAM (random access memory) or ROM (read-only memory). The RAM is, for example, SRAM (static random access memory) or DRAM (dynamic random access memory). The ROM is, for example, EEPROM (electrically erasable programmable read-only memory). The storage unit 22 functions, for example, as a main memory, auxiliary memory, or cache memory. The storage unit 22 stores data used for the operation of the terminal device 20 and data obtained by the operation of the terminal device 20.
[0029] The input unit 23 includes at least one input interface. The input interface may be, for example, a physical key, a capacitive key, a pointing device, or a touchscreen integrated with a display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 23 accepts operations to input data used for the operation of the terminal device 20. Instead of being provided in the terminal device 20, the input unit 23 may be connected to the terminal device 20 as an external input device. Any connection method can be used, for example, USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), or Bluetooth (registered trademark).
[0030] The output unit 24 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as sound. The display is, for example, an LCD (liquid crystal display) or an organic EL (electroluminescence) display. The output unit 24 outputs data obtained by the operation of the terminal device 20. Instead of being provided in the terminal device 20, the output unit 24 may be connected to the terminal device 20 as an external output device. Any connection method can be used, for example, USB, HDMI (registered trademark), or Bluetooth (registered trademark).
[0031] The communication unit 25 includes at least one external communication interface. The communication interface may be either a wired or wireless communication interface. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface compatible with mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface compatible with short-range wireless communication such as Bluetooth (registered trademark). The communication unit 25 receives data used for the operation of the terminal device 20 and transmits data obtained by the operation of the terminal device 20.
[0032] The functions of the terminal device 20 are realized by executing the program according to this embodiment on a processor corresponding to the control unit 21. In other words, the functions of the terminal device 20 are realized by software. The program causes the computer to perform the operations of the terminal device 20, thereby causing the computer to function as the terminal device 20. That is, the computer functions as the terminal device 20 by performing the operations of the terminal device 20 according to the program.
[0033] Some or all of the functions of the terminal device 20 may be implemented by a dedicated circuit corresponding to the control unit 21. In other words, some or all of the functions of the terminal device 20 may be implemented by hardware.
[0034] (Operation of the information processing device 10) Referring to Figure 2, the operation of the information processing device 10 according to this embodiment will be explained. Here, we will mainly explain an example in which the voice data is voice data related to a business negotiation, and the business negotiation is related to the sale of a vehicle.
[0035] Step S10: The control unit 11 of the information processing device 10 acquires audio data.
[0036] Any method can be used for acquiring audio data. For example, the control unit 11 may acquire audio data from external devices, including the terminal device 20, via the communication unit 15 and the network 30. Alternatively, the control unit 11 may acquire audio data via the input unit 13.
[0037] Step S20: The control unit 11 detects a specific expression in the audio data acquired in step S10, and characteristic information related to the pronunciation of that specific expression.
[0038] Specific expressions include dialects, non-standard languages, slang, and other specific expressions used in actual speech.
[0039] Any method can be used to detect specific expressions in audio data. For example, a method may be employed in which keywords, phrases, etc., that have been registered in advance as specific expressions are matched using a speech recognition engine. Alternatively, for example, a method may be employed in which speech patterns are classified using a machine learning model such as deep learning and a predetermined expression is identified.
[0040] Feature information related to the pronunciation of a specific expression includes, for example, the intonation, rhythm of pronunciation, pitch variation, and pronunciation speed. Such feature information often reflects the speaker's vocal habits, emotions, and intentions, and is considered important for capturing in detail how a specific expression was pronounced. Any method can be used to detect feature information related to the pronunciation of a specific expression. For example, a method that analyzes intonation variation using a pitch detection algorithm may be employed. Alternatively, for example, a method that extracts temporal and frequency features of the speech signal using spectrogram analysis may be employed. Furthermore, for example, feature information related to pronunciation may be detected by comprehensively capturing the characteristics of speech using machine learning models such as deep learning.
[0041] Step S30: The control unit 11 converts the specific expression into the corresponding standard language based on the detected specific expression and feature information.
[0042] Any method may be used for the conversion process to standard language. For example, the control unit 11 may perform a process to convert a specific expression to standard language based on the detected specific expression and feature information pair and a conversion rule between non-standard language and standard language. The conversion rule may be stored in, for example, the memory unit 12, and the control unit 11 may perform the above conversion process by referring to the conversion rule in the memory unit 12. In this case, the conversion rule may be a conversion table between the specific expression and feature information pair and the corresponding standard language. The corresponding standard language is the standard language to be converted for each specific expression. In other words, the corresponding standard language is the linguistic expression obtained as a result of the conversion process.
[0043] Step S40: The control unit 11 outputs text information related to the audio data.
[0044] The text information is information obtained by converting audio data into text, specifically information in which the aforementioned specific expressions within the audio data have been converted into standard Japanese. In other words, the output text information is a standard Japanese sentence obtained as a result of speech recognition, and is text after the detection of specific expressions and conversion to standard Japanese. For example, for audio data containing the Aichi dialect phrase "Kono michi wa tomenko," the control unit 11 uses a conversion rule based on the specific expression "tomenko" and its tone to output the text information "Kono michi wa tou-ko-shi" (This road is closed). Also, for example, for audio data containing the Australian English dialect phrase "Servo," the control unit 11 uses a conversion rule based on the specific expression "Servo" and its tone to output the text information "Service Station." Furthermore, sounds that are pronounced "a" in standard English are pronounced "ai" in Australian English (for example, "today," "say," and "face"). Even differences in pronunciation can be addressed by applying conversion rules based on specific expressions and tones, which convert those expressions to "today," "say," "face," etc., and output them accordingly. In this way, even if specific expressions are dialects or phrases used in different regions or cultures, they can be converted to standard Japanese using appropriate conversion rules, providing text information in a form that is understandable to a wider range of users.
[0045] Any method can be used for processing the output of text information. For example, the control unit 11 may transmit data to the terminal device 20 via the communication unit 15, and the terminal device 20 may output the judgment result via a user interface displayed by the output unit 24. Alternatively, the control unit 11 may output the judgment result via a user interface displayed by the output unit 14.
[0046] With this configuration, the information processing device 10 detects specific expressions in the audio data and characteristic information related to the pronunciation of those specific expressions, and converts the specific expressions into standard Japanese based on the detected specific expressions and characteristic information. Therefore, the audio data text conversion technology is improved in that it can appropriately extract specific expressions and convert them into standard Japanese.
[0047] While this disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art may make various modifications and alterations based on this disclosure. Therefore, it should be noted that these modifications and alterations are within the scope of this disclosure. For example, the functions, etc., included in each component or step can be rearranged in a logically consistent manner, and multiple components or steps can be combined into one or divided into two.
[0048] For example, the audio data may be audio from a business negotiation related to a predetermined offering, and the control unit 11 of the information processing device 10 may identify regional information corresponding to the speaker from the audio data. In this case, the control unit 11 may present suggestions related to the predetermined offering based on the regional information. For example, for audio data containing the specific expression "ketta," which is a dialect of Aichi Prefecture, the control unit 11 converts it to standard Japanese based on the expression and intonation of "ketta" and outputs the text information "bicycle." Here, the regional information of the Aichi region includes the information that it is common to load bicycles into cars. In this case, the control unit 11 may suggest vehicles with a large carrying capacity (e.g., minivans, SUVs, etc.) based on the regional information and the conversion process. Also, for example, for audio data containing the specific expression "barihaee," which is a dialect of Fukuoka Prefecture, the control unit 11 converts it to standard Japanese based on the expression and intonation of "hari" and outputs the text information "very fast." Here, the regional information of the Fukuoka region includes the information that expressways are frequently used. In this case, the control unit 11 may suggest vehicles with superior fuel efficiency, acceleration, and high-speed driving performance (e.g., hybrid cars, sports sedans, etc.) based on the regional information and the conversion process. By considering regional information in this way, it becomes possible to make more appropriate and convenient suggestions to the user, and it is expected that the efficiency of business negotiations will improve. Any method may be used to identify the regional information. For example, a method may be used in which a speech recognition engine is used to identify the speaker's regionality based on a list of specific expressions or dialects contained in the voice data.
[0049] Furthermore, in the embodiment described above, it is also possible to have an embodiment in which the configuration and operation of the information processing device 10 are distributed among multiple computers that can communicate with each other.
[0050] Some embodiments of the present disclosure are described below. However, it should be noted that the embodiments of the present disclosure are not limited to these. [Note 1] A method for converting audio data to text, performed by an information processing device, Acquiring audio data and, To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, Based on the detected specific expression and the characteristic information, the specific expression is converted into the corresponding standard language. To output text information related to the aforementioned audio data, A method for converting audio data containing text into text. [Note 2] The method for converting audio data to text as described in Appendix 1, A method for converting speech data to text, which converts the specific expression to the corresponding standard language based on the detected pair of specific expression and feature information and the conversion rules between non-standard language and standard language. [Note 3] A method for converting audio data to text as described in Appendix 1 or 2, The characteristic information related to the aforementioned vocalization is tone information, and the method is for converting speech data into text. [Note 4] A method for converting audio data to text as described in any one of the appendices 1 to 3, The aforementioned specific expression is a method for converting audio data to text, including dialects and non-standard languages. [Note 5] A method for converting audio data to text as described in any one of the appendices 1 to 4, wherein the audio data is audio from a business negotiation relating to a specified product, and further, From the aforementioned audio data, regional information corresponding to the speaker is identified. A method for converting audio data into text, which presents suggestions related to the predetermined offerings based on the aforementioned regional information. [Note 6] An information processing device comprising a control unit, The control unit, Acquire audio data, The system detects specific expressions in the aforementioned audio data and characteristic information related to the pronunciation of those specific expressions. Based on the detected specific expression and the characteristic information, the specific expression is converted to the corresponding standard language. An information processing device that outputs text information related to the aforementioned audio data. [Note 7] On the computer, Acquiring audio data and, To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, Based on the detected specific expression and the characteristic information, the specific expression is converted into the corresponding standard language. To output text information related to the aforementioned audio data, A program that executes the command. [Explanation of Symbols]
[0051] 10 Information Processing Devices 11 Control Unit 12 Storage section 13 Input section 14 Output section 15 Communications Department 20 Terminal devices 21 Control Unit 22 Memory section 23 Input section 24 Output section 25 Communications Department 30 Networks
Claims
1. A method for converting audio data to text, performed by an information processing device, Acquiring audio data and To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, Based on the detected specific expression and the characteristic information, the specific expression is converted into the corresponding standard language. To output text information related to the aforementioned audio data, A method for converting audio data containing text into text.
2. A method for converting audio data to text according to claim 1, A method for converting speech data to text, which converts the specific expression to the corresponding standard language based on the detected pair of specific expression and feature information and the conversion rules between non-standard language and standard language.
3. A method for converting audio data to text according to claim 1, The characteristic information related to the aforementioned vocalization is tone information, and the method is for converting speech data into text.
4. A method for converting audio data to text according to claim 1, The aforementioned specific expression is a method for converting audio data to text, including dialects and non-standard languages.
5. A method for converting audio data to text according to claim 1, wherein the audio data is audio from a business negotiation relating to a predetermined product, and further, From the aforementioned audio data, regional information corresponding to the speaker is identified. A method for converting audio data into text, which presents suggestions related to the predetermined offerings based on the aforementioned regional information.
6. An information processing device comprising a control unit, The control unit, Acquire audio data, The system detects specific expressions in the aforementioned audio data and characteristic information related to the pronunciation of those specific expressions. Based on the detected specific expression and the characteristic information, the specific expression is converted to the corresponding standard language. An information processing device that outputs text information related to the aforementioned audio data.
7. On the computer, Acquiring audio data and To detect specific expressions in the aforementioned audio data, and characteristic information related to the pronunciation of said specific expressions, Based on the detected specific expression and the characteristic information, the specific expression is converted into the corresponding standard language. To output text information related to the aforementioned audio data, A program that executes the command.