Text conversion method for voice data, information processing device, and program

By detecting and converting specific expressions and vocal features in sound data, the information processing device converts them into standard language, solving the transcription problem of non-standard languages ​​and improving the applicability and comprehensibility of text transformation.

CN122024733APending Publication Date: 2026-05-12TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2025-10-31
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, there is room for improvement in the text transformation of audio data, especially in the transcription of non-standard languages ​​such as dialects and accents, as it is difficult to effectively convert them into standard languages ​​for wider understanding.

Method used

The information processing device detects specific expressions and vocal characteristics in sound data, and uses transformation rules to convert them into standard language and output text information.

Benefits of technology

It enables the appropriate extraction and transformation of specific expressions into standard language, improving the text conversion technology of audio data and adapting to the understanding of users from different regions and cultures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024733A_ABST
    Figure CN122024733A_ABST
Patent Text Reader

Abstract

The invention provides a text conversion method for voice data, an information processing apparatus, and a program. The text conversion method of voice data executed by an information processing apparatus includes: acquiring voice data; detecting a specific expression in the sound data and feature information related to the sound production of the specific expression; according to the detected specific expression and the feature information, converting the specific expression into a corresponding standard language; and outputting text information related to the sound data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to methods for text transformation of sound data, information processing apparatus, and programs. Background Technology

[0002] Techniques for analyzing the content of business conversations are known. For example, Japanese Patent Application Publication 2019-28910 discloses a dialogue analysis system that, during business conversations with customers, checks whether sales personnel explain what should be explained and avoid stating what should not be stated. Additionally, for example, in Horimoto et al., "Transformation of Standard Japanese Based on Tomiyama Ben's Deep Learning Voice Recognition," The 38th paper... th At the Annual Conference of the Japanese Society for Artificial Intelligence (2024), the sound recognition technology for the Toyama dialect was unveiled. Summary of the Invention

[0003] Japanese Patent Application Publication No. 2019-28910 demonstrates a technique for analyzing the content of business conversations through machine learning. However, Japanese Patent Application Publication No. 2019-28910, as well as Hori Gen et al., "Transformation of Toyama Ben's Voice Recognition Based on Deep Learning" (The 38th issue), also presents a different technique. th The Annual Conference of the Japanese Society for Artificial Intelligence (2024) did not mention text-to-text conversion technology for audio data, such as audio used in business conversations. In particular, there is room for improvement in audio transcription technology for audio data containing non-standard languages, including dialects and accents. On the other hand, for the analysis and feedback of content in business conversations, it is preferable to improve text-to-text conversion technology for audio data. Thus, there is room for improvement in text-to-text conversion technology for audio data used in business conversations.

[0004] This disclosure provides a text transformation technique for audio data.

[0005] The first aspect of this disclosure relates to a text conversion method for sound data executed by an information processing apparatus, comprising: acquiring sound data; detecting a specific expression in the sound data and feature information related to the pronunciation of the specific expression; converting the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and outputting text information related to the sound data.

[0006] The information processing apparatus involved in the second aspect of this disclosure includes one or more processors, which are configured to: acquire sound data; detect a specific expression in the sound data and feature information related to the pronunciation of the specific expression; transform the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and output text information related to the sound data.

[0007] The third method of this disclosure involves a program that causes one or more processors to perform the following functions: acquiring sound data; detecting a specific expression in the sound data and feature information related to the pronunciation of the specific expression; transforming the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and outputting text information related to the sound data.

[0008] According to one embodiment of this disclosure, a text conversion technique for improving audio data is provided. Attached Figure Description

[0009] The features, advantages, and technical and industrial significance of the preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings, in which the same symbols represent the same parts, and wherein:

[0010] Figure 1 This is a block diagram illustrating the general structure of the system involved in this embodiment.

[0011] Figure 2 This is a flowchart illustrating the operation of an information processing device. Detailed Implementation

[0012] The following describes the implementation of this disclosure.

[0013] (Summary of the implementation method)

[0014] Reference Figure 1 This section describes the overview and structure of System 1 according to this embodiment. System 1 according to this embodiment includes an information processing device 10 and a terminal device 20. The information processing device 10 and the terminal device 20 are communicatively connected to, for example, a network 30 including a mobile communication network and the Internet.

[0015] Information processing device 10 is, for example, a server device installed in a data center. For instance, information processing device 10 is a server belonging to a cloud computing system or other computing system. Furthermore, in Figure 1 The example shown is of a system 1 having one information processing device 10, but it is not limited to this. System 1 may also have two or more information processing devices 10.

[0016] Terminal device 20 is any device used by users such as business personnel involved in vehicle sales transactions. For example, general-purpose electronic devices such as personal computers, smartphones, tablets, and wearable devices, or specialized electronic devices, can be used as terminal device 20. Furthermore, in Figure 1 The example shown is of a single terminal device 20 in System 1, but it is not limited to this. System 1 may also have two or more terminal devices 20.

[0017] First, an overview of the text conversion technology for the audio data involved in this embodiment will be provided, with details to follow later. Furthermore, the audio data can be, for example, audio data from a business conversation. In this embodiment, the business conversation is, for example, a business conversation related to vehicle sales, and the offering related to the business conversation is a vehicle, but it is not limited to this. For example, the business conversation could also be a meeting aimed at concluding various types of contracts, such as the sale of real estate, insurance contracts, or the sale of financial products. Additionally, the offering related to the business conversation in this embodiment can be goods, services, digital content, licenses, data / information, financial products, real estate, intangible assets, or other tradable rights.

[0018] The information processing device 10 acquires sound data. Furthermore, the information processing device 10 detects specific expressions in the sound data and feature information related to the pronunciation of those specific expressions. Based on the detected specific expressions and feature information, the information processing device 10 converts the specific expressions into standard language. Finally, the information processing device 10 outputs text information related to the sound data.

[0019] Thus, according to this embodiment, the information processing device 10 detects specific expressions in sound data and feature information related to the pronunciation of those specific expressions, and transforms the specific expressions into standard language based on the detected specific expressions and feature information. Therefore, the text conversion technology for sound data is improved in terms of being able to appropriately extract specific expressions and transform them into standard language.

[0020] Next, the structures of the information processing device 10 and the terminal device 20 will be described in detail.

[0021] (Structure of information processing device 10)

[0022] like Figure 1 As shown, the information processing device 10 includes a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15.

[0023] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specifically designed for particular processing. The dedicated circuit is, for example, a FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). The control unit 11 controls the various parts of the information processing device 10 while performing processing related to the operation of the information processing device 10.

[0024] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of them. The semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory). RAM is, for example, SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory). ROM is, for example, EEPROM (Electrically Erasable Programmable Read Only Memory). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data used in the operation of the information processing device 10 and data obtained through the operation of the information processing device 10.

[0025] The input unit 13 includes at least one input interface. The input interface may be, for example, a physical key, a capacitive key, a pointing device, or a touchscreen integrated with the display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts the input of data used in the operation of the information processing device 10. The input unit 13 may also be connected to the information processing device 10 as an external input device, instead of being installed on the information processing device 10. As a connection method, for example, any method such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface), or Bluetooth (Bluetooth) can be used.

[0026] The output unit 14 includes at least one output interface. The output interface may be, for example, a display that outputs image information or a speaker that outputs sound information. The display may be, for example, an LCD (Liquid Crystal Display) or an OLED (Electro Luminescence) display. The output unit 14 outputs data obtained through the operation of the information processing device 10. Alternatively, the output unit 14 may be connected to the information processing device 10 as an external output device, instead of being installed on the information processing device 10. As a connection method, for example, any method such as USB, HDMI, or Bluetooth can be used.

[0027] The communication unit 15 includes at least one external communication interface. This communication interface can be any interface used in wired or wireless communication. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface corresponding to mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface corresponding to short-range wireless communication such as Bluetooth. The communication unit 15 receives data used in the operation of the information processing device 10 and transmits data obtained through the operation of the information processing device 10.

[0028] The functions of the information processing device 10 are implemented by a processor, equivalent to the control unit 11, executing the program described in this embodiment. That is, the functions of the information processing device 10 are implemented through software. The program causes the computer to perform the actions of the information processing device 10, thereby enabling the computer to function as the information processing device 10. In other words, the computer functions as the information processing device 10 by executing the actions of the information processing device 10 according to the program.

[0029] In this embodiment, the program can be pre-recorded on a computer-readable recording medium. Computer-readable recording media include non-transitory computer-readable media, such as magnetic recording devices, optical discs, optical-magnetic recording media, or semiconductor memory. Program distribution can be achieved, for example, through the sale, transfer, or rental of removable recording media such as DVDs (Digital Versatile Discs) or CD-ROMs (Compact Disc Read Only Memory) containing the program. Alternatively, program distribution can also be achieved by storing the program on the storage device of an external server and sending the program from the external server to other computers. Furthermore, the program can also be provided as a program product.

[0030] Some or all of the functions of the information processing device 10 can also be implemented by a dedicated circuit equivalent to the control unit 11. That is, some or all of the functions of the information processing device 10 can also be implemented by hardware.

[0031] (Structure of terminal device 20)

[0032] like Figure 1 As shown, the terminal device 20 includes a control unit 21, a storage unit 22, an input unit 23, an output unit 24, and a communication unit 25.

[0033] The control unit 21 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specifically designed for particular processing. The dedicated circuit is, for example, a FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). While controlling various parts of the terminal device 20, the control unit 21 performs processing related to the operation of the terminal device 20.

[0034] The storage unit 22 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of them. The semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory). RAM is, for example, SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory). ROM is, for example, EEPROM (Electrically Erasable Programmable Read Only Memory). The storage unit 22 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 22 stores data used during the operation of the terminal device 20 and data obtained through the operation of the terminal device 20.

[0035] The input unit 23 includes at least one input interface. The input interface may be, for example, a physical key, a capacitive key, a pointing device, or a touchscreen integrated with the display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 23 accepts the input of data used in the operation of the terminal device 20. The input unit 23 may also be connected to the terminal device 20 as an external input device, instead of being installed on the terminal device 20. As a connection method, for example, any method such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface), or Bluetooth (Bluetooth) can be used.

[0036] The output unit 24 includes at least one output interface. The output interface may be, for example, a display for outputting image information or a speaker for outputting sound information. The display may be, for example, an LCD (Liquid Crystal Display) or an OLED (Electro Luminescence) display. The output unit 24 outputs data obtained through the operation of the terminal device 20. Alternatively, the output unit 24 may be connected to the terminal device 20 as an external output device, instead of being installed on the terminal device 20. As a connection method, for example, any method such as USB, HDMI, or Bluetooth can be used.

[0037] The communication unit 25 includes at least one external communication interface. This communication interface can be any interface used in wired or wireless communication. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface corresponding to mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface corresponding to short-range wireless communication such as Bluetooth. The communication unit 25 receives data used in the operation of the terminal device 20 and transmits data obtained through the operation of the terminal device 20.

[0038] The functions of the terminal device 20 are implemented by a processor, equivalent to the control unit 21, executing the program described in this embodiment. That is, the functions of the terminal device 20 are implemented through software. The program causes the computer to perform the actions of the terminal device 20, thereby enabling the computer to function as the terminal device 20. In other words, the computer functions as the terminal device 20 by executing the actions of the terminal device 20 according to the program.

[0039] Some or all of the functions of the terminal device 20 can also be implemented by a dedicated circuit equivalent to the control unit 21. That is, some or all of the functions of the terminal device 20 can also be implemented by hardware.

[0040] (Operation of information processing device 10)

[0041] Reference Figure 2 This section explains the operation of the information processing apparatus 10 according to this embodiment. Here, we will mainly describe an example where the voice data is related to business conversation, and the business conversation is related to vehicle sales.

[0042] Step S10: The control unit 11 of the information processing device 10 acquires the sound data.

[0043] In the acquisition and processing of audio data, any method can be used. For example, the control unit 11 can also acquire audio data from external devices, including the terminal device 20, via the communication unit 15 and the network 30. Alternatively, the control unit 11 can also acquire audio data via the input unit 13.

[0044] Step S20: The control unit 11 detects a specific expression in the sound data obtained in step S10 and feature information related to the vocalization of that specific expression.

[0045] Specific expressions include dialects, slang, and other specific expressions used in actual speech.

[0046] In the detection and processing of specific expressions in sound data, any method can be employed. For example, a method can be used to compare the sound recognition engine with keywords, phrases, etc., that are pre-registered as specific expressions. Alternatively, a method can be used, for example, to classify sound patterns and identify predetermined expressions using machine learning models such as deep learning.

[0047] Feature information related to the vocalization of a specific expression includes, for example, tone, rhythm, pitch variation, and vocal speed. These features often reflect the speaker's vocal habits, emotions, and intentions, and are considered important for accurately capturing how a specific expression is pronounced. Any method can be used in the detection and processing of feature information related to the vocalization of a specific expression. For example, pitch detection algorithms can be used to analyze tone variations. Alternatively, methods such as spectrogram analysis can be used to extract time / frequency features of the sound signal. Furthermore, machine learning models such as deep learning can be utilized to capture overall sound features and detect vocalization-related feature information.

[0048] Step S30: The control unit 11 transforms the specific expression into the corresponding standard language based on the detected specific expression and feature information.

[0049] In the transformation process to standard language, any method can be used. For example, the control unit 11 can perform a process of transforming a specific expression into standard language based on the pairing of detected specific expressions and feature information, and the transformation rules between non-standard language and standard language. The transformation rules can be stored in the storage unit 12, for example, and the control unit 11 can perform the above transformation process by referring to the transformation rules in the storage unit 12. Furthermore, in this case, the transformation rules can be a transformation table of pairing specific expressions and feature information with corresponding standard languages. The corresponding standard language is the standard language that should be transformed for each specific expression. In other words, the corresponding standard language is the language expression obtained as a result of the transformation process.

[0050] Step S40: The control unit 11 outputs text information related to the sound data.

[0051] The text information is information obtained by transforming voice data into text and transforming the above-mentioned specific expressions in the voice data into standard language. In other words, the output text information is an article in standard language obtained as a result of voice recognition, and is text after detection of specific expressions and transformation processing into standard language. For example, for voice data including the Aichi dialect such as "この道はとーめんこ", through a specific expression such as "とーめんこ" and intonation, using transformation rules, the control unit 11 outputs text information "この道は通行止め". "とーめんこ" refers to the Aichi dialect meaning "通行止め". Additionally, for example, for voice data including the Australian English dialect such as "Servo(サーヴォ)", through a specific expression such as "Servo" and intonation, using transformation rules, the control unit 11 outputs text information "Service Station(ガソリンスタンド)". "Servo" refers to the Australian English dialect meaning "Service Station". Furthermore, in standard English, the letter "a" pronounced as "eɪ" becomes pronounced as "aɪ" in Australian English. For example, the word "today" is pronounced as "tʊdaɪ". The word "say" is pronounced as "saɪ". The word "face" is pronounced as "faɪs". Such pronunciation differences are also output by applying transformation rules based on specific expressions and intonation to transform specific expressions into "today", "say", "face", etc. Thus, even if specific expressions are dialects, locutions, etc. used in different regions and cultures, by transforming them into standard language using appropriate transformation rules, text information is provided in a form that can be understood by a wider range of users.

[0052] In the output processing of text information, any method can be adopted. For example, the control unit 11 can send data to the terminal device 20 via the communication unit 15 and use the output unit 24 of the terminal device 20 to display the output user interface to output the determination result. Or, the control unit 11 can use the output unit 14 to display the output user interface to output the determination result.

[0053] According to the above structure, the information processing device 10 detects specific expressions in voice data and characteristic information related to the pronunciation of the specific expressions, and transforms the specific expressions into standard language according to the detected specific expressions and characteristic information. Therefore, in terms of being able to appropriately extract and transform specific expressions into standard language, the text transformation technology of voice data is improved.

[0054] The present disclosure has been described with reference to the accompanying drawings and embodiments, but it should be noted that those skilled in the art can also make various modifications and changes based on the present disclosure. Therefore, it is desired to note that these modifications and changes are included in the scope of the present disclosure. For example, the functions, etc. included in each structural part or each step, etc. can be reconfigured logically without contradiction, and multiple structural parts or steps, etc. can be combined into one or divided.

[0055] For example, the voice data can be the voice in a commercial conversation related to a predetermined offering, and the control unit 11 of the information processing apparatus 10 can also determine the regional information corresponding to the speaker based on the voice data. In this case, the control unit 11 can also present a recommendation related to the predetermined offering based on the regional information. For example, for voice data including a specific expression such as "ケッタ" which is a dialect of Aichi Prefecture, the control unit 11 converts it into standard language according to the expression such as "ケッタ" and the tone, and outputs the text information "自転車". "ケッタ" means the dialect of Aichi Prefecture that means "自転車". Here, the regional information of the Aichi region includes information such as generally loading a bicycle into the vehicle. In this case, the control unit 11 can also recommend a vehicle with a large loading capacity (examples: minivan, SUV, etc.) based on this regional information and this conversion process. Additionally, for example, for voice data including a specific expression such as "ばりはえー" which is a dialect of Fukuoka Prefecture, the control unit 11 converts it into standard language according to the expression such as "ばり" and the tone, and outputs the text information "とても速い". "ばり" means the dialect of Fukuoka Prefecture that means "とても". "はえー" means a non-standard language that means "速い". Here, the regional information of the Fukuoka region includes information such as a high frequency of using the highway. In this case, the control unit 11 can also recommend a vehicle with excellent fuel consumption rate, vehicle acceleration, and high-speed driving performance (examples: hybrid vehicle, sports sedan, etc.) based on this regional information and this conversion process. Thus, by considering regional information, it is possible to make more appropriate and convenient recommendations for users, and it is expected to improve the efficiency of commercial conversations. In addition, in the process of determining regional information, any method can be adopted. For example, a method of using a voice recognition engine to identify the region of the speaker based on a list of specific expressions, dialects, etc. included in the voice data can be adopted.

[0056] Additionally, for example, in the above-described embodiment, an embodiment in which the structure and operation of the information processing apparatus 10 are distributed to multiple computers capable of communicating with each other can also be implemented.

[0057] Hereinafter, a part of the embodiments of the present disclosure will be exemplified. However, it is desired to note that the embodiments of the present disclosure are not limited to these.

[0058] [Supplementary Note 1]

[0059] A text transformation method is a method for transforming sound data executed by an information processing device. The text transformation method includes: acquiring sound data; detecting a specific expression in the sound data and feature information related to the pronunciation of the specific expression; transforming the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and outputting text information related to the sound data.

[0060] [Postscript 2]

[0061] According to the text transformation method for the sound data described in Appendix 1, the method further includes transforming the specific expression into the corresponding standard language based on the pairing of the detected specific expression and the feature information, and the transformation rules between non-standard language and standard language.

[0062] [Postscript 3]

[0063] According to the text transformation method of the sound data recorded in Appendix 1 or 2, the feature information related to the vocalization is tone information.

[0064] [Postscript 4]

[0065] The text transformation method of the sound data is described in any one of the appendices 1 to 3, wherein the specific expression includes dialects and slang.

[0066] [Postscript 5]

[0067] A text transformation method for the audio data described in any one of Appendices 1 to 4, wherein the audio data is audio from a business conversation related to a predetermined offer, and the text transformation method includes: determining, based on the audio data, geographic information corresponding to the speaker; and presenting, based on the geographic information, suggestions related to the predetermined offer.

[0068] [Postscript 6]

[0069] An information processing apparatus includes one or more processors configured to perform: acquiring sound data; detecting a specific expression in the sound data and feature information related to the pronunciation of the specific expression; converting the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and outputting text information related to the sound data.

[0070] [Postscript 7]

[0071] A program that causes one or more processors to perform the following functions, wherein the functions include: acquiring sound data; detecting a specific expression in the sound data and feature information related to the pronunciation of the specific expression; converting the specific expression into a corresponding standard language based on the detected specific expression and the feature information; and outputting text information related to the sound data.

Claims

1. A text transformation method, which is a method for transforming audio data into text performed by an information processing device, characterized in that, The text transformation method includes: Obtain sound data; Detect specific expressions in the sound data and feature information related to the pronunciation of those specific expressions; Based on the detected specific expression and the feature information, the specific expression is transformed into the corresponding standard language; and Output text information related to the sound data.

2. The text transformation method according to claim 1, characterized in that, Also includes: Based on the detected specific expression and the pairing of the feature information, and the transformation rules between non-standard and standard languages, the specific expression is transformed into the corresponding standard language.

3. The text transformation method according to claim 1, characterized in that, The characteristic information related to the vocalization is tone information.

4. The text transformation method according to claim 1, characterized in that, The specific expressions mentioned include dialects and slang.

5. The text transformation method according to claim 1, characterized in that, The audio data is audio from business conversations related to the predetermined offering, and, The text transformation method includes: Based on the sound data, determine the geographical information corresponding to the speaker; and Based on the aforementioned regional information, suggestions related to the predetermined offerings are presented.

6. An information processing device, characterized in that, It includes one or more processors, wherein the one or more processors are configured as follows: Obtain sound data; Detect specific expressions in the sound data and feature information related to the pronunciation of those specific expressions; Based on the detected specific expression and the feature information, the specific expression is transformed into the corresponding standard language; and Output text information related to the sound data.

7. A program that causes one or more processors to perform the following functions, characterized in that, The functions include: Obtain sound data; Detect specific expressions in the sound data and feature information related to the pronunciation of those specific expressions; Based on the detected specific expression and the feature information, the specific expression is transformed into the corresponding standard language; and Output text information related to the sound data.