Text conversion method for voice data, information processing device, and program
By using an information processing device and an error correction table to correct error patterns in business meeting voice data, the shortcomings of existing voice data-to-text conversion technologies are solved, achieving more efficient text conversion and more accurate analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-26
AI Technical Summary
In the existing technology, there is room for improvement in the technology of converting voice data to text in business conversations, especially in the analysis and feedback of business conversation content, where the technology of converting voice data to text has not been sufficiently improved.
Voice data is acquired through an information processing device, text data is generated, and an error correction table is used to correct error patterns in the text data. The error correction table includes a first error pattern, at least one second error pattern, and a corresponding correction pattern. The error correction table generates the second error pattern through various processing methods to improve error correction efficiency.
It improves the error correction capabilities of voice data to text conversion technology, enhancing the accuracy of business meeting content analysis and feedback.
Smart Images

Figure CN122090848A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to methods, information processing apparatus, and programs for text conversion of speech data. Background Technology
[0002] Techniques for analyzing the content of business conversations are known. For example, Japanese Patent Application Publication No. 2019-28910 discloses a dialogue analysis system that examines whether, in a business conversation with a customer, the sales manager has stated what should be stated and what should not be stated. Summary of the Invention
[0003] Japanese Patent Application Publication 2019-28910 illustrates a technique for analyzing the content of business meetings using machine learning, but it does not mention transcription of speech from business meetings, i.e., text-to-speech conversion technology. On the other hand, for the analysis and feedback of content from business meetings, it is preferable to improve text-to-speech conversion technology for speech data. Thus, there is room for improvement in text-to-speech conversion technology for speech data from business meetings.
[0004] This disclosure provides a method, information processing apparatus, and program for converting speech data into text, which improves the text conversion technology of speech data.
[0005] The first embodiment of the speech data text conversion method disclosed herein is performed by an information processing device. The speech data text conversion method includes: acquiring speech data; generating text data based on the speech data; and correcting error patterns contained in the text data based on an error correction table. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern, wherein the at least one second error pattern is generated based on speech data that generated the first error pattern, and the correction pattern corresponds to the first error pattern and the at least one second error pattern.
[0006] Based on the text conversion method for speech data in the first aspect of this disclosure, the at least one second error pattern may also be generated based on object data in the speech data that generated the first error pattern, wherein the object data is data within a predetermined time range that includes the period during which the speech corresponding to the first error pattern is emitted.
[0007] Based on the text conversion method for speech data of the first aspect of this disclosure, the at least one second error pattern can also be generated by inputting the processed data after processing the object data into the speech recognition engine.
[0008] Based on the text conversion method for speech data of the first aspect of this disclosure, the processed data may also be data obtained by processing at least one of noise addition processing, noise removal processing, frequency change processing, and volume change processing on the object data.
[0009] Based on the text conversion method for speech data of the first aspect of this disclosure, the at least one second error pattern can also be generated by inputting the object data into a speech recognition engine with adjusted parameters.
[0010] The information processing apparatus of the second aspect of this disclosure includes a control unit. The control unit is configured to acquire voice data; generate text data based on the voice data; and correct error patterns contained in the text data based on an error correction table. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern, wherein the at least one second error pattern is generated based on voice data that generated the first error pattern, and the correction pattern corresponds to the first error pattern and the at least one second error pattern.
[0011] The procedure of the third aspect of this disclosure causes a computer to perform the following actions: acquire speech data; generate text data based on the speech data; and correct error patterns contained in the text data based on an error correction table. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern, wherein the at least one second error pattern is generated based on speech data that generated the first error pattern, and the correction pattern corresponds to the first error pattern and the at least one second error pattern.
[0012] According to one embodiment of this disclosure, the text conversion technology for voice data is improved. Attached Figure Description
[0013] The features, advantages, and technical and industrial significance of exemplary embodiments of the present invention will now be described with reference to the accompanying drawings, in which the same reference numerals denote the same elements, and wherein:
[0014] Figure 1 This is a block diagram illustrating the general structure of the system according to this embodiment.
[0015] Figure 2 It is a flowchart illustrating the operation of an information processing device.
[0016] Figure 3 It is a flowchart illustrating the operation of an information processing device. Detailed Implementation
[0017] The embodiments of this disclosure will now be described.
[0018] Summary of Implementation Methods
[0019] Reference Figure 1 This section describes the overview and structure of System 1 according to this embodiment. System 1 of this embodiment includes an information processing device 10 and a terminal device 20. The information processing device 10 and the terminal device 20 are connected to a network 30, such as a mobile communication network and the Internet, in a communicative manner.
[0020] Information processing device 10 is, for example, a server device installed in a data center. For example, information processing device 10 is a server belonging to a cloud computing system or other computing system. It should be noted that... Figure 1 The example shown is of a single information processing device 10 in system 1, but it is not limited to this. System 1 may also have two or more information processing devices 10.
[0021] Terminal device 20 is any device used by the user. For example, general-purpose electronic devices or dedicated electronic devices such as personal computers, smartphones, tablets, and wearable devices can be used as terminal device 20. It should be noted that... Figure 1 The example shown is of a single terminal device 20 in system 1, but it is not limited to this. System 1 may also have two or more terminal devices 20.
[0022] First, an overview of the text conversion technology for voice data in this embodiment will be given, with details to follow later. Here, voice data can be data from a specific domain. For example, voice data can be voice data from a business meeting. In this embodiment, the business meeting is, for example, a business meeting involving the sale of a vehicle, where the offered item is a vehicle, but it is not limited to this. For example, a business meeting could also be a meeting for the purpose of signing various contracts, such as the sale of real estate, insurance contracts, or the sale of financial products. Furthermore, the offered item in the business meeting in this embodiment can be goods, services, digital content, licenses, data / information, financial products, real estate, intangible assets, or other tradable rights.
[0023] The information processing device 10 acquires voice data. Furthermore, the information processing device 10 generates text data based on the voice data. Additionally, the information processing device 10 corrects patterns (hereinafter also referred to as error patterns) such as miswritten sentences and misidentified phrases contained in the text data based on an error correction table.
[0024] Here, the error correction table includes a certain error pattern (hereinafter also referred to as the first error pattern), at least one other error pattern generated based on the speech data that generated the first error pattern (hereinafter also referred to as the second error pattern), and corrected phrases corresponding to the first error pattern and at least one second error pattern (hereinafter also referred to as corrected patterns).
[0025] Thus, according to this embodiment, the information processing apparatus 10 corrects error patterns contained in text data based on an error correction table. Specifically, the error correction table includes a first error pattern, at least one second error pattern generated based on speech data from which the first error pattern was generated, and correction patterns corresponding to the first error pattern and the at least one second error pattern. Because the error correction table contains multiple error patterns, the possibility of correcting erroneous text data is increased, and because multiple error patterns are efficiently generated based on speech data, the speech data-to-text conversion technology is improved.
[0026] Next, the structures of the information processing device 10 and the terminal device 20 will be described in detail.
[0027] Structure of Information Processing Device 10
[0028] like Figure 1 As shown, the information processing device 10 includes a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15.
[0029] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (central processing unit) or a GPU (graphics processing unit), or a dedicated processor for specific processing. The dedicated circuit is, for example, a FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). The control unit 11 controls the various parts of the information processing device 10 while performing processing related to the operation of the information processing device 10.
[0030] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or at least a combination of two of these. The semiconductor memory is, for example, RAM (random access memory) or ROM (read-only memory). RAM is, for example, SRAM (static random access memory) or DRAM (dynamic random access memory). ROM is, for example, EEPROM (electrically erasable programmable read-only memory). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data for the operation of the information processing device 10 and data obtained through the operation of the information processing device 10. Specifically, for example, a speech recognition engine is stored in the storage unit 12. The speech recognition engine has the function of converting speech input into text data, and analyzes the user's speech to generate corresponding text information. Additionally, for example, an error correction table is stored in the storage unit 12. An error correction table is used to convert error patterns in text recognized by a speech recognition engine into corrected patterns. For example, the car model name RAV4 (registered trademark) as an error pattern can be recognized by the speech recognition engine as LOVE4. The error correction table may contain information that establishes a correspondence between the error pattern LOVE4 and the corrected pattern RAV4 (registered trademark). By referring to the error correction table, proper nouns such as car model names and function names can be correctly corrected.
[0031] The input unit 13 includes at least one input interface. The input interface may be, for example, a physical button, a capacitive button, an indicator device, or a touchscreen integrated with the display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input or a camera that accepts gesture input. The input unit 13 accepts data input for the operation of the information processing device 10. The input unit 13 may also be connected to the information processing device 10 as an external input device, instead of being installed on the information processing device 10. As a connection method, any method such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface), or Bluetooth (registered trademark) can be used.
[0032] The output unit 14 includes at least one output interface. The output interface may be, for example, a display that outputs image information or a speaker that outputs sound information. The display may be, for example, an LCD (liquid crystal display) or an organic EL (electroluminescence) display. The output unit 14 outputs data obtained through the operation of the information processing device 10. Alternatively, the output unit 14 may be provided externally to the information processing device 10 instead of being installed thereon. As a connection method, any method such as USB, HDMI (registered trademark), or Bluetooth (registered trademark) can be used.
[0033] The communication unit 15 includes at least one external communication interface. This communication interface can be either a wired or wireless communication interface. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface corresponding to mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface corresponding to short-range wireless communication such as Bluetooth. The communication unit 15 receives data for the operation of the information processing device 10 and also transmits data obtained through the operation of the information processing device 10.
[0034] The functions of the information processing device 10 are implemented by the processor, which corresponds to the control unit 11, executing the program of this embodiment. That is, the functions of the information processing device 10 are implemented by software. The program causes the computer to perform the actions of the information processing device 10, thereby enabling the computer to function as the information processing device 10. In other words, the computer functions as the information processing device 10 by executing the actions of the information processing device 10 according to the program.
[0035] In this embodiment, the program can be recorded on a computer-readable recording medium. Computer-readable recording media include non-transitory computer-readable media, such as magnetic recording devices, optical discs, optical-magnetic recording media, or semiconductor memory. Program distribution can be achieved, for example, by selling, transferring, or lending removable recording media such as DVDs (digital versatile discs) or CD-ROMs (compact disc read-only memory) containing the program. Alternatively, program distribution can also be achieved by storing the program in the storage of an external server and sending the program from the external server to other computers. Furthermore, the program can also be provided as a program product.
[0036] Some or all of the functions of the information processing device 10 may also be implemented by a dedicated circuit equivalent to the control unit 11. That is, some or all of the functions of the information processing device 10 may also be implemented by hardware.
[0037] Structure of terminal device 20
[0038] like Figure 1 As shown, the terminal device 20 includes a control unit 21, a storage unit 22, an input unit 23, an output unit 24, and a communication unit 25.
[0039] The control unit 21 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a CPU (central processing unit) or a GPU (graphics processing unit), or a dedicated processor for specific processing. The dedicated circuit is, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). The control unit 21 controls the various parts of the terminal device 20 while performing processing related to the operation of the terminal device 20.
[0040] The storage unit 22 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or at least two combinations thereof. The semiconductor memory is, for example, RAM (random access memory) or ROM (read-only memory). RAM is, for example, SRAM (static random access memory) or DRAM (dynamic random access memory). ROM is, for example, EEPROM (electrically erasable programmable read-only memory). The storage unit 22 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 22 stores data for the operation of the terminal device 20 and data obtained through the operation of the terminal device 20.
[0041] The input unit 23 includes at least one input interface. The input interface may be, for example, a physical button, a capacitive button, an indicator device, or a touchscreen integrated with the display. Alternatively, the input interface may be, for example, a sound sensor that accepts voice input or a camera that accepts gesture input. The input unit 23 accepts data input for the operation of the terminal device 20. The input unit 23 may also be connected to the terminal device 20 as an external input device, instead of being installed on the terminal device 20. As a connection method, any method such as USB (Universal Serial Bus), HDMI (High-Definition Multimedia Interface), or Bluetooth (registered trademark) can be used.
[0042] The output unit 24 includes at least one output interface. The output interface may be, for example, a display that outputs image information or a speaker that outputs sound information. The display may be, for example, an LCD (liquid crystal display) or an OLED (electroluminescence) display. The output unit 24 outputs data obtained through the operation of the terminal device 20. The output unit 24 may also be connected to the terminal device 20 as an external output device, replacing the device itself. As a connection method, any method such as USB, HDMI, or Bluetooth can be used.
[0043] The communication unit 25 includes at least one external communication interface. This communication interface can be either a wired or wireless communication interface. In the case of wired communication, the communication interface may be, for example, a LAN (Local Area Network) interface or a USB (Universal Serial Bus) interface. In the case of wireless communication, the communication interface may be, for example, an interface corresponding to mobile communication standards such as LTE (Long Term Evolution), 4G (4th generation), or 5G (5th generation), or an interface corresponding to short-range wireless communication such as Bluetooth. The communication unit 25 receives data for the operation of the terminal device 20 and also transmits data obtained through the operation of the terminal device 20.
[0044] The functions of the terminal device 20 are implemented by the processor, which corresponds to the control unit 21, executing the program of this embodiment. That is, the functions of the terminal device 20 are implemented by software. The program causes the computer to perform the actions of the terminal device 20, thereby enabling the computer to function as the terminal device 20. In other words, the computer functions as the terminal device 20 by executing the actions of the terminal device 20 according to the program.
[0045] Some or all of the functions of the terminal device 20 can also be implemented by a dedicated circuit equivalent to the control unit 21. That is, some or all of the functions of the terminal device 20 can also be implemented by hardware.
[0046] Operation of information processing device 10
[0047] Reference Figure 2 The operation of the information processing apparatus 10 in this embodiment will be explained. Here, we will mainly describe an example where the voice data is voice data from a business meeting involved in vehicle sales.
[0048] S10: The control unit 11 of the information processing device 10 acquires voice data.
[0049] The acquisition and processing of voice data can be performed using any method. For example, the control unit 11 can also acquire voice data from an external device, such as the terminal device 20, via the communication unit 15 and the network 30. Alternatively, the control unit 11 can also acquire voice data via the input unit 13.
[0050] S20: The control unit 11 generates text data based on the voice data acquired in step S10.
[0051] The generation and processing of text data based on speech data can employ any method. For example, the control unit 11 can generate text data corresponding to the speech data by inputting speech data into the speech recognition engine.
[0052] S30: The control unit 11 corrects the error patterns contained in the text data based on the error correction table. For example, the control unit 11 extracts all patterns of sentences, phrases, etc. in the text data that match the error patterns in the error correction table. Then, the control unit 11 changes the extracted patterns to the corrected patterns based on the corrected patterns corresponding to the error patterns in the error correction table.
[0053] S40: Control unit 11 outputs the corrected text data.
[0054] The text information output processing can employ any method. For example, the control unit 11 can send data to the terminal device 20 via the communication unit 15 and output text data through the user interface displayed on the output unit 24 of the terminal device 20. Alternatively, the control unit 11 can output text data through the user interface displayed on the output unit 14.
[0055] Here, the error correction table includes a first error pattern, at least one second error pattern generated based on the speech data that generated the first error pattern, and a correction pattern corresponding to the first error pattern and the at least one second error pattern. (Refer to...) Figure 3 An example of the operation involved in generating the error correction table of the information processing apparatus 10 of this embodiment will be described.
[0056] S110: The control unit 11 of the information processing device 10 acquires voice data that corresponds to the voice message of the first error mode. Specifically, for example, let's assume the first error mode is LOVE4. In this case, the voice data corresponding to the first error mode is the voice message that says the car model name RAV4 (registered trademark). The control unit 11 can, for example, determine the voice data that corresponds to the first error mode by referring to an error correction table, and acquire the determined voice data.
[0057] The acquisition and processing of voice data in S110 can be performed using any method. For example, the control unit 11 can also acquire voice data from an external device including the terminal device 20 via the communication unit 15 and the network 30. Alternatively, the control unit 11 can also acquire voice data via the input unit 13.
[0058] S120: The control unit 11 generates at least one second error pattern based on data (hereinafter also referred to as object data) within a predetermined time range in the voice data that includes the period during which the voice corresponding to the first error pattern is emitted. The predetermined time may be, for example, 10 seconds. That is, the object data may be voice data that includes the period of 10 seconds before the period during which the voice corresponding to the first error pattern is emitted and the period after the period during which the voice corresponding to the first error pattern is emitted. By performing the processing described later, which also includes the period before and after the period during which the voice corresponding to the first error pattern is emitted, context and background information can also be obtained. For example, if the first error pattern is LOVE4, the second error pattern may be LAB4, LAVE4, RAB4, REV4, RAF4, RAP4, LAB4, etc.
[0059] The generation of at least one second error pattern can be performed using any method. For example, at least one second error pattern can be generated by inputting processed data (after processing the object data) into the speech recognition engine. The processed data can be data processed by at least one of the following: noise addition processing, noise removal processing, frequency alteration processing, and volume alteration processing of the object data. In this way, the control unit 11 can generate at least one second error pattern by inputting processed data (after processing the object data) into the speech recognition engine. By using the processed data, the variation of the error pattern can be increased. Therefore, the possibility of correcting erroneous text data using the error correction table is increased.
[0060] Alternatively, for example, at least one second error pattern can be generated by inputting the object data into a parameter-adjusted speech recognition engine. Parameters may include, for example, recognition accuracy, confidence thresholds, language model parameters, acoustic model parameters, and parameters related to a custom dictionary. In this way, the control unit 11 can input the object data into the parameter-adjusted speech recognition engine without processing it, thereby generating at least one second error pattern. Adjusting the parameters of the speech recognition engine can also increase the variation of the error pattern. Therefore, the possibility of correcting erroneous text data using an error correction table is increased.
[0061] S130: The control unit 11 establishes a correspondence between at least one second error mode generated in S120 and the correction mode and stores it in the error correction table. For example, as described above, if the second error mode is LAB4, LAVE4, RAB4, REV4, RAF4, RAP4, LAB4, etc., these error modes are established to correspond with the correction mode RAV4 (registered trademark) and stored in the error correction table.
[0062] According to this structure, the information processing device 10 corrects error patterns contained in text data based on an error correction table. Specifically, the error correction table includes a first error pattern, at least one second error pattern generated based on the speech data from which the first error pattern was generated, and correction patterns corresponding to the first error pattern and the at least one second error pattern. Thus, since the error correction table includes multiple different error patterns, the possibility of correcting erroneous text data is increased, and since multiple error patterns are efficiently generated based on speech data, the speech data-to-text conversion technology is improved.
[0063] This disclosure has been described based on the accompanying drawings and embodiments, but those skilled in the art should note that various modifications and alterations can be made based on this disclosure. Therefore, it should be understood that such modifications and alterations are included within the scope of this disclosure. For example, the functions contained in each component or step can be reconfigured in a logically consistent manner, and multiple components or steps can be combined into one or divided.
[0064] Alternatively, for example, in the above embodiment, it is also possible to distribute the structure and operation of the information processing device 10 to multiple computers that can communicate with each other.
[0065] The following is an example of one embodiment of the present disclosure. However, it should be noted that the embodiments of the present disclosure are not limited thereto.
[0066] Postscript 1
[0067] A method for converting speech data to text, performed by an information processing device, wherein the method for converting speech data to text includes:
[0068] Acquire voice data;
[0069] Based on the voice data, text data is generated; and
[0070] Based on the error correction table, correct the error patterns contained in the text data.
[0071] The error correction table includes a first error pattern, at least one second error pattern generated based on the speech data that generated the first error pattern, and a correction pattern corresponding to the first error pattern and the at least one second error pattern.
[0072] Appendix 2
[0073] According to the text conversion method for speech data described in Appendix 1, wherein,
[0074] The at least one second error pattern is generated based on object data within a predetermined time range that includes the period during which the speech corresponding to the first error pattern is emitted from the speech data that generated the first error pattern.
[0075] Appendix 3
[0076] According to the text conversion method for speech data described in Appendix 2, wherein,
[0077] The at least one second error pattern is generated by inputting processed data, which is the result of processing the object data, into the speech recognition engine.
[0078] Appendix 4
[0079] According to the text conversion method for speech data described in Appendix 3, wherein...
[0080] The processed data is obtained by processing the object data through at least one of the following processes: noise addition processing, noise removal processing, frequency change processing, and volume change processing.
[0081] Appendix 5
[0082] According to the text conversion method for speech data described in Appendix 2, wherein,
[0083] The at least one second error pattern is generated by inputting the object data into a parameter-adjusted speech recognition engine.
[0084] Appendix 6
[0085] An information processing device includes a control unit, wherein...
[0086] The control unit acquires voice data, generates text data based on the voice data, and corrects error patterns contained in the text data based on an error correction table.
[0087] The error correction table includes a first error pattern, at least one second error pattern generated based on the speech data that generated the first error pattern, and a correction pattern corresponding to the first error pattern and the at least one second error pattern.
[0088] Appendix 7
[0089] A program that causes a computer to perform the following actions:
[0090] Acquire voice data;
[0091] Based on the voice data, text data is generated; and
[0092] Based on the error correction table, the error patterns contained in the text data are corrected, wherein...
[0093] The error correction table includes a first error pattern, at least one second error pattern generated based on the speech data that generated the first error pattern, and a correction pattern corresponding to the first error pattern and the at least one second error pattern.
Claims
1. A method for converting speech data to text, executed by an information processing device, characterized in that the method comprises: Acquire voice data; Based on the voice data, generate text data; as well as Based on an error correction table, the error patterns contained in the text data are corrected. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern. The at least one second error pattern is generated based on the speech data that generated the first error pattern. The correction pattern corresponds to the first error pattern and the at least one second error pattern.
2. The method for converting speech data to text according to claim 1, characterized in that, The at least one second error pattern is generated based on object data in the speech data that generated the first error pattern, the object data being data within a predetermined time range that includes the period during which the speech corresponding to the first error pattern is emitted.
3. The text conversion method according to claim 2, characterized in that, The at least one second error pattern is generated by inputting processed data, which is the result of processing the object data, into the speech recognition engine.
4. The text conversion method according to claim 3, characterized in that, The processed data is obtained by processing the object data through at least one of the following processes: noise addition processing, noise removal processing, frequency change processing, and volume change processing.
5. The text conversion method according to claim 2, characterized in that, The at least one second error pattern is generated by inputting the object data into a parameter-adjusted speech recognition engine.
6. An information processing device, characterized in that, The control unit is configured as follows: Acquire voice data; Based on the voice data, text data is generated; and Based on an error correction table, the error patterns contained in the text data are corrected. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern. The at least one second error pattern is generated based on the speech data that generated the first error pattern. The correction pattern corresponds to the first error pattern and the at least one second error pattern.
7. A program in which, This program causes the computer to perform the following actions: Acquire voice data; Based on the voice data, text data is generated; and Based on the error correction table, the error patterns contained in the text data are corrected. The error correction table includes a first error pattern, at least one second error pattern, and a correction pattern. The at least one second error pattern is generated based on the speech data that generated the first error pattern. The correction pattern corresponds to the first error pattern and the at least one second error pattern.