Substitution device, substitution method, and substitution program

The system addresses the challenge of replacing speech recognition text without modifying the model by using an acquisition, selection, and replacement unit to identify and replace similar text patterns, thereby enhancing efficiency and reducing costs.

JP2025073889APending Publication Date: 2025-05-13NIPPON TELEGRAPH & TELEPHONE CORP +1

Patent Information

Application Number
JP2023185037
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing speech recognition systems struggle to replace text related to speech recognition results without modifying the underlying model, as this often requires adding text to the language model, which can be costly and complex.

Method used

A system comprising an acquisition unit for obtaining character strings related to speech recognition results and pre-registered text, a selection unit for identifying similar character strings to determine a replacement target range, and a replacement unit that replaces the target range with the pre-registered text, allowing for text replacement without altering the model.

Benefits of technology

Enables the replacement of text related to speech recognition results without touching the model, simplifying the process and reducing costs associated with model modifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025073889000001_ABST
    Figure 2025073889000001_ABST
Patent Text Reader

Abstract

To replace texts related to speech recognition results without modifying a model.SOLUTION: A replacement device 10 comprises: a string acquisition unit 111 for acquiring a first string related to a first text that is a speech recognition result and a second string related to a second text that has been registered in advance; a selection unit 112 for selecting a range to be replaced from the first text based on similarity between at least a substring of the first string and the second string; and a replacement unit 113 that replaces the range to be replaced with the second text.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a replacement device, a replacement method, and a replacement program. [Background technology]

[0002] Techniques for outputting speech recognition results as text include, for example, an acoustic model + language model method that outputs speech recognition results as text using an acoustic model and a language model, etc. Techniques for adding text to replace text related to speech recognition results to a model such as a language model include, for example, a technique for estimating the occurrence probability of a word to be added and reflecting it in a language model operated by speech recognition processing (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2012-242421 A Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the conventional technology, it may not be possible to replace text related to the speech recognition result without modifying the model. For example, in the acoustic model + language model method, in order to replace text related to the speech recognition result, text for replacing the text related to the speech recognition result must be added to the language model, and this addition may incur costs for the work of reflecting the text in the language model. For this reason, in the acoustic model + language model method, it may not be possible to replace text related to the speech recognition result without modifying the model. [Means for solving the problem]

[0005] In order to solve the above-mentioned problems, the present invention includes an acquisition unit that acquires a first character string related to a first text that is a speech recognition result and a second character string related to a second text that is registered in advance, a selection unit that selects a replacement target range from the first text based on a similarity between at least a partial character string of the first character string and the second character string, and a replacement unit that replaces the replacement target range with the second text. Effect of the Invention

[0006] The present invention allows for text replacement for speech recognition results without modifying the model. [Brief description of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of an overview of a replacement system according to a first embodiment. [Diagram 2] FIG. 2 is a diagram illustrating an example of the configuration of the replacement device according to the first embodiment. [Diagram 3] FIG. 3 is a flowchart showing an example of the flow of processing executed by the substitution system according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of a replacement device according to the second embodiment. [Diagram 5] FIG. 5 is a diagram for explaining an example of selection. [Figure 6] FIG. 6 is a diagram for explaining an example of selection. [Figure 7] FIG. 7 is a diagram for explaining an example of selection. [Figure 8] FIG. 8 is a diagram for explaining an example of selection. [Figure 9] FIG. 9 is a diagram illustrating an example of the configuration of a replacement device according to the third embodiment. [Figure 10] FIG. 10 is a diagram for explaining an example of replacement. [Figure 11] FIG. 11 is a diagram illustrating an example of an overview of a replacement system according to the fourth embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the configuration of a replacement device according to the fourth embodiment. [Figure 13] FIG. 13 is a flowchart showing an example of the flow of processing executed by the substitution system according to the fourth embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of an overview of a replacement system according to the fifth embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of the configuration of a replacement device according to the fifth embodiment. [Figure 16] FIG. 16 is a diagram for explaining an example of the conversion. [Figure 17] FIG. 17 is a flowchart showing an example of the flow of a process executed by the substitution system according to the fifth embodiment. [Figure 18] FIG. 18 is a diagram illustrating an example of the configuration of a replacement device according to the sixth embodiment. [Figure 19] FIG. 19 is a flowchart showing an example of the flow of a process executed by the substitution system according to the sixth embodiment. [Figure 20] FIG. 20 is a diagram illustrating an example of an overview of a replacement system according to the seventh embodiment. [Figure 21] FIG. 21 is a diagram illustrating an example of the configuration of a replacement device according to the seventh embodiment. [Figure 22] FIG. 22 is a flowchart showing an example of the flow of a process executed by the substitution system according to the seventh embodiment. [Diagram 23] FIG. 23 is a diagram illustrating an example of the configuration of a computer that executes a replacement program. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] Each embodiment of the present invention will be described below with reference to the drawings, but the present invention is not limited to each of the following embodiments. If it is clear to a person skilled in the art to combine, modify, or improve each of the following embodiments, such an embodiment may also be included in the technical scope of the present invention. In addition, in the description of the drawings, the same parts are given the same reference numerals, and duplicated descriptions are omitted, and descriptions of members having similar functions and similar processes are also omitted.

[0009] [0.Reference technology] Before describing the present embodiment, an example of a reference technique related to speech recognition will be described. Techniques for outputting a speech recognition result as text include, for example, the above-mentioned acoustic model+language model method and an End-to-End method for outputting a speech recognition result as text by End-to-End.

[0010] However, in the reference technology, there are cases where it is not possible to replace text related to the speech recognition result without modifying the model. For example, in the End-to-End method, one neural network model is trained from a pair of input speech and output text, so that a word dictionary, which is a database in which words such as service names and product names are registered, cannot be used in the speech recognition process. In addition, in the End-to-End method, when a word is added to the model, additional learning must be performed, for example, by transcribing the speech, which is costly and tuning such as adding a word to the model cannot be easily performed.

[0011] The present embodiment has been made to solve the above-mentioned problems, and makes it possible to replace text related to the speech recognition result without modifying the model.

[0012] [1. First embodiment] [Example of summary] An example of an overview of a substitution system 1 according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram for explaining an example of an overview of the substitution system according to the first embodiment.

[0013] As illustrated in FIG. 1, a character string acquisition unit (acquisition unit) 111 of the replacement device 10 acquires a “first character string” related to a “first text” that is a speech recognition result, and a “second character string” related to a “second text” that is registered in advance.

[0014] Here, the "first text" is, for example, text data of the speech recognition result output from a model (not shown). Also, the "first character string" refers to, for example, a character string expressing the "pronunciation" or "reading" of the text of the speech recognition result output from the model. Here, the "pronunciation" refers, for example, to a phoneme string expressing the pronunciation of the text. Also, the "reading" refers, for example, to a character string expressing the reading of the text in katakana, roman letters, phonetic symbols, or the like. The above-mentioned model may be any model, for example, it may be a model of the acoustic model + language model method described above, or it may be a model of the end-to-end method.

[0015] The "second text" is, for example, text data of a word preregistered in a word dictionary. The "second character string" is, for example, a character string representing the "pronunciation" or "reading" of a word preregistered in a word dictionary.

[0016] Next, the selection unit 112 of the replacement device 10 selects a "replacement target range" from the first text based on the "similarity" between at least a part of a substring of the first string and the second string. Here, "similarity" refers to, for example, the degree of closeness between the substring and the second string. The "replacement target range" refers to, for example, the range of the first text to be replaced with the second text.

[0017] Then, the replacing unit 113 of the replacement device 10 replaces the replacement target range selected by the selection unit 112 with the second text. As a result, the replacing unit 113 outputs the first text in which the replacement target range has been replaced with the second text.

[0018] In this way, the replacement device 10 can replace text related to the speech recognition result without modifying the model. In other words, the replacement device 10 selects a replacement target range, which is a target range to be replaced, from the first text output from the model, and replaces the replacement target range with the second text. Therefore, for example, if a word dictionary that lists the words (second text) to be added and the corresponding second character strings is prepared, it is possible to replace the first text, which is the speech recognition result, with the added words (second text) without modifying the model itself.

[0019] [Example of replacement device configuration] 2 is a diagram showing an example of the configuration of a replacement device according to the first embodiment. The replacement device 10 is realized by, for example, a device related to the fields of speech recognition processing and natural language processing. In the example shown in FIG. 2, the replacement device 10 includes a control unit 11, an input unit 12, an output unit 13, and a storage unit 14.

[0020] [Control Unit] The control unit 11 controls the entire replacement device 10. For example, the control unit 11 is configured with one or more processors having programs that define the flow of various processes and internal memory that stores control information, and the processor executes each process using the programs and internal memory. The control unit 11 is realized by, for example, electronic circuits such as a central processing unit (CPU), a micro processing unit (MPU), and a graphics processing unit (GPU), or integrated circuits such as an application specific integrated circuit (ASIC) and a field programmable gate array (FPGA). In the example shown in FIG. 2, the control unit 11 has a character string acquisition unit 111, a selection unit 112, and a replacement unit 113.

[0021] [String acquisition part] The character string acquiring unit 111 acquires a first character string related to a first text, which is a speech recognition result, and a second character string related to a second text, which is registered in advance. For example, the character string acquiring unit 111 acquires, as the first character string, a phoneme string (e.g., "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU") or a reading (e.g., "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU") corresponding to the text of the speech recognition result output from the model (e.g., the text "My workplace is Yotsuhama's automated consultation Senda") or a reading (e.g., "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU"). Note that the phoneme string corresponding to the first text, which is a speech recognition result, is hereinafter referred to as "first phoneme string" as appropriate.

[0022] Also, for example, the character string acquiring unit 111 acquires, as the second character string, a phoneme string (e.g., "jidousoudaNseNtaa") or a reading (e.g., "jidousoudaNsenta") corresponding to a second text (e.g., text of the word "automatic consultation center") added to a word dictionary stored in the word dictionary storage unit 141 described later. Note that the phoneme string corresponding to the second text is hereinafter appropriately referred to as a "second phoneme string."

[0023] [Selection Department] The selection unit 112 selects a replacement target range from the first text based on the similarity between at least a partial string of the first string and the second string. For example, the selection unit 112 selects a replacement target range based on whether the partial string and the second string are acoustically similar or whether the reading of the first text corresponding to the partial string and the reading of the second text corresponding to the second string are similar. In the following, the reading of the first text corresponding to the partial string will be appropriately referred to as the "first reading". In the following, the reading of the second text corresponding to the second string will be appropriately referred to as the "second reading".

[0024] Here, a specific example of the processing of the selection unit 112 will be described using an example in which the character string acquisition unit 111 acquires a first phoneme string "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU" as a first character string and a second phoneme string "jidousoudaNseNtaa" as a second character string. In such a case, the selection unit 112 calculates the similarity between a partial character string "jidoosoodaNseNta" of the first phoneme string "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU" and the second phoneme string "jidousoudaNseNtaa", and if the similarity between the two is higher than a predetermined threshold, selects "Automatic Consultation Senda" from the first text as a replacement target range.

[0025] A specific example of the processing of the selection unit 112 will also be described for a case where the character string acquisition unit 111 acquires the reading of the first text "watashi no shokuba yon kohama no jido osoo dan ta desu" as the first character string and the second reading "jidousou dan center" as the second character string. In such a case, the selection unit 112 determines the similarity between the first reading "jidousou dan center", which is the reading of a partial character string of the reading of the first text "watashi no shokuba yon kohama no jido osoo dan ta desu", and the second reading "jidousou dan center", and if the similarity between the two is higher than a predetermined threshold, selects "jidousou dan center" from the first text as the replacement target range.

[0026] [Replacement part] The replacement unit 113 replaces the replacement target range with the second text. For example, when the selection unit 112 selects "Automatic Consultation Senda" as the replacement target range, the replacement unit 113 replaces the "Automatic Consultation Senda" with the second text "Child Consultation Center."

[0027] [Input / Output] The input unit 12 is realized by using an input device such as a keyboard or a mouse, and inputs various instruction information in response to an input operation by an operator. The output unit 13 is realized by using an output device such as a liquid crystal display. For example, the output unit 13 displays the first text after replacement in which the word "Automatic Consultation Senda" in the first text is replaced with the word "Child Consultation Center" in the second text.

[0028] [Storage] The storage unit 14 stores an OS (Operating System) and various programs executed by the replacement device 10. The storage unit 14 is realized by, for example, a storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or an optical disk, or a semiconductor memory in which information can be rewritten, such as a RAM (Random Access Memory), a flash memory, or a NVSRAM (Non Volatile Static Random Access Memory).

[0029] The storage unit 14 also includes a word dictionary storage unit 141. The word dictionary storage unit 141 stores, as a word dictionary, a second text and a second character string corresponding to the second text in association with each other. For example, the word dictionary storage unit 141 stores a word, which is the second text, in association with the reading or phoneme string of the word, or the reading and phoneme string. To explain this by taking a specific example, the word dictionary storage unit 141 stores, in association with each other, the word "Children's Consultation Center", which is the second text, and the reading of the word, "Jidousoudan Center", and the phoneme string "jidousoudaNseNtaa". Note that the word dictionary stored in the word dictionary storage unit 141 can be updated as appropriate, and can be added, edited, deleted, and the like.

[0030] [Example of processing flow] An example of the flow of a process executed by the substitution device 10 of the substitution system 1 according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the flow of a process executed by the substitution system according to the first embodiment.

[0031] In step S11, the character string acquiring unit 111 acquires a first character string and a second character string. For example, the character string acquiring unit 111 acquires a first character string output from a language model and a second character string added to a word dictionary. For example, when speech recognition is performed by an acoustic model and a language model, the character string acquiring unit 111 acquires a first character string related to a first text output from the language model and a second character string related to a second text added to a word dictionary.

[0032] In step S12, the selection unit 112 selects a replacement target range from the first text based on the similarity between at least a partial character string of the first character string and the second character string. For example, the selection unit 112 selects the replacement target range based on whether the partial character string and the second character string are acoustically similar or whether the first reading and the second reading are similar.

[0033] In step S13, the replacement unit 113 replaces the replacement target range with the second text. For example, when the selection unit 112 selects "Automatic Consultation Senda" as the replacement target range, the replacement unit 113 replaces the "Automatic Consultation Senda" with the second text "Child Consultation Center."

[0034] [Advantages of the First Embodiment] In this way, the replacement device 10 can replace text related to the speech recognition result without modifying the model. In other words, the replacement device 10 selects a replacement target range, which is a target range to be replaced, from the first text output from the model, and replaces the replacement target range with the second text. Therefore, for example, if a word dictionary that lists the words (second text) to be added and the corresponding second character strings is prepared, it is possible to replace the first text, which is the speech recognition result, with the added words (second text) without modifying the model itself.

[0035] [2. Second embodiment] In the replacement device 10α according to the second embodiment, a character string representing the sound when the first text is spoken and a character string representing the sound when the second text is spoken are acquired, and a replacement target range is selected based on the length of the substring and the length of the second string as well as the similarity, and the replacement target range is replaced with the second text. Note that the description of the configuration and processing similar to those of the first embodiment is omitted.

[0036] In the following description of the second embodiment, a character string representing a sound when a first text is spoken refers to, for example, a phoneme string related to the pronunciation of the first text. Also, a character string representing a sound when a second text is spoken refers to, for example, a phoneme string related to the pronunciation of the second text. Also, a phoneme string is composed of one or more phonemes (phonological units that are the smallest units of speech), and refers to, for example, the pronunciation of a text expressed in Roman letters or the like.

[0037] As described above, in the second embodiment, the first character string is a character string representing the sound when the first text is spoken, and the second character string is a character string representing the sound when the second text is spoken, so that vowels and consonants of pronunciation are more easily distinguished than readings, and therefore the replacement device 10α can select the replacement target range in units finer than readings. In addition, the replacement device 10α can select the replacement target range based on the length of the substring and the length of the second string as well as the similarity, so that the replacement target range can be selected with high accuracy compared to selection of the replacement target range not based on the length of the substring and the length of the second string.

[0038] [Example of replacement device configuration] Fig. 4 is a diagram showing an example of the configuration of a replacement device according to the second embodiment. In the example shown in Fig. 4, the replacement device 10α has a control unit 11α instead of the control unit 11 in the first embodiment. The control unit 11α has a character string acquisition unit 111α and a selection unit 112α instead of the character string acquisition unit 111 and the selection unit 112 in the first embodiment.

[0039] [String acquisition part] The character string acquiring unit 111α acquires a character string representing a sound when the first text is spoken and a character string representing a sound when the second text is spoken. For example, the character string acquiring unit 111α acquires a first character string which is a first phoneme string output from a language model, and a second character string which is a second phoneme string added to a word dictionary.

[0040] As an example, the character string acquiring unit 111α acquires a first character string which is a first phoneme string "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU" corresponding to a first text related to a speech recognition result such as a transcribed text of a speech recognition result of "My workplace is Senda at Yotsuhama's automated consultation center." In this case, the character string acquiring unit 111α acquires a first character string corresponding to a first text divided into morphemes, in a state where a word such as "I" is associated with a first phoneme string such as "watashi."

[0041] Also, the character string acquiring unit 111α acquires, as an example, a second character string that is a second phoneme string "jidousoudaNseNtaa" corresponding to a pre-registered second text such as a text including the word "child consultation center" registered in a word dictionary. In this case, the character string acquiring unit 111α acquires the second character string that is a second phoneme string corresponding to the second text by, for example, reading a word dictionary in which a pair of the second text and the second character string is pre-registered.

[0042] In the above example, the first text and the first string, and the second text and the second string are separate, but the first string and the second string are not particularly limited as long as they are each composed of one or more characters; for example, the first text and the first string may be the same, and the second text and the second string may be the same.

[0043] [Selection Department] The selection unit 112α selects the replacement target range based on the length of the substring and the length of the second string as well as the similarity. For example, the selection unit 112α selects the replacement target range based on whether the length of the substring and the length of the second string are close to each other. As an example, the selection unit 112α compares the length of the substring and the length of the second string, and if these lengths are close to each other to a predetermined degree or more, determines that these lengths are close to each other and selects the range of the substring or the range of the first text corresponding to the substring as the replacement target range, and does not select the replacement target range if they are not close to each other. The selection unit 112α selects the replacement target range based on the length of the substring and the length of the second string, so that the replacement target range can be selected with high accuracy compared to selection of the replacement target range not based on the length of the substring and the length of the second string. Note that the "length of the substring" is not limited to the length of the substring itself, but includes, for example, the length of a string edited based on the substring.

[0044] Also, for example, the selection unit 112α may calculate the edit distance between the partial string and the second string as the similarity, and select the replacement target range based on the edit distance. Here, the edit distance refers to, for example, a distance indicating how similar and different the above-mentioned strings are. As an example, the selection unit 112α selects the replacement target range when the edit distance between the partial string and the second string is smaller than a preset value according to the length of the second string, and does not select it otherwise. As a result, the selection unit 112α can narrow down the replacement target range to a range of partial strings whose lengths are close to each other or a range of the first text corresponding to the range of the partial string, etc., thereby reducing the amount of calculation for selecting the replacement target range, and the replacement target range can be selected with higher accuracy than when the replacement target range is simply selected based on the lengths of the strings. Note that the "edit distance between the partial string and the second string" is not limited to the edit distance between the partial string itself and the second string, but includes, for example, the edit distance between the first string including the partial string and the second string. Moreover, the "range of a partial character string" here is not limited to the range of the partial character string itself, but includes, for example, the range of a character string obtained by editing the partial character string.

[0045] Also, for example, the selection unit 112α may select the replacement target range based on a corrected edit distance obtained by subtracting the difference between the length of the partial string and the length of the second string from the edit distance. As an example, the selection unit 112α performs DP (Dynamic Programming) matching between the first string and the second string to obtain a partial string based on the first string that matches the second string, and selects the range of the partial string or the range of the first text corresponding to the partial string as the replacement target range when the corrected edit distance obtained by subtracting the difference between the length of the partial string and the length of the second string from the edit distance is smaller than a preset value. Note that the "corrected edit distance" is not limited to the corrected edit distance obtained by subtracting the difference between the length of the partial string itself and the length of the second string from the edit distance, and includes, for example, the corrected edit distance obtained by subtracting the difference between the length of the first string including the partial string and the length of the second string from the edit distance.

[0046] Here, in DP matching, the greater the difference between the length of the substring and the length of the second string, the greater the edit distance. In response to this, the selection unit 112α selects the replacement target range based on the corrected edit distance in which the above-mentioned difference is subtracted from the edit distance, thereby narrowing down the replacement target range to a range of substrings whose lengths are close to each other or a range of the first text corresponding to the substring, even after performing DP matching once. Therefore, the selection unit 112α can reduce the amount of calculation required to select the replacement target range compared to when the replacement target range is selected based on an edit distance in which the above-mentioned difference is not subtracted from the edit distance, and can select the replacement target range with high accuracy.

[0047] For example, the selection unit 112α may select the replacement target range based on the ratio or magnitude relationship between the length of the substring and the length of the second string. As an example, when the ratio of the length of the second string to the length of the substring is greater than a preset value for determining whether the replacement target range should be selected, the selection unit 112α selects the range of the substring or the range of the first text corresponding to the substring as the replacement target range. As another example, when the length of the substring is a preset value according to the length of the second string and is greater than a preset value for determining whether the replacement target range should be selected, the selection unit 112α selects the range of the substring or the range of the first text corresponding to the substring as the replacement target range.

[0048] As a result, the selection unit 112α can select a range with fewer mismatched characters (phonemes, etc.) as the replacement range compared to when the replacement range is not selected based on the ratio or magnitude relationship between the length of the substring and the length of the second string, and therefore can select the replacement range with higher accuracy. In addition, the selection unit 112α can correct errors in speech recognition using the second text that was not learned during model learning, for example, by setting stricter preset conditions for determining whether or not to select the replacement range. As a result, the selection unit 112α can increase the matching rate between the first text when the speech is correctly transcribed and the first text after replacement compared to when the preset conditions for determining whether or not to select the replacement range are not strict.

[0049] For example, the selection unit 112α may select the replacement target range based on the edit distance between the substring and the second string, and may select the replacement target range based on the ratio or the magnitude relationship between the length of the substring and the length of the second string. As an example, the selection unit 112α performs DP matching to obtain a substring based on a first string that matches the second string, and performs preprocessing to provisionally select the range of the substring or the range of the first text corresponding to the substring as the replacement target range if the edit distance obtained by subtracting the difference between the length of the substring and the length of the second string is smaller than a preset value. Then, the selection unit 112α examines whether the ratio between the length of the substring and the length of the second string is larger than a preset value for determining whether the replacement target range should be selected, and performs postprocessing to determine whether the replacement target range provisionally selected by the preprocessing is to be selected if the ratio is larger than the preset value.

[0050] By performing the above-mentioned two-stage selection, the selection unit 112α can, for example, roughly select the replacement target range through pre-processing, and then precisely inspect the replacement target range through post-processing. This reduces replacement errors by the replacement unit 113 and enables the replacement target range to be selected with higher accuracy than when the two-stage selection is not performed.

[0051] [Example of selection by the selection department] An example of the selection by the selection unit 112α will be described with reference to FIG. 5 to FIG. 8. FIG. 5 to FIG. 8 are diagrams for explaining an example of the selection. In the example shown in FIG. 5, the selection unit 112α selects the replacement target range from the first text "My workplace is the automatic consultation Senda in Yotsuhama" regarding the speech recognition result corresponding to the speech "My workplace is the child consultation center in Yokohama". Also, in the example shown in FIG. 5, the selection unit 112α selects the replacement target range from the first text "My workplace is the automatic consultation Senda in Yotsuhama" and the first character string "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU" which is the first phoneme string in a state in which each word of the first text is associated with each phoneme string.

[0052] In the example shown in FIG. 5, the selection unit 112α selects the second character string (second phoneme string) from the second character strings (second phoneme strings) preregistered in the word dictionary in descending order of length. Next, the selection unit 112α calculates a corrected edit distance by, for example, subtracting the difference between the length of the first character string and the length of the selected second character string from the edit distance between the first character string and the second character string. Next, the selection unit 112α performs a preprocessing to provisionally select the range of the partial character string and the range of the first text corresponding to the partial character string as the replacement target range when the corrected edit distance is smaller than a preset value. Then, in the example shown in FIG. 5, the selection unit 112α examines whether the ratio between the length of the partial character string and the length of the second character string is larger than a preset value for determining whether the replacement target range should be selected, and performs a postprocessing to determine whether the replacement target range provisionally selected by the preprocessing is to be selected when the ratio is larger than the preset value.

[0053] [Example of pre-processing] First, the selection unit 112α performs DP matching between, for example, a first character string, "watashinoshokubawayoNkohamanojidoosoodaNseNtadesU," and a second character string, "jidousoudaNseNtaa." Next, the selection unit 112α obtains, for example, the position M of the character (phoneme, etc.) of the first character string that first matches the second character string ("hit" in FIG. 6) and the position N of the last matching character (phoneme, etc.). Next, the selection unit 112α obtains, for example, a character string that includes these matching characters (phonemes, etc.) from the first character string as a substring, and provisionally selects, as a replacement target range, a range of a character string that includes a substring from the character of the first character string that first matches the second character string to the last matching character, and a range of the first text that corresponds to the range of the substring, and obtains a character string that corresponds to the provisionally selected replacement target range as a substring. Here, since the replacement range is 0 when M=N, the selection unit 112α provisionally selects the replacement range when M≠N, for example.

[0054] Next, the selection unit 112α calculates a corrected edit distance by subtracting the difference between the length of the first string and the length of the second string from the edit distance between the first string and the second string, for example. As an example, the selection unit 112α calculates a corrected edit distance by subtracting the difference between the length of the first string=49 and the length of the second string=17 from the edit distance before correction (the total value of the number of "+" symbols corresponding to characters that are not present in the partial string and present in the second string in FIG. 6 and the number of "-" symbols corresponding to characters that are present in the partial string and not present in the second string)=38, as 38-(49-17)=6.

[0055] Next, the selection unit 112α calculates, for example, a preset value to be compared with the edit distance after the correction in terms of magnitude. As an example, the selection unit 112α calculates a preset value to be compared with the edit distance after the correction in terms of magnitude based on the length of the second phoneme string and a value for adjusting the length of the second phoneme string. The selection unit 112α can arbitrarily set the value for adjusting the length of the second phoneme string, but for example, when this value is set to a large value, the length of the second phoneme string can be adjusted to a large value, so that even a phoneme string with a large edit distance can be selected as a replacement target range. For example, when this value is set to a small value, the selection unit 112α can adjust the length of the second phoneme string to a small value, so that only a phoneme string with a small edit distance can be selected as a replacement target range. As an example, the selection unit 112α sets the value Y for adjusting the length of the second phoneme string "jidousoudaNseNtaa" corresponding to the word "child consultation center" in the second text to 0.5. Furthermore, the selection unit 112α calculates a preset value to be compared with the corrected edit distance as the length of the second phoneme string "jidousoudaNseNtaa"*Y=17*0.5=8 (rounded down to the nearest whole number), for example.

[0056] In the example shown in Fig. 7, the selection unit 112α determines that the preset condition for determining whether to select the replacement target range is satisfied because the corrected edit distance = 6 is smaller than the preset value = 8 that is compared with the corrected edit distance, as described above, and provisionally selects the range of the partial string "jidoousooudaNseNta" in which two characters "u" corresponding to "+" in Fig. 6 that are not in the first string in the second string are added to the partial string "jidoosoodaNseNta" as the replacement target range. Also, in the example shown in Fig. 7, the selection unit 112α provisionally selects the range of the word "Automatic Consultation Senda" in the first text corresponding to the partial string "jidoosoodaNseNta" as the replacement target range.

[0057] [Example of post-processing] In the examples shown in Fig. 6 and Fig. 7, the selection unit 112α first calculates the ratio of the length of the partial string to the length of the second string, and a preset value for determining whether or not to select a replacement range. For example, the selection unit 112α calculates the ratio of the length of the second string "jidousoudaNseNtaa" in Fig. 7 = 17 to the length of the partial string corresponding to the replacement range shown in Fig. 6, which is the length of the partial string from the character of the first string that matches the second string first ("hit" in Fig. 6) by DP matching to the character that matches last = 18, as 17 / 18 = 0.94. That is, the selection unit 112α calculates the ratio of the length of the second string "jidousoudaNseNtaa" = 17 to the length of the string "jidoousooudaNseNta" edited from the partial string "jidoosoodaNseNtaa" = 18, as 17 / 18 = 0.94. Moreover, the selection unit 112α sets the value Z for determining whether or not to select a replacement target range to, for example, 0.9.

[0058] In the above example, the ratio of the length of the second string to the length of the substring (=0.94) is greater than the value Z (=0.9) for determining whether the replacement range should be selected. Therefore, in the example shown in Fig. 7, the selection unit 112α determines to select, as the replacement range, the range of the string "jidoousooudaNseNta" edited from the substring "jidoosoodaNseNta" corresponding to the replacement range provisionally selected by preprocessing, and the range of the word "automatic consultation Senda" in the first text corresponding to the substring.

[0059] 8, the selection unit 112α continues the selection process of the replacement range in the same manner as in the above example after the above replacement range is replaced with the second character string and the second text by the replacement unit 113. For example, the selection unit 112α selects, as the replacement range, the range of the character string edited from the partial character string "yoNkohama" corresponding to the replacement range and the range of the word "yoNkohama" in the first text corresponding to the partial character string.

[0060] [Effects of the second embodiment] In the second embodiment, the first character string is a character string representing the sound when the first text is spoken, and the second character string is a character string representing the sound when the second text is spoken, so that vowels and consonants of pronunciation are more easily distinguished than readings, and therefore the replacement device 10α can select the replacement target range in units finer than readings. In addition, the replacement device 10α can select the replacement target range based on the length of the partial character string and the length of the second character string, so that the replacement target range can be selected by selecting those character strings whose lengths are close to each other as the replacement target range, and therefore the replacement target range can be selected with high accuracy compared to selection of the replacement target range not based on the length of the partial character string and the length of the second character string.

[0061] [3. Third embodiment] In the replacement system 1β according to the third embodiment, when there are a plurality of second character strings, the replacement device 10β selects the replacement range in order from the replacement range corresponding to the longest second character string, and further replaces the first character string corresponding to the replacement range replaced with the second text with a character string not to be replaced. Other than this, the replacement system 1β is the same as the replacement system 1α according to the second embodiment.

[0062] When there are a plurality of second character strings, the selection unit 112β of the replacement device 10β selects the replacement target range in order from the replacement target range corresponding to the longest second character string. The replacement unit 113β of the replacement device 10β further replaces the first character string corresponding to the replacement target range replaced with the second text with a character string indicating that it is not to be replaced.

[0063] As described above, when there are a plurality of second strings, the replacement device 10β of the replacement system 1β can easily avoid comparison with a partially short second string and replacement with the second text by selecting the replacement target range in order from the replacement target range corresponding to the second string with the longest length. In addition, the replacement device 10β further replaces the first string corresponding to the replacement target range replaced with the second text with a string indicating that it is not a replacement target. For example, when there are a plurality of second strings and the comparison of the first string with each of the plurality of second strings and the replacement are repeated (looped), the replacement device 10β replaces this first string with a string indicating that it is not a replacement target. In this way, when investigating the possibility of replacement with other second strings in the word dictionary or other second texts corresponding to the other second strings, the replacement device 10β can exclude the already replaced range from the replacement target. Note that, as a termination condition of the loop process when there are a plurality of second strings, the loop process is executed until a series of processes is completed for all the second strings.

[0064] [Example of replacement device configuration] Fig. 9 is a diagram showing an example of the configuration of a substitution device according to the third embodiment. In the example shown in Fig. 9, the substitution device 10β has a control unit 11β instead of the control unit 11α in the second embodiment. Except for this point, the substitution device 10β is similar to the substitution device 10α according to the second embodiment.

[0065] [Control Unit] The control unit 11β has a selection unit 112β and a replacement unit 113β instead of the selection unit 112α and the replacement unit 113 in the second embodiment. Except for this, the control unit 11β is similar to the control unit 11α in the second embodiment.

[0066] [Selection Department] When there are a plurality of second character strings, the selection unit 112β selects the replacement target range in order from the replacement target range corresponding to the second character string with the longest length. For example, as shown in FIG. 7, when there are a plurality of second character strings registered in advance in the word dictionary, the selection unit 112β selects the replacement target range in order from the replacement target range corresponding to the range of the second character string with the longest length, "jidousoudaNseNtaa", ..., "ji". As an example, the selection unit 112β performs DP matching between the second character strings "jidousoudaNseNtaa", ..., "ji" in this order and the first character string, obtains a partial character string based on the first character string that matches the second character string, and selects the range of the partial character string or the range of the first text corresponding to the partial character string as the replacement target range when the corrected edit distance obtained by subtracting the difference between the length of the first character string and the length of the second character string from the edit distance between the first character string and the second character string is smaller than a preset value. In this way, selection unit 112β selects the range to be replaced by comparing the second string with the first string in order of length, starting from the longest second string, and has replacement unit 113β replace it, thereby making it easier to avoid comparison with the second string, which is partially shorter, and replacement with the second text.

[0067] [Replacement part] The replacing unit 113β further replaces the first character string corresponding to the replacement target range replaced by the second text with a character string indicating that it is not a replacement target. It is preferable to use characters and symbols other than those used as phonemes and phonetic symbols as the character string indicating that it is not a replacement target. An example of replacement by the replacing unit 113β will be described below with reference to FIG. 10. FIG. 10 is a diagram for explaining an example of replacement. In the example shown in FIG. 10, the replacing unit 113β further replaces the first character string "jidousoudaNseNtaa" corresponding to the replacement target range replaced by the second text with a character string "XXXXXXXXXXXXXXXXX" indicating that it is not a replacement target. Here, "XXXXXXXXXXXXXXXXX" is a character string composed of characters equivalent to the number of characters of the second character string "jidousoudaNseNtaa" corresponding to the first character string "jidoosoodaNseNta". It is not necessary to replace the character string indicating that it is not a replacement target as described above, and for example, it is sufficient to simply record (or mark) that it has been replaced.

[0068] Here, when there are multiple second character strings, for example, the comparison of the first character string with each of the multiple second character strings by the selection unit 112β and the replacement by the replacement unit 113β are repeated (looped). In this case, the replacement unit 113β further replaces the first character string corresponding to the replacement target range replaced with the second text or the like with a character string indicating that it is not to be replaced. In this way, the replacement unit 113β can exclude the already replaced range from the target of replacement when the selection unit 112β is made to check the possibility of replacement with other second character strings in the word dictionary or other second texts corresponding to the other second character strings.

[0069] [4. Fourth embodiment] In the fourth embodiment, the replacement device 10γ according to the fourth embodiment may register the second character string. FIG. 11 is a diagram for explaining an example of an outline of the replacement system according to the fourth embodiment. As shown in FIG. 11, the replacement device 10γ of the replacement system 1γ registers the second character string. Except for this, the replacement system 1γ is similar to the replacement system 1α according to the second embodiment. In the example shown in FIG. 11, the registration unit 114 of the replacement device 10γ registers the second character string. For example, the registration unit 114 adds the second phoneme string associated with the word to a word dictionary.

[0070] [Example of replacement device configuration] 12 is a diagram showing an example of the configuration of a substitution device according to the fourth embodiment. The substitution device 10γ has a control unit 11γ instead of the control unit 11α in the second embodiment. Except for this, the substitution device 10γ is similar to the substitution device 10α in the second embodiment.

[0071] [Control Unit] The control unit 11γ further includes a registration unit 114. The registration unit 114 registers a second character string. For example, the registration unit 114 adds a second phoneme string associated with a word and the length of the phoneme string to the word dictionary as shown in Fig. 5. As an example, the registration unit 114 adds a second phoneme string "jidousoudaNseNtaa" associated with the word "child consultation center" and the length of the phoneme string "17" to the word dictionary.

[0072] For example, the registration unit 114 may exclude second phoneme strings shorter than the threshold value from the registration target. In this case, the registration unit 114 can arbitrarily set, for example, an upper limit value and a lower limit value of the threshold value. As an example, when the registration unit 114 sets the threshold value to 5, the registration unit 114 excludes "ji", which is a second phoneme string with a length of 2, from the registration target when the threshold value is set to 5. This makes it easier for the registration unit 114 to prevent replacement with a second phoneme string whose replacement target range is shorter than the threshold value.

[0073] [Example of processing flow] An example of the flow of a process executed by a substitution device 10γ of a substitution system 1γ according to the fourth embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing an example of the flow of a process executed by the substitution system according to the fourth embodiment.

[0074] In step S21, the registration unit 114 registers a second character string. For example, when it is determined that speech recognition is performed using an acoustic model and a language model, the registration unit 114 adds a second phoneme string associated with a word and the length of the phoneme string to the word dictionary as shown in Fig. 5. Note that the processes from step S22 onwards are the same as those in the second embodiment described above, and therefore will not be described.

[0075] [5. Fifth embodiment] In the fifth embodiment, the replacement device 10δ according to the fifth embodiment may acquire the reading of the first text and convert the reading of the first text into the first phoneme string. FIG. 14 is a diagram for explaining an example of an outline of the replacement system according to the fifth embodiment. As shown in FIG. 14, the replacement device 10δ of the replacement system 1δ acquires the reading of the first text and converts the reading of the first text into the first phoneme string. Except for this, the replacement system 1δ is similar to the replacement system 1α according to the second embodiment. In the example shown in FIG. 14, the reading acquisition unit 115 of the replacement device 10δ acquires the reading of the first text. The conversion unit 116 of the replacement device 10δ converts the reading of the first text into the first phoneme string.

[0076] [Example of replacement device configuration] Fig. 15 is a diagram showing an example of the configuration of a replacement device according to the fifth embodiment. In the example shown in Fig. 15, the replacement device 10δ has a control unit 11δ instead of the control unit 11α in the second embodiment, and further has a reading acquisition unit 115.

[0077] [Control Unit] Control unit 11δ further includes a reading acquisition unit 115 and a conversion unit 116. Except for this, control unit 11δ is similar to control unit 11α in the second embodiment.

[0078] [Reading acquisition section] The reading acquisition unit 115 acquires the reading of the first text. As an example, the reading acquisition unit 115 performs morpheme analysis on the first text in a language that is not segmented, such as Japanese, to acquire the reading for each morpheme. The reading acquisition unit 115 may perform the morpheme analysis using, for example, a known technique. As another example, the reading acquisition unit 115 acquires a first phoneme string as the reading of the first text from each morpheme of the first text using a pronunciation dictionary or a G2P (Grapheme-to-Phoneme) conversion library, etc., since morpheme analysis is not necessary for the first text in a language that is segmented, such as English.

[0079] [Conversion section] The conversion unit 116 converts the reading of the first text into a first phoneme string. As an example, the conversion unit 116 converts the reading of each morpheme of the first text in a language that is not segmented, such as Japanese, into a first phoneme string. In this case, the conversion unit 116 may convert the reading of each morpheme of the first text into the first phoneme string, for example, via a process of converting the reading of each morpheme of the first text into a romanized character. In addition, the conversion unit 116 may perform conversion from the reading of the first text into the first phoneme string, for example, by using a known technique. As another example, the conversion unit 116 directly obtains the first phoneme string from each morpheme of the first text in a language that is segmented, such as English, by using a pronunciation dictionary or a G2P conversion library. The conversion unit 116, for example, concatenates the first phoneme string of each morpheme. Furthermore, the conversion unit 116 causes the storage unit 14 to store, for example, the correspondence between the position of each phoneme in each first phoneme string and a morpheme.

[0080] As a result, regardless of whether the first text, which is the speech recognition result, is a first text in a language that is not segmented, such as Japanese, or a first text in a language that is segmented, such as English, the conversion unit 116 can ultimately cause the string acquisition unit 111α to acquire the first phoneme string and the selection unit 112α to select the replacement target range with a specified accuracy.

[0081] An example of conversion by the conversion unit 116 will be described with reference to Fig. 16. Fig. 16 is a diagram for explaining an example of conversion. In the example shown in Fig. 16, the conversion unit 116 converts all morphemes "watashi", ··· "desu" of the first text of the speech recognition result, "watashi", ··· "desu", into a first phoneme string "watashi", ··· "desU".

[0082] [Example of processing flow] An example of the flow of the process executed by the replacement device 10δ of the replacement system 1δ according to the fifth embodiment will be described with reference to Fig. 17. Fig. 17 is a flowchart showing an example of the flow of the process executed by the replacement system according to the fifth embodiment. Steps S31, S34, and S35 in Fig. 17 are similar to the processes in the second embodiment, and therefore will not be described.

[0083] In step S32, the reading acquisition unit 115 acquires the reading of the first text. As an example, the reading acquisition unit 115 executes morpheme analysis for the first text in a language that is not segmented, such as Japanese, to acquire the reading for each morpheme. As another example, the reading acquisition unit 115 acquires a first phoneme string as the reading of the first text from each morpheme of the first text using a pronunciation dictionary, a G2P conversion library, or the like, since morpheme analysis is not necessary for the first text in a language that is segmented, such as English.

[0084] In step S33, the conversion unit 116 converts the reading of the first text into a first phoneme string. As an example, the conversion unit 116 converts the reading of each morpheme of the first text in a language that is not segmented, such as Japanese, into the first phoneme string. As another example, the conversion unit 116 directly obtains the first phoneme string from each morpheme of the first text in a language that is segmented, such as English, by using a pronunciation dictionary or a G2P conversion library.

[0085] [Variations] 17, in step S32 after the character string acquisition unit 111α acquires the first character string and the second character string in step S31, the reading acquisition unit 115 acquires the reading of the first text, and in step S33, the conversion unit 116 converts the reading of the first text into a first phoneme string, but this embodiment is not limited to such a form. In this embodiment, it is sufficient that the processes of steps S31 to S33 are performed before step S34, and the process of step S32 is performed before step S33. For example, the processes may be performed in the order of step S32, step S33, and step S31.

[0086] [6. Sixth embodiment] In the fifth embodiment, the reading acquisition unit 115 of the replacement device 10δ acquires the reading of the first text, and the conversion unit 116 converts the reading of the first text into the first phoneme string, but the replacement device according to one aspect of the present invention is not limited to the example of the fifth embodiment. The acquisition unit of the replacement device according to one aspect of the present invention acquires at least one of the reading of the first text and the reading of the second text, and the conversion unit may convert the reading of the first text into the first phoneme string when the reading of the first text is acquired by the acquisition unit, and convert the reading of the second text into the second phoneme string when the reading of the second text is acquired by the acquisition unit.

[0087] [Example of replacement device configuration] 18 is a diagram showing an example of the configuration of a replacement device according to the sixth embodiment. The replacement device 10ε has a reading acquisition unit 115ε and a control unit 11ε instead of the reading acquisition unit 115 and the control unit 11δ in the fifth embodiment. Other than this, the replacement device 10ε is similar to the replacement device 10δ according to the fifth embodiment.

[0088] [Control Unit] The control unit 11ε has a reading acquisition unit 115ε instead of the reading acquisition unit 115 in the fifth embodiment, and has a conversion unit 116ε instead of the conversion unit 116. Except for this, the control unit 11ε is similar to the control unit 11δ in the fifth embodiment.

[0089] [Reading acquisition section] The reading acquisition unit 115ε acquires the reading of the first text and the reading of the second text. For example, the reading acquisition unit 115ε acquires the reading “JIDOUSOUDANSENTA” of the word “CHILD CONSULTATION CENTER” registered in the word dictionary of FIG. 5 by a method similar to that of the reading acquisition unit 115 in the fifth embodiment.

[0090] [Conversion section] The conversion unit 116ε converts the reading of the first text into a first phoneme string when the reading acquisition unit 115ε acquires the reading of the first text, and converts the reading of the second text into a second phoneme string when the reading acquisition unit 115ε acquires the reading of the second text. For example, when the reading acquisition unit 115ε acquires the reading "jidousoudansenta" of the word "child consultation center" of the second text registered in the word dictionary of FIG. 5 as the reading of the second text, the conversion unit 116ε converts the reading "jidousoudansenta" into a second phoneme string of "jidousoudaNseNtaa". In this way, even if the second character string registered in the word dictionary does not include the second phoneme string, the conversion unit 116ε can convert the reading of the second text into the second phoneme string, and as a result, the character string acquisition unit 111α can acquire the second phoneme string.

[0091] [Example of processing flow] An example of the flow of the process executed by the substitution device 10ε of the substitution system 1ε according to the sixth embodiment will be described with reference to Fig. 19. Fig. 19 is a flowchart showing an example of the flow of the process executed by the substitution system according to the sixth embodiment. Steps S41, S44, and S45 in Fig. 19 are the same as steps S31, S34, and S35 in the fifth embodiment.

[0092] In step S42, the reading acquisition unit 115ε acquires the reading of the first text and the reading of the second text. For example, the reading acquisition unit 115ε acquires the reading "JIDOUSOUDANSENTA" of the word "CHILD CONSULTATION CENTER" in the second text registered in the word dictionary of FIG. 5 by a method similar to that of the reading acquisition unit 115 in the fifth embodiment.

[0093] In step S43, when the reading of the first text is acquired by the reading acquisition unit 115ε, the conversion unit 116ε converts the reading of the first text into a first phoneme string, and when the reading of the second text is acquired by the reading acquisition unit 115ε, the conversion unit 116ε converts the reading of the second text into a second phoneme string. For example, when the reading acquisition unit 115ε acquires the reading of the second text "jidousoudansenta" of the word "child consultation center" registered in the electronic dictionary of Fig. 5, the conversion unit 116ε converts the reading "jidousoudansenta" into a second phoneme string "jidousoudaNseNtaa".

[0094] [Variations] 19, in step S42 after the character string acquisition unit 111α acquires the first character string and the second character string in step S41, the reading acquisition unit 115ε acquires the reading of the first text and the reading of the second text, and in step S43, the conversion unit 116ε converts the reading of the first text into a first phoneme string and converts the reading of the second text into a second phoneme string, but this embodiment is not limited to such a form. In this embodiment, it is sufficient that each process of steps S41 to S43 is performed before step S44, and each process of step S42 is performed before step S43, and for example, each process may be performed in the order of step S42, step S43, and step S41.

[0095] [7. Seventh embodiment] In the seventh embodiment, the replacement device 10ζ according to the seventh embodiment may add at least one of a period, a comma, a question mark, and an exclamation mark to the first text, and divide the first text based on the at least one of a period, a comma, a question mark, and an exclamation mark added to the first text.

[0096] FIG. 20 is a diagram for explaining an example of an outline of the replacement system according to the seventh embodiment. In the example shown in FIG. 20, a replacement device 10ζ of the replacement system 1ζ adds at least one of a period, a comma, a question mark, and an exclamation mark to a first text, and divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark added to the first text. A period refers to, for example, ".", ".", etc. A comma refers to, for example, ",", ",", etc. A question mark refers to, for example, "?", etc. An exclamation mark refers to, for example, "!", etc.

[0097] In the example shown in Fig. 20, the adding unit 117 of the replacement device 10ζ adds at least one of a period, a comma, a question mark, and an exclamation mark to the first text. In the example shown in Fig. 20, the dividing unit 118 of the replacement device 10ζ divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark added to the first text.

[0098] In this way, the replacement device 10ζ adds at least one of a period, a comma, a question mark, and an exclamation mark to the first text, and divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark added to the first text. This makes it easier for the replacement device 10ζ to perform DP matching based on the first character string and the second character string for the divided first text, making it easier to select the replacement target range even if the first text does not have a punctuation mark. Therefore, even if the first text does not have a punctuation mark, the replacement device 10ζ can replace the replacement target range with higher accuracy than when the first text is not divided.

[0099] [Example of replacement device configuration] An example of the configuration of the replacement device 10ζ according to the seventh embodiment will be described with reference to Fig. 21. Fig. 21 is a diagram showing an example of the configuration of the replacement device according to the seventh embodiment. The replacement device 10ζ has a control unit 11ζ instead of the control unit 11α in the second embodiment. Except for this point, the replacement device 10ζ is similar to the replacement device 10α according to the second embodiment.

[0100] [Control Unit] The control unit 11ζ further includes an adding unit 117 and a dividing unit 118. Except for this, the control unit 11ζ is similar to the control unit 11α in the second embodiment.

[0101] [Granting section] The adding unit 117 adds at least one of a period, a comma, a question mark, and an exclamation mark to the first text. For example, when none of a period, a comma, a question mark, and an exclamation mark is added to the first text, the adding unit 117 adds at least one of a period, a comma, a question mark, and an exclamation mark. The adding unit 117 can add at least one of a period, a comma, a question mark, and an exclamation mark to the first text by any method, regardless of the type of the method based on rules or models. For example, if the adding unit 117 is a method based on models, the adding unit 117 can create a model that learns a pair of a first text without punctuation and a first text with punctuation by a known technique such as that described in the following reference documents, and add at least one of a period, a comma, a question mark, and an exclamation mark to the first text by the created model. Reference: Spoken-to-written translation corpus for Japanese texts, Iori et al., Proceedings of the 26th Annual Conference of the Association for Natural Language Processing (March 2020)

[0102] [Divided part] The division unit 118 divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark added to the first text. For example, when at least one of a period, a comma, a question mark, and an exclamation mark is added to the first text in advance or by the adding unit 117, the division unit 118 divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark. As an example, when dividing the first text based on a period, the division unit 118 scans from the first character to the period and performs a division process treating the period as one sentence. For example, when there are multiple periods, the division unit 118 performs a division process to repeatedly divide the character next to the period to the period as one sentence. In addition, the division unit 118 treats, for example, the character next to the last period to the character at the end of the sentence as one sentence. This allows the division unit 118 to make it easier for the selection unit 112α to perform DP matching and the like, and therefore makes it easier to select the replacement target range even when the first text does not contain punctuation marks. As a result, even when the first text does not contain punctuation marks, the division unit 118 allows the replacement unit 113 to replace the replacement target range with high accuracy compared to a case where the first text is not divided.

[0103] [Example of processing flow] An example of the flow of the process executed by the substitution device 10ζ of the substitution system 1ζ according to the seventh embodiment will be described with reference to Fig. 22. Fig. 22 is a flowchart showing an example of the flow of the process executed by the substitution system according to the seventh embodiment. Steps S51, S55, and S56 in Fig. 22 are similar to the processes in the second embodiment, and therefore will not be described.

[0104] In step S52, the adding unit 117 determines whether or not the first text does not include punctuation marks, etc. If the adding unit 117 determines that the first text does not include punctuation marks, etc. (No in step S52), in step S53, the adding unit 117 adds punctuation marks, etc. to the first character string and proceeds to step S54. For example, if the adding unit 117 determines that the first text does not include any of a period, a comma, a question mark, and an exclamation mark, the adding unit 117 adds at least one of a period, a comma, a question mark, and an exclamation mark to the first text.

[0105] When the adding unit 117 determines that the first text contains punctuation marks or the like (Yes in step S52), the adding unit 117 skips step S53 and proceeds to step S54. For example, when the adding unit 117 determines that the first text contains at least one of a period, a comma, a question mark, and an exclamation mark, the adding unit 117 skips adding punctuation marks or the like to the first text.

[0106] In step S54, when at least one of a period, a comma, a question mark, and an exclamation mark is added to the first text by the adding unit 117, the dividing unit 118 divides the first text based on at least one of a period, a comma, a question mark, and an exclamation mark. For example, when dividing the first text based on a period, the dividing unit 118 scans from the first character to the period and performs a dividing process treating the first text as one sentence, and when there are multiple periods, performs a dividing process to divide the first text by repeatedly treating the first text as one sentence from the character next to the period to the period. For example, the dividing unit 118 treats the first text as one sentence from the character next to the last period to the character at the end of the sentence.

[0107] [Variations] 22, in steps S52 and S53 after the character string acquisition unit 111α acquires the first character string and the second character string in step S51, the assigning unit 117 determines whether to assign punctuation marks, etc. to the first text and assigns punctuation marks, etc. to the first text, and in step S54, the division unit 118 divides the first text based on the punctuation marks, etc., but this embodiment is not limited to such a form. In this embodiment, it is sufficient that the processes of steps S51 to S54 are executed before step S55, and the process of step S52 is executed before step S53, and for example, each process may be executed in the order of step S52, step S53, and step S51.

[0108] [8. Program] The substitution device 10, the substitution device 10α, the substitution device 10β, the substitution device 10γ, the substitution device 10δ, the substitution device 10ε, and the substitution device 10ζ can be implemented by installing a substitution program in a computer as package software or online software. For example, the computer can function as the substitution device 10, the substitution device 10α, the substitution device 10β, the substitution device 10γ, the substitution device 10δ, the substitution device 10ε, and the substitution device 10ζ by executing the substitution program.

[0109] Fig. 23 is a diagram showing an example of the configuration of a computer that executes a replacement program. In the example shown in Fig. 23, a computer 1000 has a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. Each of these units is connected by a bus 1080.

[0110] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1090. The disk drive interface 1040 is connected to a disk drive 1100. For example, a removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0111] The hard disk drive 1090 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, a program that specifies each process executed by the replacement device 10, the replacement device 10α, the replacement device 10β, the replacement device 10γ, the replacement device 10δ, the replacement device 10ε, and the replacement device 10ζ is implemented as a program module 1093 in which a code executable by a computer is described. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing a process similar to the functional configuration in the replacement device 10, the replacement device 10α, the replacement device 10β, the replacement device 10γ, the replacement device 10δ, the replacement device 10ε, and the replacement device 10ζ is stored in the hard disk drive 1090. The hard disk drive 1090 may be replaced by an SSD.

[0112] Data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, the memory 1010 or the hard disk drive 1090. The CPU 1020 reads out the program module 1093 or the program data 1094 stored in the memory 1010 or the hard disk drive 1090 into the RAM 1012 as necessary and executes them.

[0113] Note that the program module 1093 and the program data 1094 are not limited to being stored in the hard disk drive 1090, and may be stored in, for example, a removable storage medium and read by the CPU 1020 via the disk drive 1100 or the like. Alternatively, the program module 1093 and the program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or wide area network (WAN)). The program module 1093 and the program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070. [Explanation of symbols]

[0114] 1 Substitution System 10 Replacement device 11 Control section 111 String acquisition part 112 Selection Department 113 Substitution part

Claims

1. an acquisition unit that acquires a first character string related to a first text that is a speech recognition result and a second character string related to a second text that is registered in advance; a selection unit that selects a replacement target range from the first text based on a similarity between at least a part of a substring of the first string and the second string; a replacement unit that replaces the replacement target range with the second text; A replacement device comprising:

2. the first character string is a character string representing a sound when the first text is spoken, 2. The replacement device according to claim 1, wherein the second character string is a character string representing a sound when the second text is spoken.

3. 2. The replacement device according to claim 1, wherein the selection unit selects the replacement target range based on the similarity as well as the length of the substring and the length of the second string.

4. The replacement device according to claim 1 , wherein the selection unit calculates an edit distance between the partial string and the second string as the similarity, and selects the replacement target range based on the edit distance.

5. 4. The replacing device according to claim 3, wherein the selection unit selects the replacement target range based on a ratio or a magnitude relationship between a length of the partial character string and a length of the second character string.

6. the selection unit, when there are a plurality of the second character strings, selects the replacement target range in order from the replacement target range corresponding to the second character string having the longest length; The replacement device according to claim 1 , wherein the replacement unit further replaces the first character string corresponding to the replacement target range replaced with the second text with a character string indicating that the first character string is not a target for replacement.

7. A replacement method performed by a replacement device, comprising: an acquiring step of acquiring a first character string related to a first text, which is a speech recognition result, and a second character string related to a second text, which is registered in advance; a selection step of selecting a replacement target range from the first text based on a similarity between at least a part of a substring of the first string and the second string; a replacing step of replacing the replacement target range with the second text; A method for substitution comprising:

8. A replacement program for causing a computer to function as the replacement device according to any one of claims 1 to 6, the replacement program causing the computer to function as the acquisition unit, the selection unit, and the replacement unit.

Citation Information

Patent Citations

  • Word additional device, word addition method, and program thereof

    JP2012242421A

Cited By

  • Video updating method and device, equipment, medium and product

    CN120980298A

  • Information processing systems, information processing methods, and programs

    JP7864402B1