Period deletion model learning device, period deletion model, and determination device
The period deletion model learning device addresses the issue of incorrect full stop placement in speech recognition text by using machine learning to analyze consecutive sentences and output correctness probabilities, thereby enhancing the accuracy of full stop addition.
Patent Information
- Application Number
- JP2022516955
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-20
- Filing Date
- 2021-04-08
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2041-04-08
AI Technical Summary
Text obtained from speech recognition often contains errors, leading to incorrect placement of full stops, and existing techniques lack accuracy in adding full stops due to the lack of reference to speech information.
A period deletion model learning device that generates a period deletion model through machine learning, which determines the correctness of full stops in text obtained from speech recognition by analyzing pairs of consecutive sentences and outputting a probability indicating the correctness of the full stop.
The solution enables the accurate addition of full stops to text from speech recognition, improving readability and translatability by correctly identifying and correcting incorrectly placed full stops.
Smart Images

Figure 0007682862000003 
Figure 0007682862000004 
Figure 0007682862000005
Abstract
Description
Technical Field
[0001] The present invention relates to a full stop deletion model learning device, a full stop deletion model, and a determination device.
Background Art
[0002] In order to improve readability and translatability of text obtained as a result of speech recognition, it is required to appropriately add full stops. On the other hand, there is known a technique of dividing an input sentence by a division candidate point determined based on a pattern of the arrangement of clauses and a pattern of syntactic dependencies between clauses in the input document (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The text obtained from the speech recognition result often contains errors, and due to these recognition errors, full stops may be added to incorrect locations. Further, when dividing the input sentence based only on the texturized input sentence as in the technique described in the above Document 1 and adding a full stop at the end of the divided sentence, the accuracy of full stop addition was low because the information of the speech corresponding to the input sentence was not referred to.
[0005] Therefore, the present invention has been made in view of the above problems, and an object thereof is to appropriately add full stops to the text obtained by speech recognition processing.
Means for Solving the Problems
[0006] In order to solve the above problems, a period deletion model learning device according to one embodiment of the present invention is a period deletion model learning device that generates, by machine learning, a period deletion model that determines whether a period added to text obtained by speech recognition processing is correct. The period deletion model takes as input two consecutive sentences, a first sentence with a period at the end and a second sentence that follows immediately after the first sentence, and outputs a probability indicating whether the period added at the end of the first sentence is correct. The period deletion model learning device generates first learning data consisting of a pair of an input sentence including one or more sentences obtained by speech recognition processing and being text with a period added based on the information obtained by the speech recognition processing, the input sentence including a previous sentence that is a sentence with a period added at the end and a subsequent sentence that follows immediately after the period, and a label indicating whether the period is correctly added, based on a first text corpus. The period deletion model learning device includes a first learning data generation unit that generates the first learning data, and a model learning unit that updates parameters of the period deletion model based on an error between the probability obtained by inputting the input sentence of the first learning data into the period deletion model and the label associated with the input sentence.
[0007] According to the above embodiment, since two sentences consisting of the previous sentence immediately before the period inserted based on the information obtained by the speech recognition processing and the subsequent sentence immediately after the period are used as the input sentence in the first learning data and the period deletion model is learned, the information obtained from the speech is reflected in the period deletion model. Further, in the first learning data, a label indicating whether the period included in the input sentence is correct is associated with each input sentence, and the period deletion model is learned based on the error between the probability obtained by inputting the input sentence into the period deletion model and the label. Therefore, it is possible to obtain a period deletion model that can correctly delete an incorrectly added period in the text obtained by speech recognition processing.
Effects of the Invention
[0008] A period deletion model learning device, a period deletion model, and a determination device that can appropriately add periods to the text obtained by speech recognition processing are realized.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Embodiments for Carrying Out the Invention
[0010] Embodiments of the period deletion model learning device, determination device, and period deletion model according to the present invention will be described with reference to the drawings. In addition, when possible, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.
[0011] The period deletion model learning device of the present embodiment is a device that generates a period deletion model for determining the correctness of a period given to text obtained by speech recognition processing by machine learning. The period deletion model of the present embodiment determines the correctness of each period in the text including the period obtained by speech recognition processing. Specifically, the period deletion model takes as input two consecutive sentences, a first sentence with a period at the end and a second sentence following immediately after the first sentence, and outputs as output the probability indicating the correctness of the period given at the end of the first sentence, and is a model constructed by machine learning.
[0012] The determination device of the present embodiment determines the correctness of a period given to the text to be determined obtained by speech recognition processing, deletes the period determined to be incorrectly given from the text to be determined, and outputs the text to be determined with the period corrected.
[0013] FIG. 1 is a diagram showing a functional configuration of a period deletion model learning device according to the present embodiment. As shown in FIG. 1, the period deletion model learning device 10 functionally includes a first learning data generation unit 11, a second learning data generation unit 12, and a model learning unit 13. Each of these functional units 11 to 13 may be configured in one device or may be distributed and configured in a plurality of devices.
[0014] Further, the period deletion model learning device 10 is configured to be able to access storage means such as a first / second text corpus storage unit 15, a third / fourth text corpus storage unit 16, and a period deletion model storage unit 17. The first / second text corpus storage unit 15, the third / fourth text corpus storage unit 16, and the period deletion model storage unit 17 may be configured within the period deletion model learning device 10, or may be configured as another device outside the period deletion model learning device 10 as shown in FIG. 1.
[0015] FIG. 2 is a diagram showing a functional configuration of a determination device according to the present embodiment. As shown in FIG. 2, the determination device 20 functionally includes an input sentence extraction unit 21, a determination unit 22, and a period correction unit 23. Each of these functional units 21 to 23 may be configured in one device or may be distributed and configured in a plurality of devices.
[0016] Further, the determination device 20 is configured to be able to access a determination target text storage unit 25 and a period deletion model storage unit 17. The determination target text storage unit 25 and the period deletion model storage unit 17 may be configured within the determination device 20 or may be configured in another external device.
[0017] In addition, in the present embodiment, an example is shown in which the period deletion model learning device 10 and the determination device 20 are each configured as separate devices (computers), but they may also be integrally configured.
[0018] Note that the block diagrams shown in FIGS. 1 and 2 show the blocks of functional units. These functional blocks (components) are realized by any combination of at least one of hardware and software. Also, the realization method of each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or two or more physically or logically separated devices may be directly or indirectly (e.g., using wired, wireless, etc.) connected and realized using these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.
[0019] Functions include, but are not limited to, judgment, decision, determination, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, solution, selection, selection, establishment, comparison, assumption, expectation, regarded as, notification (broadcasting), notification (notifying), communication (communicating), forwarding, configuration (configuring), reconfiguration (reconfiguring), allocation (allocating, mapping), assignment (assigning), etc. For example, a functional block (component) that functions as transmission is called a transmitting unit or a transmitter. In any case, as described above, the realization method is not particularly limited.
[0020] For example, the period deletion model learning device 10 and the determination device 20 in an embodiment of the present invention may function as a computer. FIG. 3 is a diagram showing an example of the hardware configuration of the period deletion model learning device 10 and the determination device 20 according to this embodiment. The period deletion model learning device 10 and the determination device 20 may each be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.
[0021] In the following description, the term "device" can be read as a circuit, a device, a unit, etc. The hardware configurations of the period deletion model learning device 10 and the determination device 20 may be configured to include one or more of the devices shown in the figures, or may be configured without including some of the devices.
[0022] Each function in the period deletion model learning device 10 and the determination device 20 is realized by causing a processor 1001 to perform operations by loading a predetermined software (program) onto hardware such as the processor 1001, the memory 1002, etc., and controlling communication by the communication device 1004 and reading and / or writing data in the memory 1002 and the storage 1003.
[0023] The processor 1001 controls the entire computer by operating an operating system, for example. The processor 1001 may be composed of a central processing unit (CPU: Central Processing Unit) including an interface with peripheral devices, a control device, an arithmetic device, a register, etc. For example, each functional unit 11 to 13, 21 to 23 shown in FIGS. 1 and 2 may be realized by the processor 1001.
[0024] Also, the processor 1001 reads a program (program code), a software module, and data from the storage 1003 and / or the communication device 1004 into the memory 1002, and executes various processes according to these. As the program, a program for causing a computer to execute at least a part of the operations described in the above embodiments is used. For example, each functional unit 11 to 13, 21 to 23 of the period deletion model learning device 10 and the determination device 20 may be stored in the memory 1002 and realized by a control program operating on the processor 1001. Although it has been described that the above various processes are executed by one processor 1001, they may be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented with one or more chips. Note that the program may be transmitted from a network via a telecommunication line.
[0025] The memory 1002 is a computer-readable recording medium and may be composed of at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The memory 1002 may also be referred to as a register, cache, main memory (main storage device), etc. The memory 1002 can store a program (program code), software module, etc. executable for implementing the period deletion model learning method and determination method according to an embodiment of the present invention.
[0026] The storage 1003 is a computer-readable recording medium and may be composed of at least one of, for example, optical discs such as CD-ROM (Compact Disc ROM), hard disk drives, flexible disks, magneto-optical disks (e.g., compact discs, digital versatile discs, Blu-ray (registered trademark) discs), smart cards, flash memories (e.g., cards, sticks, key drives), floppy (registered trademark) disks, magnetic strips, etc. The storage 1003 may also be referred to as an auxiliary storage device. The above-described storage medium may be, for example, a database including the memory 1002 and / or the storage 1003, a server, or other appropriate media.
[0027] The communication device 1004 is hardware (a transmission / reception device) for performing communication between computers via a wired and / or wireless network and is also referred to as, for example, a network device, network controller, network card, communication module, etc.
[0028] The input device 1005 is an input device that receives external input (for example, a keyboard, mouse, microphone, switch, button, sensor, etc.). The output device 1006 is an output device that performs output to the outside (for example, a display, speaker, LED lamp, etc.). Note that the input device 1005 and the output device 1006 may have an integrated configuration (for example, a touch panel).
[0029] Also, each device such as the processor 1001 and the memory 1002 is connected by a bus 1007 for communicating information. The bus 1007 may be composed of a single bus or may be composed of different buses between devices.
[0030] Also, the period deletion model learning device 10 and the determination device 20 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented by at least one of these hardware components.
[0031] Next, each functional part of the period deletion model learning device 10 will be described. The first learning data generation unit 11 generates first learning data based on the first text corpus. The first text corpus is text including one or more sentences obtained by speech recognition processing.
[0032] In the first text corpus, full stops are added based on information obtained by speech recognition processing. The method used for adding full stops is not limited. For example, the first text corpus may be text in which full stops are inserted at the end of each speech section delimited by a silent section of a predetermined length or more. Also, the first text corpus may be text obtained by speech recognition processing using a known language model that treats a full stop as one of the words.
[0033] Specifically, the first learning data generation unit 11 generates first learning data consisting of a pair of an input sentence including a previous sentence which is a sentence with a full stop added at the end of the sentence in the text constituting the first text corpus and a subsequent sentence which is a sentence following immediately after the full stop, and a label indicating whether the full stop is added correctly.
[0034] The first learning data generation unit 11 may refer to a second text corpus for adding a label indicating whether the full stop is added correctly in the input sentence. The second text corpus consists of the same text as the first text corpus and is text in which full stops are added at the end of each sentence included in the text.
[0035] The first learning data generation unit 11 may acquire the first text corpus and the second text corpus from the first / second text corpus storage unit 15. The first / second text corpus storage unit 15 is a storage means for storing the first text corpus and the second text corpus.
[0036] Figure 4 is a diagram showing an example of the first text corpus and the corresponding second text corpus stored in the first / second text corpus storage unit 15. As shown in Figure 4, the first learning data generation unit 11 acquires the first text corpus obtained by voice recognition processing such as "Just now, I was introduced. I am DDD. Thank you. Today. This is the schedule." In the first text corpus shown in Figure 4, full stops are added to the silent intervals detected in the voice recognition processing. Also, the first text corpus includes the text "Today is" which is a misrecognition of the voice "Today's".
[0037] Also, the first learning data generation unit 11 acquires the second text corpus such as "Just now, I was introduced. I am DDD. Thank you. Today's schedule." The second text corpus is text with correct full stops added, for example, text input and generated manually.
[0038] The first learning data generation unit 11 extracts full stops from the first text corpus, and further extracts the preceding sentence, which is the sentence immediately before the full stop, and the following sentence, which is the sentence immediately after the full stop, as the input sentence. Then, in the second text corpus, the first learning data generation unit 11 refers to the part of the text corresponding to the input sentence extracted from the first text corpus. If a full stop is added to the referred text, a label indicating that the full stop added to the input sentence is correct is associated with the input sentence. If no full stop is added to the referred text, a label indicating that the full stop added to the input sentence is incorrect is associated with the input sentence.
[0039] Note that when referring to the second text corpus, the first learning data generation unit 11 may perform character-by-character or word-by-word association (alignment) between the first text corpus and the second text corpus using a known algorithm. For example, an algorithm such as DP matching may be adopted.
[0040] FIG. 5 is a diagram showing an example of first learning data generated based on the first text corpus and the second text corpus shown in FIG. 4. The first learning data generation unit 11 extracts an input sentence consisting of the foregoing sentence "Just now, I have introduced" and a period and the following sentence "It is DDD" from the first text corpus. In the second text corpus, since no period is given to the text portion "Just now, I have introduced. It is DDD" corresponding to the extracted input sentence, the first learning data generation unit 11 associates a label "FALSE" indicating that the given period is incorrect with the input sentence, and generates one piece of first learning data.
[0041] Further, the first learning data generation unit 11 extracts an input sentence consisting of the foregoing sentence "It is DDD" and a period and the following sentence "Please do me a favor" from the first text corpus. In the second text corpus, since a period is given to the text portion "It is DDD. Please do me a favor" corresponding to the extracted input sentence, the first learning data generation unit 11 associates a label "TRUE" indicating that the given period is correct with the input sentence, and generates one piece of first learning data.
[0042] Similarly, the first learning data generation unit 11 generates first learning data with the flag "TRUE" associated with the input sentence "Please do me a favor. Today is", and first learning data with the flag "FALSE" associated with the input sentence "Today is. It is a schedule".
[0043] FIG. 6 is a diagram showing an example of an English first text corpus and a corresponding second text corpus. As shown in FIG. 6, the first learning data generation unit 11 acquires a first text corpus obtained by speech recognition processing of "thank you. for that kind introduction. I am glad. to see you." In the first text corpus shown in FIG. 6, a period (full stop) is added to the silent interval detected in the speech recognition processing. Further, the first text corpus includes the text "that" in which the voice "the" is misrecognized.
[0044] In addition, the first learning data generation unit 11 acquires a second text corpus of "thank you for the kind introduction. I am glad to see you." The second text corpus is text with a period (full stop) correctly added, for example, text input and generated manually.
[0045] The first learning data generation unit 11 extracts a period from the first text corpus, and further extracts the previous sentence, which is the sentence immediately before the period, and the following sentence, which is the sentence immediately after the period, as the input sentence. Then, in the second text corpus, the first learning data generation unit 11 refers to the text portion corresponding to the input sentence extracted from the first text corpus. If a period is added to the referred text, a label indicating that the period added to the input sentence is correct is associated with the input sentence. If no period is added to the referred text, a label indicating that the period added to the input sentence is incorrect is associated with the input sentence.
[0046] FIG. 7 is a diagram showing an example of the first learning data in English generated based on the first text corpus and the second text corpus shown in FIG. 6. The first learning data generation unit 11 extracts an input sentence consisting of the previous sentence "thank you", a period, and the following sentence "for that kind introduction" from the first text corpus. In the second text corpus, since the period is not given to the text part "thank you for the kind introduction" corresponding to the extracted input sentence, the first learning data generation unit 11 associates a label "FALSE" indicating that the given period is incorrect with the input sentence, and generates one piece of first learning data.
[0047] Further, the first learning data generation unit 11 extracts an input sentence consisting of the previous sentence "for that kind introduction", a period, and the following sentence "I am glad" from the first text corpus. In the second text corpus, since the period is given to the text part "for the kind introduction. I am glad" corresponding to the extracted input sentence, the first learning data generation unit 11 associates a label "TRUE" indicating that the given period is correct with the input sentence, and generates one piece of first learning data.
[0048] Similarly, the first learning data generation unit 11 generates first learning data with the flag "FALSE" associated with the input sentence "I am glad. to see you".
[0049] The second learning data generation unit 12 generates second learning data based on the third text corpus and the fourth text corpus. The second learning data generation unit 12 may acquire the third text corpus and the fourth text corpus from the third / fourth text corpus storage unit 16.
[0050] The third text corpus consists of texts that have a period properly attached at the end of the sentence. The third text corpus may be, for example, pre-generated and stored in the third / fourth text corpus storage unit 16. Also, texts collected from the web on the Internet may be pre-stored in the third / fourth text corpus storage unit 16 as the third text corpus.
[0051] The fourth text corpus consists of texts in which periods are randomly inserted into the texts constituting the third text corpus. The fourth text corpus may be pre-generated based on each third text corpus and stored in the third / fourth text corpus storage unit 16.
[0052] Specifically, the second learning data generation unit 12 generates second learning data consisting of a pair of an input sentence including a preceding sentence that is a sentence with a period attached at the end of the sentence and a succeeding sentence that is the sentence following immediately after the period, in the text constituting the fourth text corpus, and a label indicating whether the period is correctly attached.
[0053] FIG. 8 is a diagram showing an example of the third text corpus and the corresponding fourth text corpus stored in the third / fourth text corpus storage unit 16. As shown in FIG. 8, the second learning data generation unit 12 acquires the third text corpus "This show is by lottery. How do you feel about it? Please make a reservation for two people." The third text corpus shown in FIG. 8 consists of two sentences each with a period properly attached.
[0054] Also, the second learning data generation unit 12 acquires the fourth text corpus "This show is by lottery. How do you feel about it? A reservation for two people. Please." The fourth text corpus shown in FIG. 8 is a text in which periods are randomly inserted between the words of the corresponding third text corpus, and periods are attached at four positions.
[0055] The second learning data generation unit 12 extracts a period from the fourth text corpus, and further extracts the preceding sentence, which is the sentence immediately before the period, and the following sentence, which is the sentence immediately after the period, as the input sentence. Then, in the third text corpus, the second learning data generation unit 12 refers to the text portion corresponding to the input sentence extracted from the fourth text corpus. If a period is given in the referred text, a label indicating that the period given in the input sentence is correct is associated with the input sentence. If no period is given in the referred text, a label indicating that the period given in the input sentence is incorrect is associated with the input sentence.
[0056] Note that when referring to the third text corpus, the second learning data generation unit 12 may perform character-by-character or word-by-word association (alignment) between the fourth text corpus and the third text corpus using a known algorithm. For example, an algorithm called DP matching may be adopted.
[0057] FIG. 9 is a diagram showing an example of the second learning data generated based on the third text corpus and the fourth text corpus shown in FIG. 8. The second learning data generation unit 12 extracts from the fourth text corpus an input sentence consisting of the preceding sentence "This show is by lottery", a period, and the following sentence "What do you think?" In the third text corpus, since no period is given in the text portion "This show is by lottery What do you think?" corresponding to the extracted input sentence, the second learning data generation unit 12 associates a label "FALSE" indicating that the given period is incorrect with the input sentence, and generates one piece of second learning data.
[0058] Further, the second learning data generation unit 12 extracts an input sentence consisting of the previous sentence "How do you do?" (in Japanese), a period, and the following sentence "Reservation for two people" from the fourth text corpus. In the third text corpus, since a period is given to the text part "How do you do? 2-person reservation" corresponding to the extracted input sentence, the second learning data generation unit 12 associates a label "TRUE" indicating that the given period is correct with the input sentence and generates one piece of second learning data.
[0059] Similarly, the second learning data generation unit 12 generates second learning data in which the flag "FALSE" is associated with the input sentence "Reservation for two people. Please."
[0060] FIG. 10 is a diagram showing an example of the third English text corpus and the corresponding fourth text corpus stored in the third / fourth text corpus storage unit 16. As shown in FIG. 10, the second learning data generation unit 12 acquires the third text corpus "you need to share this table with other guests. what would you like to order." The third text corpus shown in FIG. 10 consists of two sentences each properly given a period.
[0061] Further, the second learning data generation unit 12 acquires the fourth text corpus "you need to share this. table with other guests. what would you. like to order." The fourth text corpus shown in FIG. 10 is text in which periods are randomly given between the words of the corresponding third text corpus, and periods are given at four positions.
[0062] The second learning data generation unit 12 extracts a period from the fourth text corpus, and further extracts the preceding sentence, which is the sentence immediately before the period, and the following sentence, which is the sentence immediately after the period, as the input sentence. Then, in the third text corpus, the second learning data generation unit 12 refers to the part of the text corresponding to the input sentence extracted from the fourth text corpus. If a period is given in the referred text, the second learning data generation unit 12 associates with the input sentence a label indicating that the period given in the input sentence is correct. If no period is given in the referred text, the second learning data generation unit 12 associates with the input sentence a label indicating that the period given in the input sentence is incorrect.
[0063] FIG. 11 is a diagram showing an example of the second learning data generated based on the third text corpus and the fourth text corpus shown in FIG. 10. The second learning data generation unit 12 extracts from the fourth text corpus an input sentence consisting of the preceding sentence "you need to share this", a period, and the following sentence "table with other guests". In the third text corpus, since no period is given to the part of the text corresponding to the extracted input sentence, "you need to share this table with other guests", the second learning data generation unit 12 associates with the input sentence a label "FALSE" indicating that the given period is incorrect, and generates one piece of the second learning data.
[0064] Also, the second learning data generation unit 12 extracts from the fourth text corpus an input sentence consisting of the preceding sentence "table with other guests", a period, and the following sentence "what would you". In the third text corpus, since a period is given to the part of the text corresponding to the extracted input sentence, "table with other guests. what would you", the second learning data generation unit 12 associates with the input sentence a label "TRUE" indicating that the given period is correct, and generates one piece of the second learning data.
[0065] Similarly, the second learning data generation unit 12 generates second learning data in which the flag "FALSE" is associated with the input sentence "what would you. like to order".
[0066] For generating the first learning data, in order to associate an appropriate label with the input sentence, text appropriately punctuated with a period corresponding to the first text corpus is required, and for example, the second text corpus is applied to this text. Since the second text corpus is, for example, text input and generated manually, it takes time to obtain a large amount of the second text corpus. By generating the second learning data as described above, the second learning data can be used together with the first learning data for learning the period deletion model, so that it is possible to supplement the amount of learning data for model generation, and as a result, a highly accurate period deletion model can be obtained.
[0067] Referring to FIG. 1 again, the model learning unit 13 performs machine learning to update the parameters of the period deletion model based on the error between the probability obtained by inputting the input sentence of the first learning data into the period deletion model and the label associated with the input sentence.
[0068] When the second learning data generation unit 12 generates the second learning data, the model learning unit 13 may use the second learning data for learning the period deletion model in addition to the first learning data.
[0069] As described above, the period deletion model may be any model as long as it takes an input sentence consisting of a preceding sentence, a period, and a succeeding sentence as input and outputs a probability indicating the correctness of the given period. For example, it may be constructed as a model including a well-known neural network.
[0070] Further, the full stop deletion model may be configured by a recurrent neural network. A recurrent neural network is a neural network configured to update a current hidden state vector using a hidden state vector at a previous time and an input vector at the current time. Learning is performed by comparing the output obtained by sequentially inputting the word sequence constituting the input sentence along time steps with a label so that the error becomes small.
[0071] Also, in the present embodiment, the full stop deletion model may be constructed by a long short-term memory network (LSTM: Long Short-Term Memory). A long short-term memory network is a type of recurrent neural network and is a model capable of learning long-term dependencies in time series data. In a long short-term memory network, by inputting the word sequence of the input sentence one word at a time in time series order to an LSTM block, it is possible to calculate an output while continuously remembering all the words input so far.
[0072] Furthermore, in the present embodiment, the full stop deletion model may be configured by a bidirectional long short-term memory network (BLSTM: Bidirectional Long Short-Term Memory) that uses information in the forward direction (forward) of the word sequence in addition to information in the reverse direction (backward).
[0073] FIG. 12 is a diagram showing an example of learning of a full stop deletion model configured by the bidirectional long short-term memory network of the present embodiment. As shown in FIG. 12, the full stop deletion model md includes an embedding layer el, a forward layer lf configured by a long short-term memory network for inputting the word sequence of the input sentence in the forward direction, a backward layer lb configured by a long short-term memory network for inputting the word sequence of the input sentence in the reverse direction, a hidden layer hd, and an output layer sm. In the example shown in FIG. 12, an input sentence is1, "Please take care. Today is", is input to the full stop deletion model md.
[0074] First, in the embedding layer el, the model learning unit 13 converts each word in the input sentence is1 into a word vector representation to obtain word vectors el1~el7. Here, the method of converting a word into a vector may be a known algorithm such as word2vec or glove, for example.
[0075] The model learning unit 13 inputs each word in the word sequence of the input sentence converted into a vector into each of the LSTM blocks lf(n) (n is an integer greater than or equal to 1 and corresponds to the number of words in the input sentence) of the forward layer lf and the LSTM blocks lb(n) of the backward layer lb. In the example shown in FIG. 12, the forward layer lf is shown as LSTM blocks lf1~lf7 unfolded in the time series forward direction. Also, the backward layer lb is shown as LSTM blocks lb7~lb1 unfolded in the time series reverse direction.
[0076] The model learning unit 13 inputs the word vector el1 into the LSTM block lf1 of the forward layer lf. Next, the model learning unit 13 inputs the word vector el2 and the output of the LSTM block lf1 into the LSTM block lf2. Then, the model learning unit 13 inputs the output of the LSTM block lf(n-1) at the previous time and the word vector el(n) at the current time into the LSTM block lf(n) at the current time to obtain the output of the LSTM block lf(n).
[0077] Also, the model learning unit 13 inputs the word vector el7 into the LSTM block lb7 of the backward layer lb. Next, the model learning unit 13 inputs the word vector el6 and the output of the LSTM block lb7 into the LSTM block lb6. Then, the model learning unit 13 inputs the output of the LSTM block lb(n+1) at the previous time and the word vector el(n) at the current time into the LSTM block lb(n) at the current time to obtain the output of the LSTM block lb(n).
[0078] The model learning unit 13 combines the output obtained by inputting the word sequence of the input sentence into the forward layer lf in the order of the array and the output obtained by inputting the word sequence of the input sentence into the backward layer lb in the reverse order of the array, and outputs the hidden layer hd. When the outputs of the LSTM blocks lf(n) and lb(n) are n-dimensional vectors respectively, as follows, the output of the hidden layer hd becomes a 2n-dimensional vector. Output of the forward layer lf: [f 1 , f 2 ,..., f n Output of the backward layer lb: [b 1 , b 2 ,..., b n => Output of the hidden layer hd: [f 1 , f 2 ,..., f n , b 1 , b 2 ,..., b n
[0079] The model learning unit 13 may input the word sequence from the beginning to the end of the previous text into the forward long short-term memory network in the order of the array in the input sentence, and input the word sequence from the end to the beginning of the subsequent text into the reverse long short-term memory network in the reverse order of the array in the input sentence.
[0080] That is, in the example shown in FIG. 12, the model learning unit 13 sequentially inputs the word vectors el1 to el4 from the beginning to the end of the previous text "Please take care of me" of the input sentence is1 into the LSTM blocks lf1 to lf4 of the forward layer lf in the order of the array, and uses the output of the LSTM block lf4 as the input to the hidden layer hd. Further, the model learning unit 13 sequentially inputs the word vectors el7 to el6 from the end to the beginning of the subsequent text "Today is" of the input sentence is1 into the LSTM blocks lb7 to lb6 of the backward layer lb in the order of the array, and uses the output of the LSTM block lb6 as the input to the hidden layer hd.
[0081] In this way, since the word sequence from the beginning to the end of the previous text is input into the forward long short-term memory network to train the full stop deletion model md, the model md learns the end-of-sentence characteristics just before the full stop. On the other hand, since the word sequence from the end to the beginning of the subsequent text is input into the backward long short-term memory network to train the full stop deletion model md, the model md learns the beginning-of-sentence characteristics just after the full stop. Also, for each of the forward layer lf and the backward layer lb, it is not necessary to input all the word sequences of the input sentence, so the computational complexity is reduced.
[0082] FIG. 13 is a diagram schematically showing the configuration of the hidden layer hd and the output layer sm of the full stop deletion model md. As shown in FIG. 13, the hidden layer h includes the output h of the j-th unit. j Based on the output of the hidden layer hd, the model learning unit 13 calculates the output value z in the output layer sm by the following equations (1) and (2). i
Equation
Equation
[0083] In the above equations (1) and (2), w represents the weight between the j-th unit in the hidden layer hd and the i-th unit in the output layer, and b represents the bias in the i-th unit in the output layer sm. The output layer sm may be composed of a so-called softmax function. The value of each output z in the output layer sm is a probability value whose sum is 1. z is the probability of "should delete the full stop", that is, the probability that the full stop is incorrect, and z is the probability of "should keep the full stop", that is, the probability that the full stop is correct. i,j i 1 2
[0084] In the training phase of the full stop deletion model md, the model learning unit 13 uses the probability values z, z output from the output layer sm. 1 ,z 2 Based on the error between the input sentence is1 and the label associated with it in the training data, the parameters of the full stop deletion model md are updated. For example, the model learning unit 13 updates the parameters of the long short-term memory network by the error backpropagation method so that the difference between the output probability value and the label becomes smaller.
[0085] Therefore, when the full stop deletion model md is applied to the stage of determining whether to delete a full stop, z 1 and z 2 Based on the larger probability value of the two, it is determined whether to delete the full stop. For example, if z 1 > z 2 then it is determined that the full stop should be deleted.
[0086] In this way, the learned full stop deletion model md learned by updating the parameters of the long short-term memory network may be stored in the full stop deletion model storage unit 17. The full stop deletion model storage unit 17 is a storage means for storing the learned and learning process full stop deletion models md.
[0087] FIG. 14 is a diagram showing an example of learning of an English training data of a full stop deletion model constituted by the bidirectional long short-term memory network of the present embodiment. As shown in FIG. 14, the full stop deletion model md includes an embedding layer el, a forward layer lf constituted by a long short-term memory network for inputting a word sequence of an input sentence in the forward direction, a backward layer lb constituted by a long short-term memory network for inputting a word sequence of the input sentence in the backward direction, a hidden layer hd, and an output layer sm. In the example shown in FIG. 14, an input sentence is2, "it makes no difference. to me", is input to the full stop deletion model md.
[0088] First, the model learning unit 13 converts each word of the input sentence is2 into a word vector representation in the embedding layer el, and obtains word vectors el21 to el27. Here, the method of converting a word into a vector may be a known algorithm such as word2vec or glove.
[0089] The model learning unit 13 inputs each word in the word sequence of the input sentence converted into a vector to each of the LSTM blocks lf(n) of the forward layer lf (n is an integer greater than or equal to 1 and corresponds to the number of words in the input sentence) and the LSTM blocks lb(n) of the backward layer lb. In the example shown in FIG. 14, the forward layer lf is shown as LSTM blocks lf1 to lf7 expanded in the time series forward direction. Also, the backward layer lb is shown as LSTM blocks lb7 to lb1 expanded in the time series reverse direction.
[0090] The model learning unit 13 inputs the word vector el21 to the LSTM block lf1 of the forward layer lf. Next, the model learning unit 13 inputs the word vector el22 and the output of the LSTM block lf1 to the LSTM block lf2. Then, the model learning unit 13 inputs the output of the LSTM block lf(n - 1) at the previous time and the word vector el(n) at the current time to the LSTM block lf(n) at the current time, thereby obtaining the output of the LSTM block lf(n).
[0091] Also, the model learning unit 13 inputs the word vector el27 to the LSTM block lb7 of the backward layer lb. Next, the model learning unit 13 inputs the word vector el26 and the output of the LSTM block lb7 to the LSTM block lb6. Then, the model learning unit 13 inputs the output of the LSTM block lb(n + 1) at the previous time and the word vector el(n) at the current time to the LSTM block lb(n) at the current time, thereby obtaining the output of the LSTM block lb(n).
[0092] The model learning unit 13 combines the output obtained by inputting the word sequence of the input sentence to the forward layer lf in the array order and the output obtained by inputting the word sequence of the input sentence to the backward layer lb in the reverse order of the array order, and outputs the hidden layer hd.
[0093] Note that, similar to the example shown in FIG. 12, the model learning unit 13 may input the word sequence from the beginning to the end of the previous sentence, which is the sentence immediately before the period, into the forward long short-term memory network in the order of arrangement in the input sentence, and input the word sequence from the end to the beginning of the subsequent sentence, which is the sentence immediately after the period, into the reverse long short-term memory network in the order reverse to the order of arrangement in the input sentence.
[0094] That is, in the example shown in FIG. 14, the model learning unit 13 sequentially inputs the word vectors el21 to el24 from the beginning to the end of the previous sentence "it makes no difference" of the input sentence is2 into the LSTM blocks lf1 to lf4 of the forward layer lf in the order of arrangement, and uses the output of the LSTM block lf4 as the input to the hidden layer hd. Also, the model learning unit 13 sequentially inputs the word vectors el27 to el26 from the end to the beginning of the subsequent sentence "to me" of the input sentence is2 into the LSTM blocks lb7 to lb6 of the backward layer lb in the order of arrangement, and uses the output of the LSTM block lb6 as the input to the hidden layer hd.
[0095] Based on the output of the hidden layer obtained by combining the output obtained by inputting the word sequence of the input sentence into the forward layer lf in the order of arrangement and the output obtained by inputting the word sequence of the input sentence into the backward layer lb in the order reverse to the order of arrangement, the model learning unit 13 calculates the output value z i in the output layer sm according to equations (1) and (2).
[0096] The full stop deletion model md, which is a model including the learned neural network, can be regarded as a program that is read or referenced by a computer, causes the computer to execute a predetermined process, and realizes a predetermined function on the computer.
[0097] That is, the learned full stop deletion model md of the present embodiment is used in a computer including a CPU and a memory. Specifically, the CPU of the computer operates to perform an operation based on the learned weighting coefficients (parameters) corresponding to each layer and a response function or the like on the input data input to the input layer of the neural network according to an instruction from the learned full stop deletion model md stored in the memory, and output a result (probability) from the output layer.
[0098] Referring again to FIG. 2, the functional units of the determination device 20 will be described. The input sentence extraction unit 21 extracts a determination input sentence consisting of two consecutive sentences, i.e., a first sentence with a full stop attached at the end of the sentence and a second sentence following immediately after the first sentence, from the text to be determined. Specifically, the input sentence extraction unit 21 may acquire the text to be determined from the text to be determined storage unit 25. The text to be determined storage unit 25 is a storage means for storing the text to be determined.
[0099] The text to be determined is the text to be determined including one or more sentences obtained by speech recognition processing. The text to be determined is, for example, text generated based on uttered speech by a conventional speech recognition algorithm and is provided with full stops. The full stops are, for example, provided in a silent section having a predetermined length or more. Also, the full stops may be provided according to a language model used in the speech recognition algorithm. Therefore, the input sentence extraction unit 21 extracts the full stops from the text including the full stops obtained as a result of speech recognition, sets the sentence with the extracted full stop as the first sentence, sets the sentence following the full stop as the second sentence, and extracts a determination input sentence consisting of the first sentence, the full stop, and the second sentence.
[0100] The determination unit 22 inputs the determination input sentence to the full stop deletion model md and determines whether the full stop attached at the end of the first sentence is correct. As described above, the full stop deletion model md outputs a probability indicating whether the full stop in the input sentence consisting of the first sentence with a full stop attached at the end of the sentence and the second sentence following the full stop is correct.
[0101] Specifically, the determination unit 22 inputs the determination input sentence into the learned full stop deletion model md stored in the full stop deletion model storage unit 17, and determines the probability z that the full stop in the determination input sentence is incorrect 1 and the probability z that the full stop in the determination input sentence is correct 2 and obtains them. When the probability z 2 is greater than the probability z 1 , the determination unit 22 determines that it is incorrect to assign the full stop in the determination input sentence. When the probability z 2 is less than the probability z 1 , the determination unit 22 determines that it is correct to assign the full stop in the determination input sentence.
[0102] The full stop correction unit 23 deletes the full stop determined by the determination unit 22 as being incorrect to be assigned from the determination target text, and outputs the determination target text with the full stop corrected. The output mode of the determination target text is not limited, and it may be stored in a predetermined storage means or displayed on a predetermined display device.
[0103] FIG. 15 is a flowchart showing the processing contents of the full stop deletion model learning method in the full stop deletion model learning device 10.
[0104] In step S1, the first learning data generation unit 11 acquires the first text corpus. In the subsequent step S2, the first learning data generation unit 11 acquires the second text corpus.
[0105] In step S3, the first learning data generation unit 11 extracts an input sentence for generating the first learning data from the first text corpus. The input sentence consists of a preceding sentence, a full stop, and a following sentence.
[0106] In the subsequent step S4, the first learning data generation unit 11 refers to the part of the text corresponding to the input sentence extracted from the first text corpus in the second text corpus. When a period is given at the end of the previous sentence of the referred text, the first learning data generation unit 11 assigns a label indicating that the period given in the input sentence is correct to the input sentence and generates first learning data. On the other hand, when no period is given at the end of the previous sentence of the referred text, the first learning data generation unit 11 assigns a label indicating that the period given in the input sentence is incorrect to the input sentence and generates first learning data.
[0107] In step S5, the model learning unit 13 learns the period deletion model md using the first learning data. Then, the model learning unit 13 stores the learned period deletion model md in the period deletion model storage unit 17.
[0108] FIG. 16 is a flowchart showing the processing content of the determination method using the learned period deletion model md in the determination device 20.
[0109] In step S11, the input sentence extraction unit 21 acquires the determination target text from, for example, the determination target text storage unit 25. In step S12, the input sentence extraction unit 21 extracts the determination input sentence from the determination target text.
[0110] In step S13, the determination unit 22 inputs the determination input sentence to the period deletion model md and determines the correctness of the period.
[0111] In step S14, the period correction unit 23 deletes the period determined to be incorrect in step S13 from the determination target text.
[0112] In step S15, the period correction unit 23 outputs the determination target text with the period corrected.
[0113] Next, with reference to FIG. 17, a period deletion model learning program for causing a computer to function as the period deletion model learning device 10 of the present embodiment and a determination program for causing the computer to function as the determination device 20 will be described.
[0114] FIG. 17(a) is a diagram showing the configuration of the period deletion model learning program. The period deletion model learning program P1A includes a main module m10 that comprehensively controls the period deletion model learning process in the period deletion model learning device 10, a first learning data generation module m11, a second learning data generation module m12, and a model learning module m13. And, by each of the modules m11 to m13, each function for the first learning data generation unit 11, the second learning data generation unit 12, and the model learning unit 13 is realized.
[0115] Note that the period deletion model learning program P1A may be transmitted via a transmission medium such as a communication line, or may be stored in the recording medium M1A as shown in FIG. 17(a).
[0116] FIG. 17(b) is a diagram showing the configuration of the determination program. The determination program P1B includes a main module m20 that comprehensively controls the determination process in the determination device 20, an input sentence extraction module m21, a determination module m22, and a period correction module m23. And, by each of the modules m21 to m23, each function for the input sentence extraction unit 21, the determination unit 22, and the period correction unit 23 is realized.
[0117] Note that the determination program P1B may be transmitted via a transmission medium such as a communication line, or may be stored in the recording medium M1B as shown in FIG. 17(b).
[0118] According to the period deletion model learning device 10, period deletion model learning method, period deletion model md, and period deletion model learning program P1A of the present embodiment described above, since learning of the period deletion model is performed with a two-sentence input consisting of the preamble immediately before the period inserted based on the information obtained by the speech recognition process and the postamble immediately after the period in the first learning data, the information obtained from the speech is reflected in the period deletion model. Also, in the first learning data, a label indicating the correctness of the period included in the input sentence is associated with each input sentence, and learning of the period deletion model is performed based on the error between the probability obtained by inputting the input sentence into the period deletion model and the label, so that the information of the correctly assigned period is reflected in the period deletion model. Therefore, it is possible to obtain a period deletion model that can correctly delete an incorrectly assigned period in the text obtained by the speech recognition process.
[0119] Also, in a period deletion model learning device according to another form, the first learning data generation unit assigns a label to the first learning data based on the presence or absence of a period at the end of the sentence corresponding to the preamble included in the input sentence in a second text corpus that consists of the same text as each first text corpus and has a period added at the end of each sentence included in the text.
[0120] According to the above form, based on the comparison with the second text corpus, which is the same text as the first text corpus and has a period added at the end of each sentence, a label associated with the input sentence extracted from the first text corpus is obtained. Since the second text corpus has a period correctly added at the end of the sentence, first learning data with an appropriate label is generated.
[0121] Also, in a period deletion model learning device according to another form, the first text corpus is text in which a period is inserted at the end of each voice section delimited by a silent section of a predetermined length or more.
[0122] According to the above-described embodiment, since the first learning data is generated based on the first text corpus which is the text with a period inserted in the silent interval in the result of speech recognition, the first learning data reflecting the information obtained by the speech recognition process can be generated.
[0123] Further, in the period deletion model learning device according to another embodiment, based on a third text corpus consisting of texts including sentences properly punctuated at the end, and a fourth text corpus consisting of texts obtained by randomly inserting periods into the texts constituting the third text corpus, second learning data is generated which consists of a pair of an input sentence including the preceding sentence with a period given at the end and the succeeding sentence following immediately after the period in the text constituting the fourth text corpus, and a label indicating whether or not the period is given, and the label of the second learning data is given based on the presence or absence of a period at the end of the sentence corresponding to the preceding sentence included in the input sentence of the second learning data in the third text corpus. The second learning data generation unit is further provided, and the model learning unit updates the parameters of the period deletion model based on the error between the probability obtained by inputting the input sentences of the first learning data and the second learning data into the period deletion model, and the label associated with the input sentence.
[0124] In order to generate the first learning data, in order to associate an appropriate label with the input sentence, a text appropriately punctuated corresponding to the first text corpus is required. For example, the second text corpus is applied to this text. Since the second text corpus is, for example, text input and generated manually, it takes time to obtain a large amount of the second text corpus. According to the above-described embodiment, since the second learning data is generated based on the third text corpus properly punctuated and the fourth text corpus consisting of texts obtained by randomly inserting periods into the texts constituting the third text corpus, and the second learning data can be used for learning the period deletion model together with the first learning data, it is possible to supplement the amount of learning data for use in model generation. As a result, a highly accurate period deletion model can be obtained.
[0125] In addition, in the full stop deletion model learning device according to another form, the full stop deletion model is configured to include a bidirectional long short-term memory network including forward and backward long short-term memory networks respectively. The model learning unit inputs the word sequence included in the previous text and the word sequence included in the subsequent text into the forward long short-term memory network in the order of arrangement in the input text from the word at the beginning of the previous text, and inputs the word sequence included in the previous text and the word sequence included in the subsequent text into the backward long short-term memory network in the reverse order of the arrangement order in the input text from the word at the end of the subsequent text, and obtains a probability based on the output of the forward long short-term memory network and the output of the backward long short-term memory network.
[0126] According to the above form, the tendency of the location where a full stop is given is learned by the full stop deletion model together with the forward and backward tendencies of the word sequences of the previous text immediately before the full stop and the subsequent text immediately after the full stop. Therefore, a full stop deletion model capable of accurately determining the necessity of deleting a full stop can be obtained.
[0127] In addition, in the full stop deletion model learning device according to another form, the model learning unit inputs the word sequence from the beginning to the end of the previous text into the forward long short-term memory network in the order of arrangement in the input text, and inputs the word sequence from the end to the beginning of the subsequent text into the backward long short-term memory network in the reverse order of the arrangement order in the input text.
[0128] According to the above form, since the word sequence from the beginning to the end of the previous text is input into the forward long short-term memory network to perform learning of the full stop deletion model, the end-of-sentence-ness immediately before the full stop is learned by the full stop deletion model. On the other hand, since the word sequence from the end to the beginning of the subsequent text is input into the backward long short-term memory network to perform learning of the full stop deletion model, the beginning-of-sentence-ness immediately after the full stop is learned by the full stop deletion model.
[0129] In order to solve the above problems, a period deletion model according to one embodiment of the present invention is a model for determining the correctness of a period added to text obtained by speech recognition processing, causing a computer to function, and is a learned period deletion model by machine learning. Two consecutive sentences, a first sentence with a period at the end and a second sentence following immediately after the first sentence, are input, and the probability indicating the correctness of the period added at the end of the first sentence is output. Based on a first text corpus that is text in which a period is added based on the information obtained by the speech recognition processing and includes one or more sentences obtained by the speech recognition processing, in the text constituting the first text corpus, an input sentence including a previous sentence that is a sentence with a period added at the end and a subsequent sentence that is a sentence following immediately after the period, and a pair with a label indicating the correctness of the addition of the period are used as first learning data. The parameters of the period deletion model are updated by machine learning based on the error between the probability output when the input sentence included in the first learning data is input to the period deletion model and the label associated with the input sentence.
[0130] In order to solve the above problems, a determination device according to one embodiment of the present invention is a determination device that determines the correctness of a period added to text obtained by speech recognition processing, and is a determination target text including one or more sentences obtained by speech recognition processing. An input sentence extraction unit that extracts a determination input sentence consisting of two consecutive sentences, a first sentence with a period added at the end of the sentence and a second sentence following immediately after the first sentence, from the determination target text; a determination unit that inputs the determination input sentence into a period deletion model and determines the correctness of the period added at the end of the first sentence; and a period correction unit that deletes, from the determination target text, a period determined by the determination unit not to be correctly added, and outputs a determination target text with the period corrected. The period deletion model is a learned model by machine learning for causing a computer to function, takes as input two consecutive sentences, a first sentence with a period added at the end of the sentence and a second sentence following immediately after the first sentence, and outputs as output a probability indicating the correctness of the period added at the end of the first sentence. Based on a first text corpus that is text including one or more sentences obtained by speech recognition processing and has a period added based on the information obtained by the speech recognition processing, a pair of an input sentence including a previous sentence that is a sentence with a period added at the end of the sentence and a subsequent sentence that is a sentence following immediately after the period, in the text constituting the first text corpus, and a label indicating the correctness of the addition of the period is used as first learning data. The period deletion model is constructed by machine learning that updates the parameters of the period deletion model based on the error between the probability output when the input sentence included in the first learning data is input into the period deletion model and the label associated with the input sentence.
[0131] According to the above-described embodiment, a period deletion model is obtained by using, as first learning data, two sentences consisting of the context immediately before a period inserted based on information obtained by voice recognition processing and the context immediately after the period, as an input sentence, and a label indicating whether the period included in the input sentence is correct, and based on the error between the probability obtained by inputting the input sentence into the period deletion model and the label. Therefore, it is possible to obtain a period deletion model that can correctly delete an erroneously added period in the text obtained by voice recognition processing. And in the determination device, the period deletion model thus obtained can be used to appropriately delete periods from the text to be determined.
[0132] As described above in detail for this embodiment, it is obvious to those skilled in the art that this embodiment is not limited to the embodiments described in this specification. This embodiment can be implemented as a modified and changed form without departing from the spirit and scope of the present invention defined by the description of the claims. Therefore, the description in this specification is for the purpose of illustrative explanation and has no restrictive meaning for this embodiment.
[0133] Each aspect / embodiment described in this specification may be applied to a system using LTE (Long Term Evolution), LTE-A (LTE-Advanced), SUPER 3G, IMT-Advanced, 4G, 5G, FRA (Future Radio Access), W-CDMA (registered trademark), GSM (registered trademark), CDMA2000, UMB (Ultra Mobile Broadband), IEEE 802.11 (Wi-Fi), IEEE 802.16 (WiMAX), IEEE 802.20, UWB (Ultra-WideBand), Bluetooth (registered trademark), other appropriate systems, and / or next-generation systems extended based on these.
[0134] The processing procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this specification may be reordered as long as there is no contradiction. For example, for the methods described in this specification, the elements of various steps are presented in an exemplary order and are not limited to the specific order presented.
[0135] The input / output information, etc. may be stored in a specific location (e.g., memory) or may be managed by a management table. The input / output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.
[0136] The determination may be made by a value represented by 1 bit (0 or 1), by a boolean value (Boolean: true or false), or by a numerical comparison (e.g., comparison with a predetermined value).
[0137] Each aspect / embodiment described in this specification may be used alone, in combination, or switched and used during execution. Also, the notification of predetermined information (e.g., notification of "being X") is not limited to being explicitly performed and may be performed implicitly (e.g., without performing the notification of the predetermined information).
[0138] As described above in detail about the present disclosure, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented as modified and changed aspects without departing from the spirit and scope of the present disclosure determined by the description of the claims. Therefore, the description of the present disclosure is for the purpose of illustrative explanation and has no restrictive meaning for the present disclosure.
[0139] Software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether called software, firmware, middleware, microcode, a hardware description language, or by any other name.
[0140] Also, software, instructions, etc. may be transmitted and received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies such as coaxial cables, fiber optic cables, twisted pairs, and digital subscriber lines (DSL) and / or wireless technologies such as infrared, wireless, and microwave, these wired and / or wireless technologies are included within the definition of a transmission medium.
[0141] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc., which may be referred to throughout the above description, may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0142] Note that terms described in this disclosure and / or terms necessary for understanding this specification may be replaced with terms having the same or similar meanings.
[0143] The terms "system" and "network" as used in this specification are used interchangeably.
[0144] Also, the information, parameters, etc. described in this specification may be represented by absolute values, relative values from a predetermined value, or in terms of corresponding other information.
[0145] As used in this disclosure, the terms "determining" and "determination" may encompass a wide variety of operations. "Determining" and "determination" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up (e.g., searching in a table, database, or other data structure), ascertaining, and considering something as having been "determined" or "determined". "Determining" and "determination" may also include considering something as having been "determined" or "determined" after receiving (e.g., receiving information), transmitting (e.g., transmitting information), inputting, outputting, accessing (e.g., accessing data in memory), etc. "Determining" and "determination" may further include considering something as having been "determined" or "determined" after resolving, selecting, choosing, establishing, comparing, etc. That is, "determining" and "determination" may include considering that some operation has been "determined" or "determined". Also, "determining (determination)" may be read as "assuming", "expecting", "considering", etc.
[0146] As used in this disclosure, the phrase "based on" does not mean "based only on" unless otherwise specified. In other words, the phrase "based on" means both "based only on" and "based at least on".
[0147] When terms such as "first", "second", etc. are used in this specification, any reference to those elements does not generally limit the quantity or order of those elements. These terms can be used in this specification as a convenient way to distinguish between two or more elements. Therefore, a reference to the first and second elements does not mean that only two elements can be employed there, or that the first element must precede the second element in any way.
[0148] As long as the terms "include", "including", and their variants are used in this specification or in the claims, these terms are intended to be inclusive, just like the term "comprising". Furthermore, the term "or" used in this specification or in the claims is not intended to be an exclusive disjunction.
[0149] In this specification, unless the context or technology clearly indicates that there is only one device, multiple devices are also included.
[0150] Throughout this disclosure, unless the context clearly indicates a singular number, a plurality is included.
Description of Reference Signs
[0151] 10... Period deletion model learning device, 11... First learning data generation unit, 12... Second learning data generation unit, 13... Model learning unit, 15... First / Second text corpus storage unit, 16... Third / Fourth text corpus storage unit, 17... Period deletion model storage unit, 20... Judgment device, 21... Input sentence extraction unit, 22... Judgment unit, 23... Period correction unit, 25... Judgment target text storage unit, m10... Main module, m11... First learning data generation module, m12... Second learning data generation module, m13... Model learning module, M1A, M1B... Recording medium, m20... Main module, m21... Input sentence extraction module, m22... Judgment module, m23... Period correction module, md... Period deletion model, P1A... Period deletion model learning program, P1B... Judgment program.
Claims
1. A period deletion model learning device that generates, by machine learning, a period deletion model for determining the correctness of a period added to text obtained by speech recognition processing, wherein the period deletion model takes as input two consecutive sentences, a first sentence with a period at the end and a second sentence following immediately after the first sentence, and outputs a probability indicating the correctness of the period added at the end of the first sentence, and the period deletion model learning device generates first learning data consisting of a pair of an input sentence including the previous sentence, which is a sentence with a period added at the end in the text constituting the first text corpus, and the subsequent sentence, which is the sentence following immediately after the period, in the text constituting the first text corpus, which is text including one or more sentences obtained by speech recognition processing and having a period added based on information obtained by the speech recognition processing, and a label indicating the correctness of the addition of the period, based on the first text corpus, and a model learning unit that updates the parameters of the period deletion model based on the error between the probability obtained by inputting the input sentence of the first learning data into the period deletion model and the label associated with the input sentence, wherein the period deletion model learning device is provided with the above components.
2. The first learning data generation unit assigns the label of the first learning data based on the presence or absence of a period at the end of the sentence corresponding to the previous sentence included in the input sentence in a second text corpus, which consists of the same text as each first text corpus and has a period added at the end of the sentence included in the text, The period deletion model learning device according to claim 1.
3. The first text corpus is text in which a period is inserted at the end of each voice section delimited by a silent section of a predetermined length or more, The period deletion model learning device according to claim 1 or 2.
4. Based on a third text corpus consisting of text containing sentences properly punctuated with full stops at the end, and a fourth text corpus consisting of text with full stops randomly inserted into the text constituting the third text corpus, in the text constituting the fourth text corpus, an input sentence including the foregoing text to which the full stop is attached at the end of the sentence and the subsequent text following immediately after the full stop, and a second learning data consisting of a pair with a label indicating whether the full stop is properly attached are generated. The label of the second learning data is assigned based on the presence or absence of a full stop at the end of the sentence corresponding to the foregoing text included in the input sentence of the second learning data in the third text corpus. A second learning data generation unit is further provided. The model learning unit updates the parameters of the full stop deletion model based on the probability obtained by inputting the input sentences of the first learning data and the second learning data into the full stop deletion model, and the error between the probability and the label associated with the input sentence. The full stop deletion model learning device according to any one of claims 1 to 3.
5. The full stop deletion model includes a bidirectional long short-term memory network including forward and backward long short-term memory networks respectively. The model learning unit inputs the word sequence included in the foregoing text and the word sequence included in the subsequent text into the forward long short-term memory network in the order of arrangement in the input sentence from the word at the beginning of the foregoing text. inputs the word sequence included in the foregoing text and the word sequence included in the subsequent text into the backward long short-term memory network in the reverse order of the order of arrangement in the input sentence from the word at the end of the subsequent text. acquires the probability based on the output of the forward long short-term memory network and the output of the backward long short-term memory network. The full stop deletion model learning device according to any one of claims 1 to 4.
6. The full stop deletion model includes a hidden layer that combines the output of the forward long short-term memory network and the output of the backward long short-term memory network, and an output layer that generates the probability based on the output of the hidden layer. The full stop deletion model learning device according to claim 5.
7. The model learning unit inputs the word sequence from the beginning to the end of the foregoing text into the forward long short-term memory network in the order of arrangement in the input sentence. Input the word sequence from the end to the beginning of the subsequent text into the long short-term memory network in the reverse direction according to the reverse order of the array in the input sentence. The period deletion model learning device according to claim 6.
8. A model for determining the correctness of a period assigned to text obtained by speech recognition processing, which causes a computer to function and is a learned period deletion model by machine learning. Input two consecutive sentences, namely the first sentence with a period at the end and the second sentence following immediately after the first sentence, and output the probability indicating the correctness of the period assigned to the end of the first sentence. Based on a first text corpus which is text including one or more sentences obtained by speech recognition processing and having periods assigned based on the information obtained by the speech recognition processing, pairs of an input sentence including a previous sentence which is a sentence with a period assigned at the end and a subsequent sentence which is a sentence following immediately after the period in the text constituting the first text corpus, and a label indicating the correctness of the assignment of the period are used as first training data. It is constructed by machine learning that updates the parameters of the period deletion model based on the error between the probability output when the input sentence included in the first training data is input to the period deletion model and the label associated with the input sentence. Cause a computer to Receive, as input, a determination input sentence consisting of two consecutive sentences, namely the first sentence with a period assigned at the end and the second sentence following immediately after the first sentence, extracted from a determination target text which is the text to be determined including one or more sentences obtained by speech recognition processing. Perform an operation on the input determination input sentence based on at least the parameters constituting the period deletion model. Cause it to function so as to output the probability indicating the correctness of the period assigned to the end of the first sentence based on the result of the operation. A learned period deletion model.
9. A determination device for determining the correctness of a period assigned to text obtained by speech recognition processing. An input sentence extraction unit that extracts a determination input sentence consisting of two consecutive sentences, namely the first sentence with a period assigned at the end and the second sentence following immediately after the first sentence, from a determination target text which is the text to be determined including one or more sentences obtained by speech recognition processing. A determination unit that inputs the determination input sentence into a period deletion model and determines the correctness of the period assigned to the end of the first sentence. A period correction unit that deletes from the text to be determined a period determined by the determination unit not to be correctly added, and outputs the text to be determined with the corrected period; The period deletion model is a learned model by machine learning for causing a computer to function, takes as input two consecutive sentences, a first sentence with a period at the end and a second sentence following immediately after the first sentence, and outputs a probability indicating whether the period given at the end of the first sentence is correct, Based on a first text corpus that is text containing one or more sentences obtained by speech recognition processing and having periods added based on the information obtained by the speech recognition processing, in the text constituting the first text corpus, an input sentence including a preceding sentence that is a sentence with a period added at the end and a succeeding sentence that is a sentence following immediately after the period, and a label indicating whether the addition of the period is correct are used as first learning data, A model constructed by machine learning that updates parameters of the period deletion model based on an error between the probability output when the input sentence included in the first learning data is input to the period deletion model and the label associated with the input sentence. Determination device.
Citation Information
Patent Citations
Voice language processing unit conversion device
JP1999126091A
Mark insertion device and its method
JP2001083987A
Automatic speech question-answer device
JP2003263190A
Text processing device and program
JP2014164575A
Dialogue situation characteristic calculation device, sentence end mark estimation device, method thereof, and program
JP2015219480A
Cited By
Document analysis system, document analysis method, and program
JP2024046883A