Matching device, matching method, and matching program
The verification device and method address the challenge of ensuring document and speech content consistency by extracting and comparing key expressions from speech recognition results and documents, thereby enhancing communication accuracy.
Patent Information
- Application Number
- PCT/JP2023/044659
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Existing techniques struggle to determine if the content of a pre-prepared document matches the speech based on that document, particularly when the speech is not delivered exactly as described.
A verification device and method that includes an acquisition unit for obtaining speech recognition results, an extraction unit for extracting key expressions from both the speech recognition results and the document, and a verification unit that compares these expressions to determine consistency.
Enables effective verification of whether the content of a document and the associated speech are consistent, facilitating accurate communication and documentation.
Smart Images

Figure JP2023044659_19062025_PF_FP_ABST
Abstract
Description
Verification device, verification method, and verification program
[0001] The present disclosure relates to a matching device, a matching method, and a matching program.
[0002] Patent Document 1 describes a technology that refers to attributes represented by words included in the speech recognition results of an inquiry spoken by a user, and identifies an answer from among pre-prepared answers that is associated with the attribute as an answer that matches the inquiry.
[0003] Japanese Patent No. 7334293
[0004] Here, speech may be made based on a document prepared in advance. For example, a doctor may provide an explanation to a patient based on a medical document. However, speech does not always follow the prepared document. For this reason, a technology is needed to determine whether the content of a prepared document and a speech based on the document are consistent. The technology described in Patent Document 1 can identify an answer that matches the speech of the inquiry from among prepared answers, but does not take into consideration oral responses based on the identified answer. For this reason, there is a problem in that it is difficult to determine whether the prepared answer and the content of a speech based on the answer are consistent.
[0005] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology that can determine whether the content of a pre-prepared document and an utterance based on the document are consistent.
[0006] A matching device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring a speech recognition result for an utterance related to a document, an extraction means for extracting key expressions from each of the speech recognition result and the document, and a matching means for matching the speech recognition result and the document based on each of the extracted expressions.
[0007] A matching method according to an exemplary aspect of the present disclosure includes an acquisition process in which at least one processor acquires a speech recognition result for an utterance related to a document; an extraction process in which the at least one processor extracts key expressions from each of the speech recognition result and the document; and a matching process in which the at least one processor matches the speech recognition result and the document based on each of the extracted expressions.
[0008] A matching program according to an exemplary aspect of the present disclosure is a program that causes a computer to function as a matching device, and causes the computer to function as an acquisition means that acquires a speech recognition result for an utterance related to a document, an extraction means that extracts key expressions from each of the speech recognition result and the document, and a matching means that matches the speech recognition result and the document based on each of the extracted expressions.
[0009] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that it is possible to determine whether a document prepared in advance and the content of an utterance based on the document are consistent.
[0010] FIG. 1 is a block diagram showing the configuration of a matching device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of a matching method according to the present disclosure. FIG. 3 is a block diagram showing the configuration of a matching device according to the present disclosure. FIG. 4 is a diagram showing an example of expressions and meanings extracted by an extraction unit in the present disclosure. FIG. 5 is a diagram showing an example of a matching process based on paraphrased expressions in the present disclosure. FIG. 6 is a diagram showing an example of a matching process based on expressions weighted according to meaning in the present disclosure. FIG. 7 is a diagram showing an example of a matching process based on expressions weighted according to topic in the present disclosure. FIG. 8 is a flow diagram showing the flow of a matching method according to the present disclosure. FIG. 9 is a diagram showing an example of a matching result screen according to the present disclosure. FIG. 10 is a block diagram showing the configuration of a matching device according to the present disclosure. FIG. 11 is a block diagram showing an example of the hardware configuration of a computer functioning as each device according to the present disclosure.
[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.
[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.
[0013] (Configuration of the Verification Device) The configuration of the verification device 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the verification device 1. As shown in FIG. 1, the verification device 1 includes an acquisition unit 11, an extraction unit 12, and a verification unit 13. The acquisition unit 11, the extraction unit 12, and the verification unit 13 are an example of a configuration that realizes the acquisition means, the extraction means, and the verification means. For example, if the verification device 1 includes at least one processor, the acquisition unit 11, the extraction unit 12, and the verification unit 13 are realized by the at least one processor executing a program.
[0014] The acquisition unit 11 acquires a speech recognition result for an utterance related to a document. The extraction unit 12 extracts key expressions from each of the speech recognition result and the document. The key expressions may be, for example, important expressions in a field related to the document. The key expressions may be words or phrases. The number of expressions to be extracted may be at least one from each of the speech recognition result and the document, but it is preferable to extract multiple expressions from each of them. As a technique for extracting key expressions, for example, a known technique for extracting named entities from natural language sentences can be applied, but is not limited to this.
[0015] The matching unit 13 matches the speech recognition result with the document based on each extracted expression. For example, matching the speech recognition result with the document may involve identifying a sentence in the document that corresponds to the speech recognition result, or may involve identifying a sentence in the document that does not correspond to the speech recognition result. As an example, the matching unit 13 may determine whether the sentence to be matched corresponds to the speech recognition result based on the similarity between an expression vector obtained by vectorizing an expression extracted from the speech recognition result and an expression vector obtained by vectorizing an expression extracted from the sentence to be matched in the document. However, the details of the process of matching the speech recognition result with the sentence are not limited to this.
[0016] (Effects of the Verification Device) As described above, the verification device 1 employs a configuration including the above-described acquisition unit 11, extraction unit 12, and verification unit 13. Therefore, the verification device 1 has the effect of being able to grasp whether a document prepared in advance and the content of an utterance based on the document are consistent.
[0017] (Flow of the Verification Method) The flow of the verification method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of the verification method S1. As shown in Fig. 2, the verification method S1 includes an acquisition process S11, an extraction process S12, and a verification process S13. The verification method S1 is executed, for example, by the above-mentioned verification device 1 that includes at least one processor.
[0018] In an acquisition process S11, at least one processor (e.g., acquisition unit 11) acquires a speech recognition result for an utterance related to a document. In an extraction process S12, at least one processor (e.g., extraction unit 12) extracts key expressions from each of the speech recognition result and the document. In a matching process S13, at least one processor (e.g., matching unit 13) matches the speech recognition result and the document based on each extracted expression.
[0019] (Effects of the Verification Method) As described above, the verification method S1 employs a configuration including the acquisition process S11, extraction process S12, and verification process S13 described above. Therefore, the verification method S1 can achieve the same effects as the verification device 1 described above.
[0020] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.
[0021] (Configuration of the Verification Device) The configuration of the verification device 1A will be described with reference to FIG. 3. FIG. 3 is a block diagram showing the configuration of the verification device 1A. The verification device 1A includes a control unit 110, a storage unit 120, and a communication unit 130. The control unit 110 controls each unit of the verification device 1A. In addition to the acquisition unit 11, extraction unit 12, and verification unit 13 included in the verification device 1A, the control unit 110 also includes a voice data acquisition unit 14, a correction unit 15, an estimation unit 16, a paraphrase unit 17, and a presentation unit 18. The voice data acquisition unit 14, the correction unit 15, the estimation unit 16, the paraphrase unit 17, and the presentation unit 18 are an example of a configuration that realizes a voice data acquisition means, a correction means, an estimation means, a paraphrase means, and a presentation means. For example, if the verification device 1A has at least one processor, the speech data acquisition unit 14, the correction unit 15, the estimation unit 16, the paraphrase unit 17, and the presentation unit 18 are realized by the at least one processor executing a program.
[0022] The storage unit 120 stores documents prepared in advance for speech. The documents are composed of text data. The documents may be in a state called, for example, a draft, a manuscript, etc., but are not limited thereto. An example of a document is, for example, but is not limited to, a medical document.
[0023] The document is configured to include sentences related to each of a plurality of topics. In the case of a medical document, examples of topics include "regarding the medical condition" and "regarding the examination." The sentences related to each topic may be sentences that form the structural units of the document, such as chapters or sections. As an example, the sentences related to each topic may be, but are not limited to, a chapter describing "regarding the medical condition" or a chapter describing "regarding the examination." The communication unit 130 communicates with each device connected to the verification device 1A.
[0024] As shown in FIG. 3 , the verification device 1A is connected to a storage device 2. The storage device 2 stores an estimation model and an extraction model. For example, the storage device 2 may be connected to the verification device 1A via a network. Such a network may be a wireless local area network (LAN), a wired LAN, a wide area network (WAN), a public line network, a mobile data communication network, or a combination of these networks. Alternatively, the storage device 2 may be connected to the verification device 1A as a peripheral device or may be built into the verification device 1A.
[0025] As shown in FIG. 3 , the verification device 1A is connected to a voice input device 3, an input device 4, and a display device 5. Examples of the voice input device 3 include, but are not limited to, a pin microphone and a headset. Examples of the input device 4 include, but are not limited to, a mouse, a keyboard, and a touchpad. Examples of the display device 5 include, but are not limited to, a liquid crystal display and an organic electroluminescence (EL) display. Some or all of the voice input device 3, the input device 4, and the display device 5 may be connected to the verification device 1A as peripheral devices. Alternatively, some or all of the voice input device 3, the input device 4, and the display device 5 may be built into a terminal (not shown) connected to the verification device 1A via a network, or may be connected as a peripheral device.
[0026] The voice data acquisition unit 14 acquires voice data indicating an utterance. The utterance is an utterance related to a document stored in the storage unit 120. For example, the voice data acquisition unit 14 may acquire voice data for each speech section of a series of utterances, or may acquire voice data indicating a series of utterances. Note that an "utterance section" refers to a unit in a series of utterances that is used to compare the series of utterances with a document, and may be, for example, a unit that can be recognized as being spoken in one breath. For example, each speech section may include an utterance consisting of multiple sentences.
[0027] The acquisition unit 11 is configured similarly to the first exemplary embodiment, and is also configured as follows: The acquisition unit 11 acquires a speech recognition result for speech data for each speech interval. The speech recognition result is represented by text data. For example, when speech data is acquired for each speech interval, the acquisition unit 11 may acquire a speech recognition result for the speech data for each speech interval using a speech recognition model described below. Also, for example, when speech data indicating a series of utterances is acquired, the acquisition unit 11 may perform a process of dividing the speech data into multiple speech intervals and then acquire a speech recognition result for each speech interval using the speech recognition model. Also, the acquisition unit 11 may acquire a speech recognition result for each speech interval by performing a process of dividing the speech recognition result obtained for a series of utterances using the speech recognition model into speech intervals.
[0028] Furthermore, for example, the acquisition unit 11 may acquire a speech recognition result for speech data using a speech recognition model. Such a speech recognition model may be a model that has been machine-learned to output text data as a speech recognition result for input speech data. However, the process of acquiring the speech recognition result is not limited to a method using a speech recognition model, and known speech recognition techniques can be applied. Note that, instead of acquiring the speech recognition result for speech data using a speech recognition model, the acquisition unit 11 may acquire the speech recognition result using an external service that provides speech recognition results for speech data.
[0029] The correction unit 15 corrects recognition errors in the speech recognition result. For example, the correction unit 15 may correct recognition errors in the speech recognition result using a language processing model. The language processing model may be a general-purpose large-scale language model, or a pre-trained language processing model that has been fine-tuned to suit the field related to the document. However, the process of correcting the speech recognition result is not limited to a method using a language processing model, and known text correction techniques can also be applied. Note that the "speech recognition result" in the description of each part below may refer to the speech recognition result corrected by the correction unit 15.
[0030] The estimation unit 16 estimates a topic related to the speech recognition result from among the multiple topics. More specifically, the estimation unit 16 estimates a topic related to the speech recognition result for each speech section. As the multiple topics, multiple topics included in the document stored in the storage unit 120 are applied. For example, the estimation unit 16 may acquire a topic for each speech section using an estimation model stored in the storage device 2. The estimation model is a model that has been machine-learned to classify input text data into one of the multiple topics. However, the process of estimating a topic is not limited to a method using an estimation model.
[0031] The extraction unit 12 is configured similarly to the first exemplary embodiment, but is also configured as follows: The extraction unit 12 acquires the meaning of each extracted expression. For example, the extraction unit 12 extracts key expressions and acquires labels of the extracted expressions as their meanings. The labels represent, for example, the attributes or types of the expressions. For example, the extraction unit 12 may extract key expressions and labels from each of the speech recognition results and documents using an extraction model stored in the storage device 2. The extraction model is a model trained by machine learning to output key expressions and labels in input text data. However, the process of extracting expressions and meanings is not limited to the method using the extraction model.
[0032] An example of an expression and meaning extracted by the extraction unit 12 will be described with reference to FIG. 4. FIG. 4 is a diagram showing an example of an expression and meaning extracted by the extraction unit 12. In the example of FIG. 4, it is assumed that the speech recognition result includes the text "Lung cancer confirmed in the right lower lobe by chest X-ray, notified to the patient." Alternatively, it may be assumed that the text is included in a document. In this case, the extraction unit 12 extracts the key expression "chest X-ray," and acquires the label "examination" of the expression as its meaning. Furthermore, the extraction unit 12 extracts the key expression "lung cancer in the right lower lobe," and acquires the label "disease name" of the expression as its meaning.
[0033] The paraphrasing unit 17 paraphrases at least one of the extracted expressions with another expression. For example, the paraphrasing unit 17 may paraphrase at least one of the expressions extracted from the speech recognition result with another expression, but may not paraphrase the expressions extracted from the document. In this case, for example, a simple expression used in speech can be expected to be paraphrased with a technical expression used in the document. Furthermore, the paraphrasing unit 17 may paraphrase at least one of the expressions extracted from the document with another expression, but may not paraphrase the expressions extracted from the speech recognition result. In this case, for example, a technical expression used in the document can be expected to be paraphrased with a simple expression used in speech. Furthermore, the paraphrasing unit 17 may paraphrase at least one of the expressions extracted from the speech recognition result and at least one of the expressions extracted from the document with another expression.
[0034] Furthermore, for example, the paraphrase unit 17 may convert each extracted expression into another expression by referring to a predetermined paraphrase list. The paraphrase list may be a list of information associating an expression that can be extracted from a document with another expression that paraphrases the expression. In this case, one expression may be associated with multiple other expressions. The paraphrase list may also be generated in advance for expressions in a field related to the document. Furthermore, for example, the paraphrase unit 17 may convert each extracted expression into another expression using a paraphrase model. In this case, the paraphrase model is a model trained by machine learning to output one or more other expressions that paraphrase an input expression. Furthermore, for example, the paraphrase model may be trained by machine learning to be adapted to a field related to the document.
[0035] The matching unit 13 is configured similarly to the first exemplary embodiment, but is also configured as follows. For each utterance interval, the matching unit 13 calculates an expression vector (hereinafter also referred to as the expression vector of the utterance interval) by vectorizing expressions extracted from the speech recognition result of the utterance interval based on the frequency of appearance. Furthermore, for each matching unit of the document, the matching unit calculates an expression vector (hereinafter also referred to as the expression vector of the matching unit) by vectorizing expressions extracted from the matching unit based on the frequency of appearance. The matching unit may be, but is not limited to, a chapter, a section, a paragraph, etc. Furthermore, the matching unit 13 matches the speech recognition result and the document by identifying a correspondence between each utterance interval and each matching unit based on the similarity between the expression vector of each utterance interval and the expression vector of each matching unit. The matching unit 13 may also calculate a score indicating the degree of correspondence between a certain utterance interval and a certain matching unit based on the similarity described above.
[0036] Furthermore, the matching unit 13 matches the speech recognition result with the document further based on the paraphrased expression. For example, the matching unit 13 may match the speech recognition result with the document based on the paraphrased expression instead of or in addition to the expression extracted from the speech recognition result. Note that when the paraphrasing unit 17 paraphrases one expression into multiple expressions, the matching unit 13 generates an expression vector based on the multiple paraphrased expressions.
[0037] An example of a matching process based on paraphrased expressions will be described with reference to FIG. 5. FIG. 5 is a diagram schematically illustrating an example of a matching process based on paraphrased expressions. In FIG. 5, speech recognition results R1, ..., Ri, ... indicate speech recognition results for each utterance section. Document D includes matching units P1, ..., Pj, .... For example, the expressions "chest X-ray" and "stomach cancer" are extracted from speech recognition result Ri, with occurrence frequencies of 1 and 2, respectively. Furthermore, the expressions "chest X-ray," "stomach cancer," and "October 30th" are extracted from matching unit Pj, with occurrence frequencies of 2, 1, and 1, respectively. Furthermore, the expression "chest X-ray" extracted from speech recognition result Ri is rephrased to "chest X-ray" by paraphrase unit 17.
[0038] The matching unit 13 then generates an expression vector based on the occurrence frequency of the extracted expressions or paraphrased expressions for each of the speech recognition result Ri and the matching unit Pj. The expression vector N_Ri of the speech recognition result Ri and the expression vector N_Pj of the matching unit Pj each have, as elements, the occurrence frequencies of a total of three expressions extracted from at least one of the speech recognition result Ri and the matching unit Pj. In this example, three expressions (or paraphrased expressions), "chest X-ray", "stomach cancer", and "October 30th", are extracted from at least one of the speech recognition result Ri and the matching unit Pj. Therefore, the expression vector N_Ri of the speech recognition result is expressed as (1, 2, 0) with the occurrence frequencies of these three expressions as elements. Furthermore, the expression vector N_Pj of the matching unit Pj is expressed as (2, 1, 1) with the occurrence frequencies of these three expressions as elements. For example, the matching unit 13 calculates the similarity between the expression vectors N_Ri and N_Pj, and if the similarity is equal to or greater than a threshold, identifies the matching unit Pj as a sentence that corresponds to the speech section indicated by the speech recognition result Ri in document D. If the similarity is less than the threshold, the matching unit 13 identifies the matching unit Pj as a sentence that does not correspond to the speech section indicated by the speech recognition result Ri in document D. Note that while Fig. 5 shows an example in which matching is performed by referring to the paraphrased expression "chest X-ray" but not the pre-paraphrased expression "chest X-ray," the matching unit 13 may perform matching by referring to both the pre-paraphrased expression and the paraphrased expression.
[0039] The matching unit 13 may also match the speech recognition result and the document based on each expression weighted according to its meaning. For example, the magnitude of the weight according to its meaning may be predetermined or determined by input. An example of matching processing based on expressions weighted according to their meaning will be described with reference to FIG. 6. FIG. 6 is a diagram schematically illustrating an example of matching processing based on expressions weighted according to their meaning. In FIG. 6, elements that are the same as those in FIG. 5 are denoted by the same reference numerals, and detailed description will not be repeated. For example, if weights of "10," "5," and "1" are assigned according to the meanings of the extracted expressions "examination," "disease name," and "time," respectively, the weight vector W (examination, disease name, time) is expressed as (10, 5, 1). Therefore, the expression vector N'_Ri obtained by multiplying the expression vector N_Ri of the speech recognition result Ri by the weight vector W is expressed as (10, 10, 0) as shown in FIG. 6. Furthermore, the representation vector N'_Pj obtained by multiplying the representation vector N_Pj of the matching unit Pj by the weight vector W is expressed as (20, 5, 1) as shown in Figure 6. The matching unit 13 then calculates the similarity between the weighted representation vectors N'_Ri and N'_Pj. Then, depending on whether the similarity is equal to or greater than a threshold, the matching unit 13 determines whether the matching unit Pj corresponds to the speech section indicated by the speech recognition result Ri.
[0040] The matching unit 13 may also weight expressions extracted from the speech recognition results according to the estimated topic, and weight expressions extracted from the document according to the topic related to the sentence containing the expression, thereby matching the speech recognition results and the document based on the weighted expressions. For example, the weight assigned to an expression extracted from the speech recognition results may be the accuracy of the estimated topic. Furthermore, the weight assigned to an expression extracted from the document may be 1 if the topic is related to the sentence containing the expression, and 0 if the topic is not related. Whether the topic is related to the sentence can be determined by referring to configuration information that indicates the configuration of each topic defined for the document containing the sentence. However, if the document is not written in accordance with such configuration information, the sentence contained in the document may be related to multiple topics. In this case, the weight assigned to an expression extracted from the document may be a value corresponding to the likelihood that the expression is a topic related to the sentence containing the expression (e.g., a value closer to 1 indicates a higher likelihood, and a value closer to 0 indicates a lower likelihood).
[0041] An example of matching processing based on expressions weighted according to a topic will be described with reference to FIG. 7 . FIG. 7 is a diagram schematically illustrating an example of matching processing based on expressions weighted according to a topic. In FIG. 7 , elements that are similar to those in FIGS. 5 and 6 are denoted by the same reference numerals, and detailed descriptions will not be repeated. For example, assume that document D includes a matching unit P1, which is a sentence related to the topic "disease name," and a matching unit Pj, which is a sentence related to the topic "surgery description." Note that FIG. 7 illustrates an example in which the sentences related to each topic correspond one-to-one with the matching units, but this is not necessarily limited to a one-to-one correspondence. For example, multiple matching units (e.g., multiple paragraphs) may be related to the same topic (e.g., disease name), or one matching unit (e.g., one chapter) may include multiple topics (e.g., disease name, surgery description).
[0042] As shown in FIG. 7 , it is assumed that the accuracy of the topic related to the speech recognition result Ri is 0.9 for "disease name" and 0.1 for "surgery description." In this case, the weight vector W_Ri assigned to each expression extracted from the speech recognition result Ri is expressed as (0.9, 0.1) based on the accuracy of the topics "disease name" and "surgery description." Therefore, the expression vector N"_Ri obtained by multiplying the expression vector N_Ri of the speech recognition result Ri by the weight vector W_Ri corresponding to the topic is expressed as (0.9, 1.8, 0, 0.1, 0.2, 0) as shown in FIG. 7 . In this example, the expression vector N"_Ri includes six elements obtained by multiplying the appearance frequencies of three expressions by weights for each of the two topics: disease name (chest X-ray), disease name (stomach cancer), disease name (October 30th), surgery description (chest X-ray), surgery description (stomach cancer), and surgery description (October 30th). That is, the expression vector N''_Ri is expressed by concatenating vectors for each topic, each of which is obtained by multiplying each element of the expression vector N_Ri by a weight corresponding to the topic.
[0043] As shown in FIG. 7 , since the matching unit Pj is not related to the topic "disease name" but to the topic "surgery explanation," the weight vector W_Pj assigned to each expression extracted from the matching unit Pj is expressed as (0, 1). Furthermore, the expression vector N"_Pj obtained by multiplying the expression vector N_Pj of the matching unit Pj by the weight vector W_Pj according to the topic is expressed as (0, 0, 0, 2, 1, 1). The six elements of the expression vector N"_Pj are explained in the same way as the expression vector N"_Ri. The matching unit 13 calculates the similarity between the expression vectors N"_Ri and N"_Pj weighted according to the topic. Furthermore, the matching unit 13 determines whether the matching unit Pj corresponds to the speech section indicated by the speech recognition result Ri depending on whether the similarity is equal to or greater than a threshold.
[0044] The matching unit 13 may assign both weights according to meaning and weights according to topic to expressions extracted from each of the speech recognition result and the document. For example, in the examples of FIGS. 6 and 7 , the matching unit 13 may calculate an expression vector N'"_Ri by multiplying the expression vector N'_Ri of the speech recognition result Ri, which has been weighted according to meaning, by a weight vector W_Ri according to the topic. The matching unit 13 may also calculate an expression vector N'"_Pj by multiplying the expression vector N'_Pj of the matching unit Pj, which has been weighted according to meaning, by a weight vector W_Pj according to the topic. The matching unit 13 may then determine whether the matching unit Pj corresponds to the speech section indicated by the speech recognition result Ri by calculating the similarity between the expression vector N'"_Ri and the expression vector N'"_Pj.
[0045] Furthermore, the matching unit 13 generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition results but not extracted from the document. For example, when an expression extracted from the speech recognition results of a certain speech section is not extracted from any of the matching units that make up the document, the matching unit 13 may generate candidates for additional sentences based on the expression extracted from the speech recognition results. Furthermore, for example, the matching unit 13 may generate candidates for additional sentences from the expression using a language processing model. The language processing model may be a general-purpose large-scale language model, or a pre-trained language processing model that has been fine-tuned to suit the field related to the document.
[0046] Furthermore, the matching unit 13 may identify a speech section that includes an expression extracted from the speech recognition result but not extracted from the document. In this case, the presentation unit 18, which will be described later, may highlight and present to the user a speech section that includes a key expression among the speech sections for which a correspondence relationship with the document has not been identified. In this case, even among speech sections for which a correspondence relationship with the document has not been identified, speech sections that do not include a key expression, such as casual conversation, are not highlighted. This allows the user to understand speech sections that are highly important among speech sections for which a correspondence relationship with the document has not been identified. Furthermore, only key expressions may be highlighted in speech sections for which a correspondence relationship with the document has not been identified. This allows the user to understand key expressions in speech sections for which a correspondence relationship with the document has not been identified.
[0047] Furthermore, the matching unit 13 identifies unspoken sentences that are not extracted from the speech recognition result and do not correspond to the speech recognition result in the document based on expressions extracted from the document. For example, the matching unit 13 may identify, as unspoken sentences, matching units that do not correspond to the speech recognition result of any speech section among the matching units that make up the document.
[0048] The presentation unit 18 presents the matching result obtained by the matching unit 13 to the user. For example, the presentation unit 18 presents the matching result to the user by displaying a screen showing the matching result on the display device 5. For example, the matching result screen showing the matching result may include both the speech recognition result and the document, or may include only one of them. Furthermore, the speech recognition results included in the matching result screen may include all of the speech recognition results corresponding to each of the multiple speech segments, or may include only the speech recognition results filtered according to the above-mentioned score calculated by the matching unit 13. Furthermore, the document included in the matching result screen may include all of the matching units constituting the document, or may include only the matching units filtered according to the above-mentioned score calculated by the matching unit 13.
[0049] Furthermore, the matching result screen may be capable of recognizing the correspondence between the speech recognition results of each speech section and each matching unit. For example, when the matching result screen includes both speech recognition results and documents, the speech recognition results and matching units of corresponding speech sections may be connected by a graphic such as a line, or the speech recognition results of corresponding speech sections may be displayed in the same background color or text color. Furthermore, the matching result screen may highlight speech recognition results of speech sections that are not associated with any matching units. Furthermore, the matching result screen may highlight matching units (examples of unspoken sentences) that are not associated with any speech sections. Examples of highlighting methods include, but are not limited to, using a different background color, text color, or text size from the rest of the text, applying text decoration (bold, underline, etc.), or surrounding the text with a frame. Furthermore, the matching result screen may hide the speech recognition results and matching units of the associated speech sections, and display only the speech recognition results and matching units of the unassociated speech sections. The matching result screen may also include candidates for the additional sentences described above. In addition, the expressions extracted by the extraction unit 12 may be displayed in a identifiable manner (for example, highlighted) in the speech recognition results or documents included in the matching result screen.
[0050] The above-described acquisition unit 11, extraction unit 12, matching unit 13, correction unit 15, estimation unit 16, paraphrase unit 17, and presentation unit 18 may function at any point while the speech data acquisition unit 14 is acquiring speech data, based on the speech data acquired up to that point. In other words, the matching device 1A may be executed in real time while speech is continuing. In this case, the matching result screen is updated in real time while speech is continuing.
[0051] (Matching Method S1A) The matching device 1A configured as described above executes matching method S1A. The flow of matching method S1A will be described with reference to FIG. 8. FIG. 8 is a flow diagram showing the flow of matching method S1A. As shown in FIG. 8, matching method S1A includes steps S101 to S112. Note that matching method S1A may be repeatedly executed in real time while utterance related to the document is continuing, or may be executed after the utterance related to the document has ended.
[0052] Step S101 is an example of a voice data acquisition process. In step S101, the voice data acquisition unit 14 acquires voice data indicating speech related to a document. As an example, the document is a medical document in a draft state. Furthermore, speech is made by a doctor to a patient based on the draft. Hereinafter, a medical document in a draft state will also be referred to as a medical document draft, or simply a draft, etc.
[0053] Step S102 is an example of an acquisition process. In step S102, the acquisition unit 11 acquires a speech recognition result for the speech data for each utterance section. For example, the speech recognition result is acquired using the above-described speech recognition model.
[0054] Step S103 is an example of a correction process. In step S103, the correction unit 15 corrects a recognition error in the speech recognition result. For example, the correction of the recognition error is performed using the above-described language processing model.
[0055] In step S104, the estimation unit 16 acquires a topic for each matching unit that constitutes the document. The topic for each matching unit may be acquired by referring to predetermined configuration information of the document. For example, if the document is a medical document, the storage unit 120 may store configuration information that specifies that the medical document includes a chapter on the topic "symptoms," a chapter on "surgery explanation," and a chapter on "risks associated with surgery."
[0056] Step S105 is an example of an estimation process. In step S105, the estimation unit 16 estimates a topic related to the speech recognition result for each utterance section. For example, the topic estimation for each utterance section is performed using the estimation model described above.
[0057] Step S106 is an example of extraction processing. In step S106, the extraction unit 12 extracts key expressions and the meanings of the expressions from the speech recognition result and the document. For example, the extraction of each expression is performed using the extraction model described above.
[0058] Step S107 is an example of a paraphrase process. In step S107, the paraphrase unit 17 paraphrases at least one of the expressions extracted from the speech recognition result and the document into another expression. For example, the paraphrase list or the paraphrase model described above is used for paraphrasing the expression.
[0059] Steps S108 to S111 are an example of matching processing. In step S108, the matching unit 13 calculates an expression vector for each utterance section and an expression vector for each matching unit of the document, each of which has the extracted expressions and / or paraphrased expressions as elements. The matching unit 13 also weights the expression vector for each utterance section and the expression vector for each matching unit of the document according to the meaning and / or the topic. Specific examples of the expression vectors and weightings are as described with reference to FIGS. 5 to 7.
[0060] In step S109, the matching unit 13 matches the speech recognition result with the document by identifying the correspondence between each utterance section and each matching unit based on each weighted expression vector.
[0061] In step S110, the matching unit 13 generates candidates for additional sentences to be added to the document based on expressions that are extracted from the speech recognition results of at least one of the speech sections and that are not extracted from any of the matching units of the document.
[0062] In step S111, the matching unit 13 identifies unspoken sentences in the document that do not correspond to the speech recognition results, based on expressions that have not been extracted from the speech recognition results of any speech section and that have been extracted from at least one matching unit of the document.
[0063] Step S112 is an example of a presentation process. In step S112, the presentation unit 18 displays the matching results obtained in steps S108 to S111 on the display device 5. An example of the matching result screen displayed on the display device 5 in step S112 will be described with reference to FIG. 9.
[0064] FIG. 9 is a diagram illustrating an example of a matching result screen. In FIG. 9 , a doctor (an example of a user) is assumed to be explaining (speaking) to a patient based on a draft of a medical document (an example of a document). As shown in FIG. 9 , the matching result screen G1 includes areas G11 to G13. Area G11 is an area where the speech recognition result for the speech is displayed. Area G12 is an area where the draft of the medical document is displayed. Area G13 is an area where candidates for additional sentences are displayed. Area G11 includes speech recognition results R1, R2, and R3 corresponding to three speech sections. Area G12 includes three matching units P1 (topic "disease name"), P2 (topic "surgery explanation"), and P3 (topic "risks associated with surgery") contained in the document. In the matching result screen G1, the speech recognition result R1 and the matching unit P1 are connected by a line, thereby displaying a recognizable correspondence between them. Similarly, the speech recognition result R3 and the matching unit P2 are also connected by a line, thereby displaying a recognizable correspondence between them. By checking the collation result screen G1, the doctor can confirm that he has explained to the patient the topics of collation units P1 and P2, "disease name" and "surgery explanation."
[0065] Furthermore, since verification unit P3 could not be associated with the speech recognition results of any of the speech sections, it is highlighted and surrounded by a bold frame as an unspoken sentence. The doctor can confirm that among the contents written in the draft of the medical document, verification unit P3 has not been explained (has not been spoken). If verification method S1A is being executed in real time during speech, the doctor will explain (speak) the unspoken verification unit P3 (topic "risks associated with surgery") to the patient. As a result, the verification result screen G1 is updated in real time while the doctor is explaining to the patient, and the highlighting of verification unit P3 is removed. It is also expected that the speech recognition results of the newly added speech section and verification unit P3 will be connected by a line and displayed.
[0066] Furthermore, the speech recognition result R2 is highlighted and surrounded by a bold frame because it was not associated with any matching unit in the document. The candidate additional sentences displayed in area G13 are sentences generated based on expressions extracted from the speech recognition result R2 (i.e., expressions extracted from the speech recognition result but not from the document). A document editing area (not shown) may be displayed by operating object G14 "Edit Document" included in area G13. The document editing area is an area for editing the draft of the medical document displayed in area G12. This allows the doctor to edit the draft based on the candidate additional sentences and complete the medical document.
[0067] 9 has been described assuming that matching method S1A is executed in real time, but matching method S1A does not necessarily have to be executed in real time and may be executed after the doctor has finished explaining to the patient. Also, in the description of FIG. 9 , it has been described assuming that the doctor who made the utterance views the matching result screen G1 and updates the document. However, the user who views the matching result screen G1 and updates the document may be a user other than the speaker (for example, a related person such as the speaker's boss or secretary).
[0068] (Effects of the Verification Device 1A) As described above, the verification device 1A is configured to further include a presentation unit 18 that presents the verification result from the verification unit 13 to the user. Therefore, in addition to the effects of the verification device 1, the verification device 1A provides the effect that the user can easily understand whether the document and the utterance based on the document are consistent by viewing the presented verification result. Furthermore, by viewing the presented verification result, the user can make a decision, for example, on whether to change the content of the utterance or the content of the document. In other words, the verification device 1A can support the user's decision-making.
[0069] Furthermore, the matching device 1A further includes a paraphrasing unit 17 that paraphrases at least one of the extracted expressions into another expression, and the matching unit 13 matches the speech recognition result and the document based on the paraphrased expression. Therefore, in addition to the effects of the matching device 1, the matching device 1A can achieve an effect of being able to perform matching with high accuracy based on whether the meanings are the same between the speech recognition result and the document, regardless of whether the expressions themselves used between them are the same.
[0070] Furthermore, in the matching device 1A, the extraction unit 12 acquires the meaning associated with each extracted expression, and the matching unit 13 matches the speech recognition result with the document based on each expression weighted according to its meaning. Therefore, in addition to the effects of the matching device 1, the matching device 1A can achieve the effect of being able to perform matching with high accuracy by placing emphasis on whether more important expressions are the same between the speech recognition result and the document.
[0071] In addition, in the matching device 1A, the document is configured to include sentences related to each of a plurality of topics, and an estimation unit 16 is further provided for estimating a topic related to the speech recognition result from the plurality of topics, and the matching unit 13 weights the expressions extracted from the speech recognition result according to the estimated topic, and weights the expressions extracted from the document according to the topics related to the sentences containing the expressions, thereby matching the speech recognition result and the document based on each weighted expression. Therefore, in addition to the effects achieved by the matching device 1, the matching device 1A can achieve the effect of being able to perform accurate matching between the speech recognition result and the document by emphasizing what topics the contents relate to.
[0072] Furthermore, in the collation device 1A, the collation unit 13 generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition result but not extracted from the document. Therefore, in addition to the effects of the collation device 1, the collation device 1A can also provide the effect of supporting the document update process, which involves adding missing content to the document in response to speech.
[0073] Furthermore, in the collation device 1A, the collation unit 13 is configured to identify unspoken sentences in the document that are not extracted from the speech recognition result and do not correspond to the speech recognition result, based on expressions extracted from the document. Therefore, in addition to the effects of the collation device 1, the collation device 1A can also provide the effect of being able to check whether the content described in the document has been spoken.
[0074] Furthermore, the verification device 1A is further provided with a correction unit 15 that corrects recognition errors in the speech recognition result. Therefore, in addition to the effects of the verification device 1, the verification device 1A can achieve the effect of being able to more accurately verify the speech recognition result with the document after the recognition errors have been corrected.
[0075] Furthermore, verification device 1A further includes a voice data acquisition unit 14 that acquires voice data representing an utterance, and is configured such that acquisition unit 11, extraction unit 12, and verification unit 13 function at any point in time while voice data acquisition unit 14 is acquiring voice data based on the voice data acquired up to that point in time. Therefore, in addition to the effects of verification device 1, verification device 1A has the effect of allowing the user to grasp the verification result in real time while the speaker is speaking.
[0076] [Variation 1] A verification device 1B, which is a variation of the verification device 1A, will be described with reference to Fig. 10. Fig. 10 is a block diagram showing the configuration of the verification device 1B. As shown in Fig. 10, the verification device 1B is configured such that, in addition to the same configuration as the verification device 1A, the control unit 110 includes a determination unit 19. Below, differences from the verification device 1A will be described, and similar points will not be described repeatedly.
[0077] The acquisition unit 11 is configured similarly to the collation device 1A, and is also configured as follows. The acquisition unit 11 acquires speech recognition results for the speech of the explainer and the speech of the person being explained. The explainer is a person who gives an explanation related to the document. The person being explained is a person who receives an explanation related to the document from the explainer. For example, in an example where the document is a medical document, the explainer may be a doctor, and the person being explained may be a patient or a family member.
[0078] For example, the acquisition unit 11 may estimate the speaker (the speaker or the person being explained) of each part in the voice data to divide the voice data according to the speaker, and acquire a voice recognition result for the voice data for each speaker. The acquisition unit 11 may also estimate the speaker of each part in the voice recognition result for the voice data to divide the voice recognition result according to the speaker.
[0079] The matching unit 13 is configured similarly to the matching device 1A, and is also configured as follows: The matching unit 13 identifies a sentence in a document that corresponds to the utterance of the explainer. For example, the matching unit 13 identifies a correspondence between the utterance of the explainer and each matching unit of the document based on expressions extracted from each speech recognition result for the utterance of the explainer and expressions extracted from each matching unit of the document.
[0080] The determination unit 19 determines whether the recipient agrees to a sentence corresponding to the explainer's utterance based on the speech recognition result of the recipient's utterance. For example, the determination unit 19 may determine whether the recipient agrees based on whether a predetermined expression indicating agreement is included in the speech recognition result of the recipient's utterance. Furthermore, for example, the determination unit 19 may determine whether the recipient agrees using a consent determination model. The consent determination model may be a machine-learned model that outputs whether or not consent is indicated when text data indicating the speech recognition result of the recipient's utterance is input. Furthermore, the determination unit 19 identifies a pair of the speech recognition result of the explainer's utterance and the speech recognition result of the recipient's utterance responding to the explainer's utterance. This allows the determination unit 19 to determine whether or not the recipient agrees to a sentence corresponding to the explainer's utterance in the document.
[0081] The result of determining whether the recipient has consented to the corresponding verification unit may be included in the verification result displayed by the presentation unit 18. This allows the person providing the explanation to update the document by adding whether or not the recipient has consented to sentences in the document that require confirmation of the recipient's consent. In this case, for example, a doctor can complete the medical document by adding whether or not the patient has consented to each chapter in the draft of the medical document.
[0082] (Effects of the Verification Device 1B) As described above, the verification device 1B is configured such that the acquisition unit 11 acquires the speech recognition results for the explainer's utterance and the speech recognition results for the recipient's utterance, the verification unit 13 identifies a sentence in the document that corresponds to the explainer's utterance, and the determination unit 19 determines whether the recipient agrees with the sentence that corresponds to the explainer's utterance based on the speech recognition result for the recipient's utterance. Therefore, in addition to the effects achieved by the verification devices 1 and 1A, the verification device 1B can easily determine whether the recipient agrees with the sentence that corresponds to the explainer's utterance in the document.
[0083] [Variation 2] The matching device 1A or the matching device 1B can be modified as follows. The processing by the paraphrasing unit 17 may be performed before the processing by the extraction unit 12. In this case, the paraphrasing unit 17 rephrases at least one expression included in at least one of the speech recognition result and the document with another expression. The paraphrase list or paraphrase model described above can be used for the processing by the paraphrasing unit 17. Furthermore, the extraction unit 12 extracts key expressions from each of the speech recognition result and the document in which at least one expression in at least one of them has been rephrased with another expression. In this variation, the matching unit 13 can also perform matching based on the paraphrased expression, thereby achieving the same effects as the matching device 1A or the matching device 1B.
[0084] [Application Example] Although the collation devices 1, 1A, and 1B have been described mainly with reference to examples in which the documents are medical documents, the documents are not limited to these. For example, the document may be a press release. In this case, when a public relations officer holds a press conference based on the press release, the collation devices 1, 1A, and 1B can determine whether the press release and the contents of the conference are consistent. Furthermore, for example, the document may be a presentation material. In this case, when a speaker gives a speech based on the presentation material, the collation devices 1, 1A, and 1B can determine whether the presentation material and the speech content are consistent.
[0085] [Example of Software Implementation] Some or all of the functions of the collation devices 1, 1A, and 1B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.
[0086] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 11. Figure 11 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.
[0087] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.
[0088] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0089] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0090] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0091] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0092] (Appendix A1) A matching device comprising: an acquisition means for acquiring a speech recognition result for an utterance related to a document; an extraction means for extracting key expressions from each of the speech recognition result and the document; and a matching means for matching the speech recognition result and the document based on each of the extracted expressions.
[0093] (Appendix A2) The verification device according to Appendix A1, further comprising a presentation unit that presents a verification result obtained by the verification unit to a user.
[0094] (Appendix A3) The matching device according to Appendix A1 or A2, further comprising a paraphrase means for paraphrasing at least one of the extracted expressions into another expression, wherein the matching means matches the speech recognition result and the document based on the paraphrased expression.
[0095] (Appendix A4) The matching device according to Appendix A1 or A2, further comprising a paraphrase means for paraphrasing at least one expression contained in at least one of the speech recognition result and the document into another expression, wherein the extraction means extracts the key expression from each of the speech recognition result and the document in which at least one expression has been paraphrased into the other expression.
[0096] (Appendix A5) The matching device according to any one of Appendices A1 to A4, wherein the extraction means obtains a meaning associated with each of the extracted expressions, and the matching means matches the speech recognition result and the document based on each expression weighted according to the meaning.
[0097] (Appendix A6) The matching device according to any one of Appendices A1 to A5, wherein the document is configured to include sentences related to each of a plurality of topics, and further comprises an estimation means for estimating a topic related to the speech recognition result from the plurality of topics, and the matching means weights expressions extracted from the speech recognition result according to the estimated topic, and weights expressions extracted from the document according to topics related to sentences including the expressions, thereby matching the speech recognition result and the document based on the weighted expressions.
[0098] (Appendix A7) The matching device according to any one of Appendices A1 to A6, wherein the matching means generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition result but not extracted from the document.
[0099] (Appendix A8) The matching device according to any one of Appendices A1 to A7, wherein the matching means identifies an unspoken sentence in the document that is not extracted from the speech recognition result and does not correspond to the speech recognition result, based on an expression extracted from the document.
[0100] (Appendix A9) The matching device according to any one of Appendices A1 to A8, wherein the acquisition means acquires a speech recognition result for the speech of the explainer and a speech recognition result for the speech of the person being explained to; the matching means identifies a sentence in the document that corresponds to the speech of the explainer; and the matching device further comprises a determination means that determines whether the person being explained to agrees with the sentence that corresponds to the speech of the explainer based on the speech recognition result for the speech of the person being explained to.
[0101] (Appendix A10) The matching device according to Appendix A9, wherein the determination means determines whether the person receiving the explanation consents by using a consent determination model that has been machine-learned to output whether the person receiving the explanation consents when a speech recognition result for the person receiving the explanation is input.
[0102] (Supplementary Note A11) The verification device according to any one of Supplementary Notes A1 to A10, further comprising a correction means for correcting a recognition error in the speech recognition result.
[0103] (Appendix A12) The verification device according to any one of Appendices A1 to A11, further comprising a voice data acquisition means for acquiring voice data indicating the utterance, wherein the acquisition means, the extraction means, and the verification means function at any point in time while the voice data acquisition means is acquiring the voice data based on the voice data acquired up to that point in time.
[0104] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0105] (Appendix B1) A matching method comprising: an acquisition process in which at least one processor acquires a speech recognition result for an utterance related to a document; an extraction process in which the at least one processor extracts key expressions from each of the speech recognition result and the document; and a matching process in which the at least one processor matches the speech recognition result and the document based on each of the extracted expressions.
[0106] (Supplementary Note B2) The matching method according to Supplementary Note B1, further comprising a presentation process in which the at least one processor presents a matching result obtained by the matching process to a user.
[0107] (Appendix B3) The matching method described in Appendix B1 or B2, further including a paraphrase process in which the at least one processor paraphrases at least one of the extracted expressions into another expression, and in the matching process, the at least one processor matches the speech recognition result and the document further based on the paraphrased expression.
[0108] (Appendix B4) The matching method described in Appendix B1 or B2, further comprising a paraphrase process in which the at least one processor paraphrases at least one expression contained in at least one of the speech recognition result and the document into another expression, and in the extraction process, the at least one processor extracts the key expression from each of the speech recognition result and the document in which at least one expression has been paraphrased into the other expression.
[0109] (Appendix B5) The matching method according to any one of Appendices B1 to B4, wherein in the extraction process, the at least one processor obtains a meaning associated with each of the extracted expressions, and in the matching process, the at least one processor matches the speech recognition result and the document based on each expression weighted according to the meaning.
[0110] (Appendix B6) The matching method according to any one of Appendices B1 to B5, wherein the document is configured to include sentences related to each of a plurality of topics, and the at least one processor further includes an estimation process for estimating a topic related to the speech recognition result from the plurality of topics, and in the matching process, the at least one processor weights expressions extracted from the speech recognition result according to the estimated topic, and weights expressions extracted from the document according to topics related to sentences including the expressions, thereby matching the speech recognition result and the document based on the weighted expressions.
[0111] (Appendix B7) The matching method according to any one of Appendices B1 to B6, wherein in the matching process, the at least one processor generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition result but not extracted from the document.
[0112] (Appendix B8) The matching method according to any one of Appendices B1 to B7, wherein in the matching process, the at least one processor identifies unspoken sentences in the document that are not extracted from the speech recognition result and do not correspond to the speech recognition result, based on expressions extracted from the document.
[0113] (Appendix B9) The matching method described in any one of Appendices B1 to B8, wherein in the acquisition process, the at least one processor acquires a speech recognition result for the utterance of the explainer and a speech recognition result for the utterance of the person being explained to; in the matching process, the at least one processor identifies a sentence in the document that corresponds to the utterance of the explainer; and the at least one processor further includes a determination process in which the at least one processor determines whether the person being explained to agrees with the sentence that corresponds to the utterance of the explainer based on the speech recognition result for the utterance of the person being explained to.
[0114] (Appendix B10) The matching method described in Appendix B9, wherein in the judgment process, the at least one processor judges whether or not the recipient of the explanation has consented using a consent judgment model that has been machine-learned to output whether or not the recipient of the explanation has consented when a voice recognition result for the recipient's utterance is input.
[0115] (Supplementary Note B11) The verification method according to any one of Supplementary Notes B1 to B10, further comprising a correction process performed by the at least one processor to correct a recognition error in the speech recognition result.
[0116] (Appendix B12) The matching method described in any one of Appendices B1 to B11, further comprising a voice data acquisition process in which the at least one processor acquires voice data indicating the utterance, and the at least one processor performs the acquisition process, the extraction process, and the matching process at any point in time while the voice data is being acquired in the voice data acquisition process based on the voice data obtained up to that point.
[0117] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0118] (Appendix C1) A program that causes a computer to function as a matching device, causing the computer to function as: an acquisition means that acquires a speech recognition result for an utterance related to a document; an extraction means that extracts key expressions from each of the speech recognition result and the document; and a matching means that matches the speech recognition result and the document based on each of the extracted expressions.
[0119] (Appendix C2) The collation program according to Appendix C1, further causing the computer to function as a presentation unit that presents a collation result by the collation unit to a user.
[0120] (Appendix C3) A matching program according to Appendix C1 or C2, further causing the computer to function as a paraphrasing means for paraphrasing at least one of the extracted expressions into another expression, and the matching means for matching the speech recognition result and the document based further on the paraphrased expression.
[0121] (Appendix C4) A matching program as described in Appendix C1 or C2, further causing the computer to function as a paraphrase means for paraphrasing at least one expression contained in at least one of the speech recognition result and the document into another expression, and the extraction means extracts the key expression from each of the speech recognition result and the document in which at least one expression has been paraphrased into the other expression.
[0122] (Appendix C5) The matching program according to any one of Appendices C1 to C4, wherein the extraction means obtains a meaning associated with each of the extracted expressions, and the matching means matches the speech recognition result and the document based on each expression weighted according to the meaning.
[0123] (Appendix C6) A matching program according to any one of Appendices C1 to C5, wherein the document is configured to include sentences related to each of a plurality of topics, and the computer is further made to function as an estimation means for estimating a topic related to the speech recognition result from the plurality of topics, and the matching means weights expressions extracted from the speech recognition result according to the estimated topic, and weights expressions extracted from the document according to topics related to sentences including the expressions, thereby matching the speech recognition result and the document based on the weighted expressions.
[0124] (Appendix C7) The matching program according to any one of Appendices C1 to C6, wherein the matching means generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition result but not extracted from the document.
[0125] (Appendix C8) The matching program according to any one of Appendices C1 to C7, wherein the matching means identifies an unspoken sentence in the document that is not extracted from the speech recognition result and does not correspond to the speech recognition result, based on an expression extracted from the document.
[0126] (Appendix C9) A matching program described in any one of Appendices C1 to C8, wherein the acquisition means acquires speech recognition results for the speech of the explainer and the speech of the person being explained to; the matching means identifies a sentence in the document that corresponds to the speech of the explainer; and the computer further functions as a determination means that determines whether the person being explained to agrees with the sentence that corresponds to the speech of the explainer based on the speech recognition results for the speech of the person being explained to.
[0127] (Appendix C10) The matching program described in Appendix C9, wherein the determination means determines whether the recipient of the explanation has consented or not using a consent determination model that has been machine-learned to output whether the recipient of the explanation has consented or not when a speech recognition result for the recipient's utterance is input.
[0128] (Appendix C11) The collation program according to any one of appendices C1 to C10, further causing the computer to function as a correction means for correcting a recognition error in the speech recognition result.
[0129] (Appendix C12) A matching program described in any one of Appendices C1 to C11, further causing the computer to function as a voice data acquisition means for acquiring voice data indicating the utterance, and causing the acquisition means, extraction means, and matching means to function at any point while the voice data acquisition means is acquiring the voice data based on the voice data obtained up to that point.
[0130] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0131] (Appendix D1) A matching device comprising at least one processor, the at least one processor performing an acquisition process to acquire a speech recognition result for an utterance related to a document, an extraction process to extract key expressions from each of the speech recognition result and the document, and a matching process to match the speech recognition result and the document based on each of the extracted expressions.
[0132] The collation device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.
[0133] (Supplementary Note D2) The verification device according to Supplementary Note D1, wherein the at least one processor further executes a presentation process of presenting a verification result obtained by the verification process to a user.
[0134] (Appendix D3) The matching device described in Appendix D1 or D2, wherein the at least one processor further performs a paraphrase process to rephrase at least one of the extracted expressions into another expression, and in the matching process, the at least one processor matches the speech recognition result and the document further based on the paraphrased expression.
[0135] (Appendix D4) The matching device described in Appendix D1 or D2, wherein the at least one processor further performs a paraphrase process to paraphrase at least one expression included in at least one of the speech recognition result and the document into another expression, and in the extraction process, the at least one processor extracts the key expression from each of the speech recognition result and the document in which at least one expression has been paraphrased into the other expression.
[0136] (Appendix D5) The matching device described in any one of Appendices D1 to D4, wherein in the extraction process, the at least one processor obtains a meaning associated with each of the extracted expressions, and in the matching process, the at least one processor matches the speech recognition result and the document based on each expression weighted according to the meaning.
[0137] (Appendix D6) The matching device according to any one of Appendices D1 to D5, wherein the document is configured to include sentences related to each of a plurality of topics, and the at least one processor further performs an estimation process to estimate a topic from the plurality of topics that is related to the speech recognition result, and in the matching process, the at least one processor weights expressions extracted from the speech recognition result according to the estimated topic, and weights expressions extracted from the document according to topics related to sentences containing the expressions, thereby matching the speech recognition result and the document based on the weighted expressions.
[0138] (Appendix D7) The matching device described in any one of Appendices D1 to D6, wherein in the matching process, the at least one processor generates candidates for additional sentences to be added to the document based on expressions extracted from the speech recognition result but not extracted from the document.
[0139] (Appendix D8) The matching device according to any one of Appendices D1 to D7, wherein in the matching process, the at least one processor identifies unspoken sentences in the document that are not extracted from the speech recognition result and do not correspond to the speech recognition result, based on expressions extracted from the document.
[0140] (Appendix D9) The matching device described in any one of Appendices D1 to D8, wherein in the acquisition process, the at least one processor acquires a speech recognition result for the speech of the explainer and a speech recognition result for the speech of the person being explained to; in the matching process, the at least one processor identifies a sentence in the document that corresponds to the speech of the explainer; and the at least one processor further performs a determination process to determine whether the person being explained to agrees with the sentence that corresponds to the speech of the explainer, based on the speech recognition result for the speech of the person being explained to.
[0141] (Appendix D10) In the judgment process, the at least one processor determines whether the recipient of the explanation has consented or not using a consent judgment model that has been machine-learned to output whether the recipient of the explanation has consented or not when a voice recognition result for the recipient's speech is input. This is the comparison device described in Appendix D9.
[0142] (Supplementary Note D11) The verification device according to any one of Supplementary Notes D1 to D10, wherein the at least one processor further executes a correction process for correcting a recognition error in the speech recognition result.
[0143] (Appendix D12) The matching device described in any one of Appendices D1 to D11, wherein the at least one processor further executes a voice data acquisition process to acquire voice data indicating the utterance, and the acquisition process, the extraction process, and the matching process function at any point while the voice data acquisition process is acquiring the voice data based on the voice data obtained up to that point.
[0144] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0145] (Appendix E1) A non-transitory recording medium having recorded thereon a matching program that causes a computer to function as a matching device, the matching program causing the computer to execute: an acquisition process that acquires speech recognition results for utterances related to a document; an extraction process that extracts key expressions from each of the speech recognition results and the document; and a matching process that matches the speech recognition results and the document based on each of the extracted expressions.
[0146] REFERENCE SIGNS LIST 1, 1A, 1B Collation device 2 Storage device 3 Speech input device 4 Input device 5 Display device 11 Acquisition unit 12 Extraction unit 13 Collation unit 14 Speech data acquisition unit 15 Correction unit 16 Estimation unit 17 Paraphrase unit 18 Presentation unit 19 Determination unit 110 Control unit 120 Storage unit 130 Communication unit C1 Processor C2 Memory
Claims
1. An acquisition unit that acquires a speech recognition result for speech related to a document, an extraction unit that extracts key expressions from each of the speech recognition result and the document, and a collation unit that collates the speech recognition result and the document based on each of the extracted expressions. A collation device comprising the same.
2. The collation device according to claim 1, further comprising a presentation unit that presents the collation result by the collation unit to a user.
3. The collation device according to claim 1 or 2, further comprising a paraphrasing unit that paraphrases at least one of the extracted expressions into another expression, and the collation unit collates the speech recognition result and the document based on the paraphrased expression.
4. The collation device according to claim 1 or 2, further comprising a paraphrasing unit that paraphrases at least one expression included in at least one of the speech recognition result and the document into another expression, and the extraction unit extracts the key expression from each of the speech recognition result and the document in which at least one expression is paraphrased into another expression in at least one of them.
5. The extraction unit acquires the meaning associated with each of the extracted expressions, and the collation unit collates the speech recognition result and the document based on each expression weighted according to the meaning. The collation device according to any one of claims 1 to 4.
6. The document is configured to include articles related to each of a plurality of topics, and further comprises an estimation unit that estimates a topic related to the speech recognition result among the plurality of topics, and the collation unit weights the expressions extracted from the speech recognition result according to the estimated topic, and weights the expressions extracted from the document according to the topic related to the article including the expression, and collates the speech recognition result and the document based on each weighted expression. The collation device according to any one of claims 1 to 5.
7. The collation device according to any one of claims 1 to 6, wherein the collation unit generates a candidate for an additional article to be added to the document based on an expression extracted from the speech recognition result and not extracted from the document.
8. The verification device according to any one of claims 1 to 7, wherein the verification means identifies a non-spoken sentence in the document that does not correspond to the speech recognition result based on an expression extracted from the document and not extracted from the speech recognition result.
9. The acquisition means acquires a speech recognition result for the speaker's utterance and a speech recognition result for the listener's utterance; the verification means identifies a sentence corresponding to the speaker's utterance in the document; and the verification device according to any one of claims 1 to 8 further includes a determination means for determining the presence or absence of the listener's agreement with the sentence corresponding to the speaker's utterance based on the speech recognition result for the listener's utterance.
10. The verification device according to claim 9, wherein the determination means determines the presence or absence of the listener's agreement using an agreement determination model that is machine-learned to output the presence or absence of agreement when the speech recognition result for the listener's utterance is input.
11. The verification device according to any one of claims 1 to 10 further includes a correction means for correcting a recognition error in the speech recognition result.
12. The verification device according to any one of claims 1 to 11 further includes speech data acquisition means for acquiring speech data indicating the utterance, and the acquisition means, the extraction means, and the verification means function based on the speech data obtained up to that point at any point in time when the speech data acquisition means is acquiring the speech data.
13. A verification method including: an acquisition process in which at least one processor acquires a speech recognition result for an utterance related to a document; an extraction process in which the at least one processor extracts key expressions from each of the speech recognition result and the document; and a verification process in which the at least one processor verifies the speech recognition result and the document based on the extracted expressions.
14. A program that causes a computer to function as a collation device, the program causing the computer to function as: an acquisition means for acquiring a speech recognition result for speech related to a document; an extraction means for extracting key expressions from each of the speech recognition result and the document; and a collation means for collating the speech recognition result and the document based on the extracted expressions. A collation program.
Citation Information
Patent Citations
Speech recognition device and speech recognition method therefor
JP2008268571A
Document analysis device and program
JP2014170377A
Document search apparatus, word presentation apparatus, and method and program thereof
JP2019012382A
Information processing device, control method, and program
JP2022059732A
Document creation / reading method, document creation / reading device, document creation / reading robot, and document creation / reading program
WO2005027092A1