Information processing device, information processing method, and program

The information processing apparatus improves speech information extraction accuracy by segmenting and prioritizing speech units based on speaker information and delimiter estimation, addressing the limitations of existing techniques.

WO2025115166A1PCT designated stage expired Publication Date: 2025-06-05NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/042890
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing techniques for segmenting speech at silent intervals struggle to improve the accuracy of information extraction from speech.

Method used

An information processing apparatus and method that acquire text from recognized speech, assign speaker information, estimate delimiters of speech units, and set priorities for each speech unit based on speaker information and delimiter estimation.

Benefits of technology

The proposed solution enhances the accuracy of information extraction from speech by effectively segmenting and prioritizing speech units based on speaker information and delimiter estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023042890_05062025_PF_FP_ABST
    Figure JP2023042890_05062025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device includes: an acquisition means for acquiring text obtained by recognizing speech that includes a vocalization by one or more speakers; an attaching means for attaching, to each vocalization included in the text, information about the speaker who made the vocalization; an estimation means for estimating a break in a constituent unit composed of one or more vocalizations by referring to an expression included in the text and / or the information about the speaker; and a setting means for setting a priority for each vocalization included in the text by referring to the processing results of the attaching means and / or the estimation means.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and program

[0001] The present disclosure relates to an information processing device, an information processing method, and a program.

[0002] A technique is known in which weight information is assigned to each utterance unit generated by dividing an utterance into segments, and a summary of the utterance is generated by assigning weight information to each utterance unit. For example, Patent Literature 1 discloses a technique in which weight information is assigned to each utterance unit generated by dividing an utterance into segments, and a summary of the utterance is generated.

[0003] Japanese Patent Application Publication No. 2020-071675

[0004] The technology described in Patent Document 1 has a problem in that, because speech units are generated by dividing speech into silent sections, it is difficult to improve the accuracy of information extraction from speech.

[0005] The present disclosure has been made in view of the above-mentioned problems, and an exemplary purpose thereof is to provide a technology that can improve the accuracy of information extraction from speech.

[0006] An information processing device according to an exemplary aspect of the present disclosure includes an acquisition means for acquiring text that recognizes speech including utterances by one or more speakers; an assignment means for assigning information about the speaker who made the utterance to each utterance included in the text; an estimation means for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting means for setting a priority for each utterance included in the text by referring to processing results of at least one of the assignment means and the estimation means.

[0007] An information processing method according to an exemplary aspect of the present disclosure includes an acquisition process for acquiring text that recognizes speech including utterances by one or more speakers; an assignment process for assigning information about the speaker who made the utterance to each utterance included in the text; an estimation process for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting process for setting a priority for each utterance included in the text by referring to at least one of the processing results of the assignment process and the estimation process.

[0008] A program according to an exemplary aspect of the present disclosure causes a computer to execute an acquisition process for acquiring text that recognizes speech including utterances by one or more speakers; an assignment process for assigning information about the speaker who made the utterance to each utterance included in the text; an estimation process for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting process for setting a priority for each utterance included in the text by referring to the processing results of at least one of the assignment process and the estimation process.

[0009] According to an exemplary aspect of the present disclosure, an exemplary effect is achieved in that the accuracy of information extraction from utterances can be improved.

[0010] FIG. 1 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 3 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 4 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 5 is a diagram showing an example of processing content according to the present disclosure. FIG. 6 is a block diagram showing a configuration of an information processing device according to the present disclosure. FIG. 7 is a flow diagram showing the flow of an information processing method according to the present disclosure. FIG. 8 is a block diagram showing the configuration of a computer functioning as an information processing device according to the present disclosure.

[0011] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.

[0012] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.

[0013] (Overview of Information Processing Device 1) An overview of the information processing device 1 according to this exemplary embodiment will be described. As an example, the information processing device 1 is a device that sets a priority for each utterance by one or more speakers in order to extract information from the utterances.

[0014] (Configuration of information processing device 1) The configuration of the information processing device 1 will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the information processing device 1. As shown in Fig. 1, the information processing device 1 includes an acquisition unit 11, an assignment unit 12, an estimation unit 13, and a setting unit 14.

[0015] (Acquisition unit 11) The acquisition unit 11 acquires text obtained by recognizing speech including utterances by one or more speakers. For example, the series of utterances may be a dialogue between multiple speakers. Furthermore, when there are multiple speakers, there may be a difference in the amount of knowledge each speaker has about the series of utterances, such as between an expert and a non-expert. Furthermore, when recognizing the speech and converting it into text, a known technology such as a speech recognition engine may be used.

[0016] (Assignment unit 12) The assignment unit 12 assigns, to each utterance included in the text, information about the speaker who made the utterance. The information about the speaker may be, for example, information indicating the attributes of the speaker. For example, if each utterance included in the text is made by speaker A or B, the information about the speaker may include information about whether the speaker is A or B.

[0017] (Estimation unit 13) The estimation unit 13 estimates divisions of constituent units made up of one or more utterances by referring to at least one of expressions included in the text and information about the speaker assigned to each text by the assignment unit 12. The divisions of the constituent units may be, for example, groups of topics in a series of utterances.

[0018] The estimation unit 13 may estimate the division of the structural units by, for example, referring to the meanings of words included in the text. For example, if the text includes a word indicating a question, the estimation unit 13 may estimate the division of the structural units consisting of the question and its answer.

[0019] Furthermore, the estimation unit 13 may be configured to, for example, refer to information about speakers assigned to each text and estimate a point where a speaker switches as a division of the structural unit. For example, if there is a speaker who mainly asks questions and a speaker who mainly answers the questions, the timing of the switch from the answering speaker to the questioning speaker may be estimated as a division of the structural unit. In a similar case, for example, if the speaker who is asking the question repeats the answer of the speaker who is answering the question, and then the speaker who is asking the question asks another question, the timing of the switch from the repetition by the speaker who is asking the question to the other question may be estimated as a division of the structural unit.

[0020] Furthermore, the estimation unit 13 may estimate the division of constituent units by, for example, referring to the degree of association between multiple expressions included in the text. As an example, the estimation unit 13 may estimate the division of constituent units consisting of a range including a series of words having similar meanings in the text. As a specific example, consider a case where the words "date of birth," "birthday," and "hospital" are included in this order in separate utterances in a series of utterances. In this case, since "date of birth" and "birthday" have the same meaning, the estimation unit 13 may, for example, estimate the utterance including "date of birth" and "birthday" as the same constituent unit, and estimate the division of constituent units between the utterance and the utterance including "hospital."

[0021] (Setting Unit 14) The setting unit 14 refers to the processing results of at least one of the assigning unit 12 and the estimating unit 13, and sets a priority to each utterance included in the text.

[0022] (Example of Priority Setting with Reference to Processing Results of Assignment Unit 12) The setting unit 14 may, for example, set a priority for each utterance included in the text with reference to information about the speaker. With respect to a series of utterances by multiple speakers, for example, a speaker with more reliable knowledge is considered to have clearer utterances and higher reliability of the utterance content. In this case, for example, the setting unit 14 may set a higher priority for an utterance by a speaker with more reliable knowledge.

[0023] (Example of Priority Setting Referring to Processing Results of Estimation Unit 13) Here, as an example, consider a case where a series of utterances is composed of questions and answers, and speaker A has more reliable knowledge about the series of utterances. In this case, for example, when speaker A asks a question and speaker B answers it, and speaker A subsequently repeats the answer, it is considered that the question and the repetition contain more valid information than the answer. In this case, the setting unit 14 may, for example, set a higher priority to the utterances related to the question and the repetition than to the answer. On the other hand, for example, when speaker A does not repeat the answer, it is considered that both the question and the answer contain valid information. In this case, the setting unit 14 may, for example, set the same priority to the utterances related to the question and the answer.

[0024] (Effects of Information Processing Device 1) As described above, the information processing device 1 is configured to acquire text obtained by recognizing speech including utterances by one or more speakers, assign information about the speaker who made the utterance to each utterance included in the text, estimate divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speaker, and set priorities for each utterance included in the text by referring to processing results of at least one of the assigning means and the estimating means. Therefore, the information processing device 1 has the effect of improving the accuracy of information extraction from utterances.

[0025] (Flow of Information Processing Method S1) The flow of information processing method S1 will be described with reference to Fig. 2. Fig. 2 is a flow diagram showing the flow of information processing method S1. As shown in Fig. 2, information processing method S1 includes an acquisition process (step) S11, an assignment process (step) S12, an estimation process (step) S13, and a setting process (step) S14.

[0026] (Step S11) In step S11, the acquisition unit 11 acquires text obtained by recognizing speech including speech by one or more speakers. The specific processing by the acquisition unit 11 has been described above, and therefore will not be described here.

[0027] (Step S12) In step S12, the assigning unit 12 assigns, to each utterance included in the text, information about the speaker who made the utterance. The specific processing by the assigning unit 12 has been described above, and therefore will not be described here.

[0028] (Step S13) In step S13, the estimation unit 13 estimates divisions of structural units each consisting of one or more utterances, by referring to at least one of expressions included in the text and information about a speaker assigned to each text by the assignment unit 12. The specific processing by the estimation unit 13 has been described above, and therefore will not be described here.

[0029] (Step S14) In step S14, the setting unit 14 sets a priority for each utterance included in the text by referring to the processing results of at least one of steps S12 and S13. The specific processing by the setting unit 14 has been described above, and therefore will not be described here.

[0030] (Effects of Information Processing Method S1) As described above, information processing method S1 is configured to obtain text obtained by recognizing speech including utterances by one or more speakers, assign information about the speaker to each utterance included in the text, estimate divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speaker, and set priorities for each utterance included in the text by referring to at least one processing result of the assignment process and the estimation process. Therefore, information processing method S1 has the effect of improving the accuracy of information extraction from utterances.

[0031] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0032] (Overview of Information Processing Device 1A) An overview of the information processing device 1A according to this exemplary embodiment will be described. As an example, similar to the information processing device 1, the information processing device 1A is a device that sets priorities for each utterance by one or more speakers in order to extract information from the utterances.

[0033] The following describes an example in which a doctor and a patient communicate about a medical examination through a series of utterances, but this does not limit the present exemplary embodiment. For example, when a doctor creates a medical document based on the details of a medical examination, he or she is required to organize the content and write a concise summary. However, it is difficult to record the doctor-patient interaction on the spot. However, simply automatically converting the audio data of the doctor-patient interaction into text may result in a lot of redundant information, requiring time for editing. Therefore, the information processing device 1A may, for example, set a priority for each utterance included in the audio data of the doctor-patient interaction. Then, the information processing device 1A may, for example, record only utterances whose set priority exceeds a predetermined threshold in the medical document. In this way, the information processing device 1A, for example, sets a priority for each utterance included in the doctor-patient interaction and narrows down the utterances to be recorded, thereby improving the accuracy of extracting information to be recorded from each utterance.

[0034] Specific examples of situations in which the information processing device 1A is used include situations in which information is extracted from a dialogue regarding an examination and recorded in a medical record or the like, situations in which information is extracted from a dialogue regarding consent to a treatment plan and recorded in a consent form or the like, and situations in which information is extracted from a dialogue regarding test results and recorded in an explanatory document or the like for the patient.

[0035] (Configuration of information processing device 1A) The configuration of information processing device 1A will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of information processing device 1A. Information processing device 1A includes a control unit 10, a storage unit 20, and an input / output unit 30.

[0036] (Control unit 10) The control unit 10 controls each unit of the information processing device 1A in an integrated manner. The control unit 10 includes a conversion unit 15 and a text generation unit 16 in addition to the acquisition unit 11, the assignment unit 12, the estimation unit 13, and the setting unit 14 that are included in the information processing device 1A.

[0037] (Acquisition Unit 11) The acquisition unit 11 acquires text T obtained by recognizing speech including utterances by one or more speakers. As described in the exemplary embodiment 1, when there are multiple speakers, there may be a difference in the amount of knowledge each speaker has about a series of utterances, such as between an expert and a non-expert. Specific examples of such speakers include a medical professional such as a doctor and a patient, a reporter belonging to a research institute or a news organization and a general interviewee, etc.

[0038] (Assignment Unit 12) The assignment unit 12 assigns information about the speaker who made each utterance to each utterance included in the text T. As a specific example of the information about the speaker described in the exemplary embodiment 1, if each utterance included in the text T is made by a doctor or a patient, the information about the speaker may include information about whether the speaker is a doctor or a patient.

[0039] (Estimation unit 13) The estimation unit 13 estimates divisions of structural units consisting of one or more utterances by referring to at least one of expressions contained in the text T and information about speakers assigned to each text by the assignment unit 12.

[0040] In addition to the example of the method for the estimation unit 13 to estimate the division of the constituent units described in the exemplary embodiment 1, the estimation unit 13 may, for example, determine the type of utterance and estimate the division of the constituent units by referring to the determination result. The type of utterance may include, for example, information regarding whether or not the answer to the question is subject to the constraints of predetermined options.

[0041] For example, when a series of utterances is composed of a question and answer session, the types of the utterances may include a "closed question," in which the answer to the question is restricted by predetermined options, and an "open question," in which the answer to the question is not restricted by predetermined options. In the case of a "closed question," the answer is generally restricted by options indicated as affirmative or negative, or one of the options in the question. On the other hand, in the case of an "open question," the answer is generally not restricted by any options. As such, since the content of the answer to a question depends on whether it is a "closed question" or an "open question," for example, the estimation unit 13 may determine whether the type of the utterance, which is a question, is a "closed question" or an "open question," and, based on the determination result, estimate the division of the structural unit composed of the question and its answer.

[0042] For example, a "closed question" may include a fact check or an agreement check. Furthermore, the fact check and agreement check may take the form of a tag question, such as "Isn't it...?"

[0043] Furthermore, the estimation unit 13 may estimate the division of the structural units by referring to the degree of association between multiple expressions included in the text T, for example.

[0044] (Setting Unit 14) The setting unit 14 sets a priority for each utterance included in the text T by referring to the processing results of at least one of the assigning unit 12 and the estimating unit 13.

[0045] (Example of Priority Setting with Reference to Processing Results of Assignment Unit 12) The setting unit 14 may, for example, set a priority for each utterance included in the text T with reference to information about the speaker. As described in the exemplary embodiment 1, with respect to a series of utterances by multiple speakers, for example, a speaker with more certain knowledge is considered to have clearer utterances and higher reliability for the content of the utterances. As a specific example, in the aforementioned exchange regarding a medical examination, a doctor with more certain knowledge is considered to have higher reliability for the content of the utterances. In this case, for example, the setting unit 14 may set a higher priority for the utterances of a doctor with more certain knowledge.

[0046] (Example of Priority Setting Referring to Processing Results of the Estimation Unit 13) In the exemplary embodiment 1, as an example, a case was considered in which a series of utterances consisted of a question and answer session, and speaker A had more reliable knowledge of the series of utterances. As a specific example of this case, consider the case of a question and answer session during the aforementioned medical consultation. In this case, speaker A with more reliable knowledge is considered to be a doctor. Furthermore, for example, when a patient answers a question from the doctor and the doctor subsequently repeats the answer, the question and the repetition are considered to contain more useful information than the answer. In this case, the setting unit 14 may, for example, set a higher priority to the utterances related to the question and the repetition than the answer. Specific examples of when a doctor repeats a patient's answer include when the doctor corrects the patient's answer to an inaccurate answer and repeats it with an appropriate expression, or when the doctor repeats the patient's answer because the patient's answer is unclear due to old age or other reasons. On the other hand, for example, when a doctor does not repeat an answer, both the question and the answer are considered to contain useful information. At this time, the setting unit 14 may set the same priority to the utterances relating to the question and the answer, for example.

[0047] (Conversion Unit 15) The conversion unit 15 may, for example, convert an expression included in the text T into an expression specific to a predetermined field. For example, the expression specific to the predetermined field may be a highly specialized expression such as technical terminology in the field.

[0048] As a specific example, in the aforementioned exchange regarding medical examination, it is considered that there is a difference in the amount of knowledge possessed by the doctor and the patient, and therefore, the doctor and the patient may use simple expressions. In this case, the conversion unit 15 may, for example, convert such simple expressions into expressions specific to the medical field, i.e., technical terms such as medical terminology. However, for example, as the expressions used between the doctor and the patient become simpler, there is a possibility that variations in the expressions may occur. In this case, the conversion unit 15 may, for example, use the structured model SM to convert expressions including variations into technical terms.

[0049] (Text Generation Unit 16) The text generation unit 16 may, for example, recognize and convert into text speech including speech by one or more speakers. For example, the text generation unit 16 may use a known technology such as a voice recognition engine.

[0050] (Storage Unit 20) The storage unit 20 stores various types of data referenced by the control unit 10 and various types of data generated by the control unit 10. As an example, the storage unit 20 stores the following: text T; speaker information SP; type information TY; delimiter information D; priority information PR; and structured model SM.

[0051] The text T is, for example, a text obtained by recognizing speech including speech by one or more speakers using a known technology such as a speech recognition engine and converting it into text. For example, the text T may be generated by converting the speech into text in a configuration internal to the information processing device 1A, such as the aforementioned text conversion unit 16. Alternatively, for example, the text T may be generated by converting the speech into text outside the information processing device 1A. For example, the text T may include the recognition result text VT and the processed text PT, which will be described later in the description of FIG. 5 . For example, the text T may further include a unique expression resulting from the conversion unit 15 converting an expression included in the processed text PT into a unique expression in a predetermined field.

[0052] The speaker information SP is, for example, information about the speaker who made the utterance included in the text T. The speaker information SP may be, for example, information indicating the attributes of the speaker. For example, if each utterance included in the text T is made by speaker A or B, the speaker information SP may include information about whether the speaker is A or B. As a specific example, if each utterance included in the text T is made by a doctor or a patient, the speaker information SP may include information about whether the speaker is a doctor or a patient.

[0053] The type information TY is, for example, information about the type of utterance included in the text T. The type information TY may include, for example, information about whether or not the answer to a question is restricted by predetermined options. For example, when a series of utterances is composed of questions and answers, the type information TY may include information about whether the answer to the question is a "closed question" restricted by predetermined options, or an "open question" not restricted by predetermined options.

[0054] The delimiter information D is, for example, information about delimiters of constituent units formed of one or more utterances. For example, the delimiter information D may be information about the result of estimation of the delimiters of the constituent units by the estimation unit 13 as described above.

[0055] The priority information PR is, for example, information regarding the priority set for each utterance included in the text T. For example, the priority information PR may be information regarding the priority set for each utterance with reference to at least one of the speaker information SP and the delimiter information D.

[0056] The structured model SM is a model used, for example, when converting an expression contained in the text T into an expression specific to a predetermined field. For example, if an expression contained in the text T is a simple expression including variations, the expression may be converted into technical terminology using the structured model SM. The structured model SM may be, for example, a model using known technology. Furthermore, for example, the structured model SM may be an internal configuration of the information processing device 1A as shown in FIG. 3, or may be an external configuration connected to the information processing device 1A via a network or the like.

[0057] (Input / Output Unit 30) The input / output unit 30 is configured to include at least one of input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. Alternatively, the input / output unit 30 may be configured to be connected to input / output devices such as a keyboard, a mouse, a display, a printer, and a touch panel. In this configuration, the input / output unit 30 accepts various types of information input to the information processing device 1A from the connected input devices. Furthermore, the input / output unit 30 outputs various types of information to connected output devices under the control of the control unit 10. An example of the input / output unit 30 is an interface such as a USB (Universal Serial Bus).

[0058] For example, the control unit 10 may output data resulting from processing by the setting unit 14 or the conversion unit 15 via the input / output unit 30 to a device external to the information processing device 1A.

[0059] (Flow of information processing method S1A executed by information processing device 1A) The flow of information processing method S1A executed by information processing device 1A will be described with reference to Fig. 4. Fig. 4 is a flow diagram showing the flow of information processing method S1A. In addition to acquisition processing (step) S11, assignment processing (step) S12, estimation processing (step) S13, and setting processing (step) S14 included in information processing method S1, information processing method S1A includes conversion processing (step) S15.

[0060] (Step S11) In step S11, the acquisition unit 11 acquires text obtained by recognizing speech including speech by one or more speakers. The specific processing by the acquisition unit 11 has been described above, and therefore will not be described here.

[0061] (Step S12) In step S12, the assigning unit 12 assigns, to each utterance included in the text, information about the speaker who made the utterance. The specific processing by the assigning unit 12 has been described above, and therefore will not be described here.

[0062] (Step S13) In step S13, the estimation unit 13 estimates divisions of structural units each consisting of one or more utterances, by referring to at least one of expressions included in the text and information about a speaker assigned to each text by the assignment unit 12. The specific processing by the estimation unit 13 has been described above, and therefore will not be described here.

[0063] (Step S14) In step S14, the setting unit 14 sets a priority for each utterance included in the text by referring to the processing results of at least one of steps S12 and S13. The specific processing by the setting unit 14 has been described above, and therefore will not be described here.

[0064] (Step S15) In step S15, the conversion unit 15 may convert, for example, expressions included in the text into expressions specific to a predetermined field. The specific processing by the conversion unit 15 has been described above, and therefore will not be described here.

[0065] (Example of processing contents by information processing device 1A and information processing method S1A) Fig. 5 is a diagram showing an example of processing contents by information processing device 1A and information processing method S1A. The recognition result text VT in the example of Fig. 5 is text obtained by recognizing speech including speech by multiple speakers. In addition, in the example of Fig. 5, processed text PT is generated from the recognition result text VT through processing by the assignment unit 12 and the estimation unit 13.

[0066] (Example of assignment process by assignment unit 12) In the example of Fig. 5, information about the speaker who made the utterance, i.e., utterance information, is assigned to each utterance included in the recognition result text VT by the assignment unit 12. As a result, in the example of the processed text PT, information about whether the utterance was made by speaker A or speaker B is assigned to each utterance. In the example of Fig. 5, speaker A indicates a medical professional such as a doctor, and speaker B indicates a patient.

[0067] (Example of Estimation Processing by Estimation Unit 13) Furthermore, the example of the recognition result text VT is in the form of a question and answer session, and the types of utterances included in the example of the recognition result text VT include either a "closed question (CQ)" or an "open question (OQ)." In the example of FIG. 5 , the estimation unit 13 determines whether the utterance type of the question sentence included in the recognition result text VT, i.e., the type of utterance, is CQ or OQ, and estimates the boundaries of the structural units consisting of the question sentence and its answer, i.e., topic breaks, by referring to the determination result. In the example of the processed text PT, the type of utterance (utterance type: CQ or OQ) is added to each structural unit consisting of a question sentence and its answer in a series of utterances, and horizontal lines are drawn at the boundaries of each structural unit (topic breaks).

[0068] Furthermore, in the example of the recognition result text VT, when the answer to the question OQ is to be accurate or when the answer is incorrect, the answer is repeated. In this case, in the example of the processed text PT, the division of the structural unit including the question to the repetition of the answer is estimated, and such structural unit is marked with "OQ / Repetition." In the example of the processed text PT, in the seventh line from the bottom, speaker A (a medical professional such as a doctor) asks a question about a prescription drug, and in the fourth line from the bottom, speaker B (a patient) answers with the name "CBBBBBD." However, speaker A determines that this answer is a misspelling of "BBBBBB," and in the third line from the bottom, speaker A repeats the correct name as "BBBBBB."

[0069] 5, the setting unit 14 sets a priority for each utterance included in the processed text PT by referring to information about the speaker of each utterance in the processed text PT and the type of utterance in each structural unit. In the example of FIG. 5, a dotted arrow is attached to each utterance in the processed text PT for which a priority higher than a predetermined threshold is set, and information is extracted from each utterance at the end of the dotted arrow. In this example, in each structural unit of the processed text PT, the setting unit 14 sets the same priority to the utterances of both speakers A and B when the utterance type is CQ or OQ, and sets a higher priority to speaker A when the utterance type is OQ / repetition.

[0070] 5, the conversion unit 15 converts the simple expressions including variations contained in the processed text PT into technical terms such as medical terms. Specifically, the conversion unit 15 converts the expressions contained in the example of the processed text PT into the expressions "prescription drugs" and "drinking, smoking."

[0071] (Effects of Information Processing Device 1A) As described above, the information processing device 1A employs a configuration in which the estimation means determines the type of utterance and estimates the division of structural units by referring to the determination result. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can achieve the effect of further improving the accuracy of information extraction from utterances by also referring to the type of utterance.

[0072] Furthermore, the information processing device 1A employs a configuration in which the type of utterance includes information regarding whether or not the answer to the question is restricted by predetermined options. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can further improve the accuracy of information extraction from the utterance by referring to information regarding the question included in the utterance.

[0073] Furthermore, the information processing device 1A employs a configuration in which the estimation means estimates the division of structural units by referring to the degree of association between multiple expressions included in the text. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can obtain an effect of further improving the accuracy of information extraction from an utterance while also referring to the degree of association between expressions included in the utterance.

[0074] Furthermore, the information processing device 1A employs a configuration in which the setting means sets a priority for each utterance included in the text by referring to information about the speaker. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can achieve the effect of further improving the accuracy of information extraction from utterances by also referring to information about the speaker.

[0075] Furthermore, the information processing device 1A is configured to convert expressions contained in text into expressions specific to a predetermined field. Therefore, in addition to the effects of the information processing device 1, the information processing device 1A can also obtain the effect of converting expressions contained in speech into expressions specific to a predetermined field.

[0076] [Third Exemplary Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.

[0077] (Overview of Information Processing Device 1B) An overview of information processing device 1B according to this exemplary embodiment will be described. As an example, similar to information processing device 1A, information processing device 1B is a device that sets priorities for each utterance by one or more speakers in order to extract information from the utterances.

[0078] (Configuration of information processing device 1B) The configuration of information processing device 1B will be described with reference to Fig. 6. Fig. 6 is a block diagram showing the configuration of information processing device 1B. Information processing device 1B includes a communication unit 40 in addition to the control unit 10, storage unit 20, and input / output unit 30 included in information processing device 1A. The control unit 10, storage unit 20, and input / output unit 30 have the same configurations as the control unit 10, storage unit 20, and input / output unit 30 included in information processing device 1A, respectively, and therefore description thereof will be omitted here.

[0079] (Communication Unit 40) The communication unit 40 communicates with devices external to the information processing device 1B. As an example, the communication unit 40 communicates with a language model LM connected to the information processing device 1B via a network N. The communication unit 40 transmits data supplied from the control unit 10 to the outside, and supplies data received from the language model LM to the control unit 10. Note that the specific configuration of the network N does not limit the present exemplary embodiment, and as an example, a wireless LAN (Local Area Network), a wired LAN, a WAN (Wide Area Network), a public line network, a mobile data communication network, or a combination of these networks can be used.

[0080] (Processing by Language Model LM) The language model LM may be, for example, a large language model (LLM) using known technology. For example, the information processing device 1B may transmit data resulting from processing by the setting unit 14 or the conversion unit 15 included in the control unit 10 to the language model LM via the communication unit 40 and the network N. At this time, for example, the language model LM may generate a document by referring to the data resulting from the processing received from the information processing device 1B.

[0081] As a specific example, consider a case where the speech of a conversation in which a doctor explains a treatment plan or the like to a patient is converted into text and the text is processed by the information processing device 1B. In this case, for example, the language model LM may receive data resulting from processing by the setting unit 14 or the conversion unit 15 from the information processing device 1B, and may generate a consent document for the patient regarding the treatment plan or the like by referring to the received data.

[0082] (Flow of information processing method S1B executed by information processing device 1B) The flow of information processing method S1B executed by information processing device 1B will be described with reference to FIG. 7. FIG. 7 is a flow diagram showing the flow of information processing method S1B. Information processing method S1B includes a generation process (step) S16 in addition to an acquisition process (step) S11, an assignment process (step) S12, an estimation process (step) S13, a setting process (step) S14, and a conversion process (step) S15 included in information processing method S1A. Steps S11 to S15 are similar to steps S11 to S15 included in information processing method S1A, and therefore will not be described here.

[0083] (Step S16) In step S16, the language model LM may, for example, acquire data resulting from processing by the information processing device 1B and generate a document by referring to the acquired data. Specific processing by the language model LM has been described above, and therefore will not be described here.

[0084] (Effects of Information Processing Device 1B) As described above, in the information processing device 1B, the language model LM is configured to acquire data resulting from processing by the information processing device 1B and generate a document by referring to the acquired data. Therefore, in addition to the effects of the information processing device 1, the information processing device 1B can also achieve the effect of improving the efficiency of identifying utterances and generating documents.

[0085] [Example of implementation by software] Some or all of the functions of the information processing devices 1, 1A, 1B (hereinafter also referred to as "each of the above devices") may be implemented by hardware such as an integrated circuit (IC chip), or by software.

[0086] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 8. Figure 8 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.

[0087] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.

[0088] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.

[0089] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.

[0090] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.

[0091] [Appendix 1] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0092] (Supplementary Note 1) An information processing device comprising: an acquisition means for acquiring text that recognizes speech including utterances by one or more speakers; an assignment means for assigning information about the speaker who made the utterance to each utterance included in the text; an estimation means for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting means for setting a priority for each utterance included in the text by referring to processing results of at least one of the assignment means and the estimation means.

[0093] (Supplementary Note 2) The information processing device according to Supplementary Note 1, wherein the estimation means determines the type of the utterance and estimates the division of the structural units by referring to the determination result.

[0094] (Supplementary Note 3) The information processing device according to Supplementary Note 2, wherein the type of utterance includes information on whether or not a response to a question is subject to a restriction of predetermined options.

[0095] (Supplementary Note 4) The information processing device according to any one of Supplementary Notes 1 to 3, wherein the estimation means estimates the division of the structural units by referring to the degree of association between a plurality of expressions included in the text.

[0096] (Supplementary Note 5) The information processing device according to any one of Supplementary Notes 1 to 4, wherein the setting means sets a priority for each utterance included in the text by referring to information about the speaker.

[0097] (Supplementary Note 6) The information processing device according to any one of Supplementary Notes 1 to 5, further comprising a conversion means for converting an expression included in the text into an expression specific to a predetermined field.

[0098] (Supplementary Note 7) An information processing method comprising: an acquisition process for acquiring text obtained by recognizing speech including utterances by one or more speakers; an assignment process for assigning information about the speaker who made the utterance to each utterance included in the text; an estimation process for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting process for setting a priority for each utterance included in the text by referring to at least one processing result of the assignment process and the estimation process.

[0099] (Supplementary Note 8) The information processing method according to Supplementary Note 7, wherein the estimation process determines the type of the utterance, and estimates the division of the structural units by referring to the determination result.

[0100] (Supplementary Note 9) The information processing method according to Supplementary Note 8, wherein the type of utterance includes information on whether or not a response to a question is subject to a restriction of predetermined options.

[0101] (Supplementary Note 10) The information processing method according to any one of Supplementary Notes 7 to 9, wherein the estimation process estimates the division of the structural units by referring to the degree of association between multiple expressions included in the text.

[0102] (Supplementary Note 11) The information processing method according to any one of Supplementary Notes 7 to 10, wherein the setting process refers to information about the speaker to set a priority for each utterance included in the text.

[0103] (Supplementary Note 12) The information processing method according to any one of Supplementary Notes 7 to 11, further comprising a conversion process for converting expressions contained in the text into expressions specific to a predetermined field.

[0104] (Supplementary Note 13) A program for causing a computer to operate as the information processing device according to any one of Supplementary Notes 1 to 6, the program causing the computer to function as each of the means.

[0105] [Appendix 2] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.

[0106] (Supplementary Note 1) An information processing device comprising at least one processor, the at least one processor performing the following operations: an acquisition process for acquiring text obtained by recognizing speech including utterances by one or more speakers; an assignment process for assigning information about the speaker who made each utterance to each utterance included in the text; an estimation process for estimating divisions of structural units consisting of one or more utterances by referring to at least one of expressions included in the text and information about the speakers; and a setting process for setting a priority for each utterance included in the text by referring to at least one processing result of the assignment process and the estimation process.

[0107] The information processing device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.

[0108] (Supplementary Note 2) The information processing device according to Supplementary Note 1, wherein in the estimation process, the at least one processor determines a type of the utterance and estimates divisions of the structural units by referring to a result of the determination.

[0109] (Supplementary Note 3) The information processing device according to Supplementary Note 2, wherein the type of utterance includes information on whether or not a response to a question is subject to a restriction of predetermined options.

[0110] (Supplementary Note 4) The information processing device according to any one of Supplementary Notes 1 to 3, wherein in the estimation process, the at least one processor estimates the division of the structural units by referring to the degree of association between multiple expressions included in the text.

[0111] (Supplementary Note 5) In the setting process, the at least one processor sets a priority for each utterance included in the text by referring to information about the speaker. The information processing device according to any one of Supplementary Notes 1 to 4,

[0112] (Supplementary Note 6) The information processing device described in any one of Supplementary Notes 1 to 5, wherein the at least one processor further performs a conversion process to convert expressions contained in the text into expressions specific to a predetermined field.

[0113] 1, 1A, 1B Information processing device 10 Control unit 11 Acquisition unit 12 Assignment unit 13 Estimation unit 14 Setting unit 15 Conversion unit 16 Text generation unit 20 Storage unit 30 Input / output unit 40 Communication unit C1 Processor C2 Memory

Claims

1. An acquisition means for acquiring text that recognizes speech including speech by one or more speakers; an imparting means for imparting information regarding the speaker who made the speech to each speech included in the text; an estimation means for estimating a delimiter of a constituent unit composed of one or more speeches by referring to at least one of the expressions included in the text and the information regarding the speaker; and a setting means for setting a priority for each speech included in the text by referring to a processing result of at least one of the imparting means and the estimation means. An information processing apparatus comprising the above.

2. The estimation means discriminates the type of the speech and estimates the delimiter of the constituent unit by referring to the discriminated result. The information processing apparatus according to claim 1.

3. The type of the speech includes information regarding whether an answer to an interrogative sentence is subject to a constraint of a predetermined option. The information processing apparatus according to claim 2.

4. The estimation means estimates the delimiter of the constituent unit by referring to the degree of association between a plurality of expressions included in the text. The information processing apparatus according to any one of claims 1 to 3.

5. The setting means sets a priority for each speech included in the text by referring to the information regarding the speaker. The information processing apparatus according to any one of claims 1 to 3.

6. The information processing apparatus according to any one of claims 1 to 3, further comprising a conversion means for converting an expression included in the text into an expression unique to a predetermined field.

7. An acquisition process for acquiring text that recognizes speech including speech by one or more speakers; an imparting process for imparting information regarding the speaker who made the speech to each speech included in the text; an estimation process for estimating a delimiter of a constituent unit composed of one or more speeches by referring to at least one of the expressions included in the text and the information regarding the speaker; and a setting process for setting a priority for each speech included in the text by referring to a processing result of at least one of the imparting process and the estimation process. An information processing method including the above.

8. A program that causes a computer to execute: an acquisition process of acquiring text obtained by recognizing speech including speech by one or more speakers; an assignment process of assigning information regarding the speaker who made the speech to each speech included in the text; an estimation process of estimating a delimiter of a constituent unit composed of one or more speeches by referring to at least one of the expressions included in the text and the information regarding the speaker; and a setting process of setting a priority for each speech included in the text by referring to at least one of the processing results of the assignment process and the estimation process.

Citation Information

Patent Citations

  • Method for recognizing topic structure and device therefor

    JP1995160712A

  • Dialogue summarization system and dialogue summarization program

    JP2013120514A