Document summarizing apparatus, document summarizing system, and document summarizing method
The document summarization device effectively links emotional information to specific sentences in a summary document, addressing the challenge of associating emotions with speech content.
Patent Information
- Application Number
- JP2024115538
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Existing systems struggle to accurately associate emotional information with specific sentences in a summary document, making it difficult to identify which part of the speech content the emotional information relates to.
A document summarization device that utilizes a memory device and calculation device to create a summary document, identifying correspondences between sentences and emotional information, and attaches emotional information to the summary based on these correspondences.
Enables the creation of a summary document that includes emotional information linked to specific sentences, enhancing the understanding of the speaker's emotions during the conversation.
Smart Images

Figure 2026014460000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a document summarization device, a document summarization system, and a document summarization method for presenting a summary document. [Background technology]
[0002] In welfare consultations such as those providing support for the independence of the needy, consultation records are created that summarize what the client says. In these consultation records, the counselor not only summarizes what the client said, but also reads and records information about the client's emotions at the time of the statement. For example, if a record such as "Things aren't going well with my husband. I'm scared" is created as a consultation record, the information "I'm scared" is information about the client's emotions that the counselor read.
[0003] As a method for analyzing the emotions of such speakers, for example, Patent Document 1 (JP 2022-20499 A) describes a minutes generation device that generates minutes data from character data that has been converted from audio information into text and speaker information, and performs emotion analysis processing on the generated minutes data. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-20499 Summary of the Invention [Problem to be solved by the invention]
[0005] However, when information about the analyzed speaker's emotions is recorded in a document that summarizes the speech content, such as a consultation record, it is difficult to identify which sentence in the summary document the information about the emotions relates to.
[0006] In view of the above problems, an object of the present invention is to identify a correspondence between acquired emotional information and sentences in a summary document, and to add the emotional information to the summary document based on the correspondence. [Means for solving the problem]
[0007] In order to solve the above problems, the present invention provides a document summarization device that presents a document that summarizes the content of a speaker's utterances, the document summarization device having a memory device and a calculation device, wherein the memory device stores emotional information, and the calculation device creates a summary document that summarizes a text including a plurality of sentences that describe the content of the utterances, obtains emotional information from one or more of the plurality of sentences, identifies correspondences between sentences in the summary document and sentences in the text that form the basis of the summary, and presents the summary document with emotional information attached based on the identified correspondences. [Effects of the Invention]
[0008] According to one aspect of the present invention, a document can be created that includes information about the speaker's emotions attached to a summary document of the speech content.
[0009] Problems, configurations, and effects other than those described above will become apparent from the following description of the preferred embodiments of the invention. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating the overall configuration of a document summarization system according to an embodiment of the present invention. [Figure 2] 10 is an example of a basic information table according to an embodiment of the present invention. [Figure 3] 10 is an example of an emotion information master table according to an embodiment of the present invention. [Figure 4] 10 is an example of a situation information master table according to an embodiment of the present invention. [Figure 5] 10 is an example of a language information master table according to an embodiment of the present invention. [Figure 6]10 is an example of a voice information master table according to an embodiment of the present invention. [Figure 7] 10 is an example of a facial expression information master table according to an embodiment of the present invention. [Figure 8] 10 is an example of a biometric information master table according to an embodiment of the present invention. [Figure 9] FIG. 10 is a sequence diagram of creating a consultation record document in an embodiment of the present invention. [Figure 10] 10 is an example of a speech record according to an embodiment of the present invention. [Figure 11] 1 is an example of a summary document according to an embodiment of the present invention. [Figure 12] 10 is an example of analysis of emotion information in an embodiment of the present invention. [Figure 13] 10 is an example of analysis of audio information in an embodiment of the present invention. [Figure 14] 10 is an example of analysis of facial expression information in an embodiment of the present invention. [Figure 15] 10 is an example of analysis of biological information in an embodiment of the present invention. [Figure 16] 10 is an example of caption generation for video data in an embodiment of the present invention. [Figure 17] 1 is a specific example of important emotional information in an embodiment of the present invention. [Figure 18] 1 is a specific example of important context information in an embodiment of the present invention. [Figure 19] 10 is a flowchart of a process for adding important information to a summary document according to an embodiment of the present invention. [Figure 20] 10 shows an example of the result of a process of adding important information to a summary document in an embodiment of the present invention. [Figure 21] 1 is an example of a GUI according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0012] The examples are illustrative of the present invention, and have been omitted or simplified as appropriate for clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.
[0013] In order to facilitate understanding of the invention, the position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings.
[0014] Although various types of information may be described using expressions such as "table" and "list" as examples, the various types of information may also be expressed using data structures other than these. For example, various types of information such as "XX table" and "XX list" may also be expressed as "XX information." When describing identification information, expressions such as "identification information," "identifier," "name," "ID," and "number" are used, but these are interchangeable.
[0015] When there are multiple components with the same or similar functions, they may be described using the same reference numeral with different subscripts. When there is no need to distinguish between these multiple components, the subscripts may be omitted.
[0016] In the embodiments, processing performed by executing a program may be described. Here, a computer executes the program using a processor (e.g., a CPU or a GPU) and performs processing defined by the program using storage resources (e.g., a memory) and interface devices (e.g., a communication port). Therefore, the entity performing the processing by executing the program may be the processor. Similarly, the entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The entity performing the processing by executing the program may be any computing unit, and may include a dedicated circuit that performs specific processing. Here, the dedicated circuit may be, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a CPLD (Complex Programmable Logic Device).
[0017] A program may be installed on a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server may include a processor and a storage resource for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. In addition, in an embodiment, two or more programs may be realized as one program, or one program may be realized as two or more programs. [Example]
[0018] (1) Structure of the document summarization system <Overall system configuration> 1 is an overall configuration diagram of a document summarization system 10, which includes a document summarization device 100, a user client terminal 200, and a client information acquisition device 300. In the document summarization system 10 of this embodiment, for example, in a welfare consultation service provided by a government, a counselor (clerk) at a consultation desk accepts a consultation from a client and creates a document (consultation record document) that describes the content of the conversation, the situation at the time, etc. The document summarization device 100 creates a summary document from a text document containing multiple sentences, acquires emotional information from one or more of the multiple sentences, identifies correspondences between sentences in the summary document and the sentences that form the basis of the acquired emotional information, and presents the emotional information in the summary document based on the identified correspondences.The device has a communication device 101, a central processing unit (CPU) 102, an input / output interface (I / O) 103, a main storage device (memory) 104, and an auxiliary storage device 105.
[0019] The auxiliary storage device 105 has a user portal unit 111, a voice recognition processing unit 112, a facial expression recognition processing unit 113, a biometric information processing unit 114, an emotion information analysis unit 115, a situation information analysis unit 116, a summary document creation unit 117, an attachment position identification unit 118, an importance calculation unit 119, and a caption generation unit 11A. Each of these units is a program, and the CPU 102 executes the program to perform the function of each unit.
[0020] The auxiliary storage device 105 also has databases (DBs) for storing various types of data, including a basic information DB (120), an emotion information DB (121), a situation information DB (122), a language information DB (123), a voice information DB (124), a facial expression information DB (125), and a biometric information DB (126).
[0021] The user client terminal 200 is operated by a counselor and performs overall control of the document summary system 10. It has a communication device 201, a CPU (202), an auxiliary storage device 203, an input / output interface (I / O) 204, and a main storage device (memory) 205, and realizes specified functions by the CPU 202 executing programs stored in the auxiliary storage device 203.
[0022] The user client terminal 200 also has a communication device 201 for communicating with other devices, and is connected to the text summarizing device 100 via a communication network 400. The communication network 400 may be wired / wireless, public line / dedicated line, or may be a direct connection via a communication cable or a wireless connection such as short-range wireless communication. The text summarizing device 100 and the user client terminal 200 may be, for example, a personal computer, and the two may be configured as a single device.
[0023] The client information acquisition device 300 is an external input means for inputting information about the client, and includes, for example, a microphone 301 for acquiring audio information, a camera 302 for acquiring video information, and a wearable device 303 capable of acquiring biometric information such as heart rate, sweat rate, and brain waves, and is connected to the user client terminal 200 via a communication network 400a. The wearable device 303 may be of a commonly used type such as a headset, glasses, shirt, or wristband. The communication network 400a may be wired or wireless, public or dedicated, and may be a direct connection via a communication cable or a wireless connection such as short-range wireless communication.
[0024] Next, the contents of each DB in the auxiliary storage device 105 of the document summarizing device 100 will be described. <Basic information table> 2 is an example of a basic information table 120A stored in the basic information DB (120), which is a table linking the name 120A2 of the speaker (client, counselor) at the time of consultation, the classification 120A3 of whether the speaker is a client or counselor, the gender 120A4 of the speaker, the date of birth 120A5 of the speaker, etc., with a speaker ID (120A1). In this embodiment, it is assumed that a client with a speaker ID (120A1) of "a" visits a consultation desk and a counselor with a speaker ID (120A1) of "b" attends to the consultation, and a consultation record is created regarding the content of the consultation.
[0025] <Emotion Information Master Table> 3 is an example of an emotion information master table 121A stored in the emotion information DB (121), in which emotion information 121A2, which is information about the emotion of the client, is linked to an emotion ID (121A1). For the emotion information, for example, the six basic emotions (joy, sadness, anger, fear, disgust, surprise) proposed by Paul Ekman may be used.
[0026] <Situation Information Master Table> 4 is an example of a situation information master table 122A stored in the situation information DB 122, in which situation information 122A2 relating to the situation of the client is linked to a situation ID 122A1. The situation information 122A2 may include, for example, the state of the client that should be understood in welfare consultation, such as "crying" or "sobbing."
[0027] <Language information master table> 5 is an example of a language information master table 123A stored in the language information DB (123), in which language features 123A3, which are features related to language information used in analyzing emotion information, are specified in association with emotion IDs (123A1) and language information IDs (123A2). As the language features 123A3, for example, a language expression representing "joy" with emotion ID (123A1) of "1" is "thank you, happy," an expression representing "sadness" with emotion ID (123A1) of "2" is "sad, lonely," and an expression representing "anger" with emotion ID (123A1) of "3" is "can't stand it, angry." Furthermore, expressions for "fear" with emotion ID (123A1) "4" may include "scary, worried," expressions for "fun" with emotion ID (123A1) "5" may include "fun, happy," and expressions for "surprise" with emotion ID (123A1) "6" may include "unexpected, unbelievable."
[0028] <Audio information master table> 6 shows an example of a voice information master table 124A stored in the voice information DB 124, in which voice features 124A2, which are features related to voice information used in analyzing emotional information, are defined in association with a voice information ID 124A1. The voice features 124A2 may include, for example, voice volume, fundamental frequency, formant frequency, etc.
[0029] <Audio information master table> 7 shows an example of the facial expression information master table 125A stored in the facial expression information DB 125, in which facial expression feature quantities 125A2, which are features of facial expression information used to analyze emotional information, are linked to facial expression information IDs 125A1. For example, gaze, pupil size, eye movement, etc. may be used as the facial expression feature quantities 125A2.
[0030] <Biometric information master table> 8 shows an example of a biometric information master table 126A stored in the biometric information DB 126, in which biometric feature values 126A2, which are feature values related to biometric information used in analyzing emotion information, are linked to biometric information IDs 126A1. The biometric feature values 126A2 may include, for example, heart rate, brain waves, and sweat rate.
[0031] (2) Sequence for creating consultation records In this embodiment, the process of creating a consultation record document in welfare consultation, in which emotional information and situation information at the time of the consultation with a client are added to a document summarizing the content of the consultation with the client, will be described with reference to the sequence diagram shown in Fig. 9. In this embodiment, a person who has a problem and seeks consultation is called a client, and a person who responds to the consultation from the client and creates the consultation record document is called a counselor, and the counselor is the user of this device.
[0032] <Step S101: Registering basic information> The user uses the user client terminal 200 to input basic speaker information including the speaker ID, name, category, gender, date of birth, etc. of the counselor who handled the consultation and the person being consulted. The input basic speaker information is transmitted to the document summarization device 100 via the network 400. The transmitted basic speaker information is received by the user portal unit 111 in the document summarization device 100, which generates the basic information table 120A shown in FIG. 2 and stores it in the basic information DB (120).
[0033] <Step S102: Obtaining Input Information> While receiving a consultation from a client, the user acquires information (input information) about the client using the user client terminal 200 and the client information acquisition device 300. As input information, the user can acquire audio data of a conversation between the client and the user during the consultation using the microphone 301, which is an external input means, acquire video data of the client during the consultation using the camera 302, or acquire biometric data of the client during the consultation using the wearable device 303. The acquired input information is transmitted to the user client terminal 200 via the network 400a.
[0034] <Step S103: Processing Request> The user client terminal 200 transmits the acquired input information to the document summarization device 100 via the network 400, and requests the start of processing to create a consultation record document. As described above, the input information to be transmitted includes audio data, video data, and biometric data. Upon receiving the input information, the document summarization device 100 performs the processing described below.
[0035] <Step S104: Creating a speech record> The speech recognition processing unit 112 performs speech recognition processing on the speech data included in the input information received from the user client terminal 200, and generates text data by transcribing the speech. At this time, the speaker of the utterance may be identified by the speech recognition processing. The identified speaker may also be associated with a speaker ID (120A1) registered by the user in the basic information table 120A. Alternatively, speech features for each speaker may be registered in advance in the basic information table 120A, and a speaker identified based on these features may be associated with the speaker ID (120A1), thereby performing speaker verification. The speech recognition processing may also be used to correct or complete words in the utterance. FIG. 10 shows an example of a speech record 1000 in which speaker identification is performed on text data transcribed from input speech data, and the speaker (120A1) is identified as either "a: client" or "b: counselor" for each speech content.
[0036] <Step S105: Creating a Summary Document> The summary document creation unit 117 creates a summary document 1100 that summarizes the speech record 1000. For summarization, the speech record may be summarized using, for example, a large-scale language model, or a text summarization service provided by a web service or web API. Furthermore, during summarization, only the speech content of each speaker may be extracted and summarized. FIG. 11 is an example of a summary document 1100 generated from the speech record 1000 shown in FIG. 10. Note that in step S104, text data obtained by transcribing speech data may be summarized using a general document summarization tool or the like.
[0037] <Step S106: Extraction of speech features> The voice recognition processing unit 112 performs voice recognition processing on the voice data included in the received input information, and extracts voice features 124A2 defined in the voice information master table 124A.
[0038] <Step S107: Extraction of Facial Expression Features> Facial expression recognition processing section 113 performs facial expression recognition processing on the video data included in the received input information, and extracts facial expression feature amount 125A2 defined in facial expression information master table 125A.
[0039] <Step S108: Extraction of biometric features The biometric information processing unit 114 performs biometric information processing on the biometric data included in the received input information, and extracts biometric features 126A2 defined in the biometric information master table 126A.
[0040] In the next steps S109 to S112, the emotion information analysis section 115 analyzes the emotion information of the client from a different perspective in each step. <Step S109: Sentiment Analysis of Linguistic Information> The emotion information analysis unit 115 analyzes the emotion information of the client based on the linguistic information, using the utterance record 1000 generated in step S104 and the linguistic information master table 123A stored in the linguistic information DB (123). For example, in the linguistic information master table 123A shown in FIG. 5, the degree of "joy" for which the emotion ID (123A1) is "1" as emotion information may be analyzed as shown in FIG. 12. Specifically, for the content of the utterance for each line number in the utterance record 1000, the similarity between the utterance content and each linguistic feature 123A3 related to "joy" defined in the linguistic information master table 123A is calculated. Then, the similarity calculated for each linguistic feature 123A3 is averaged for each line number in the utterance record, and the average value is used as the degree of "joy" for each line number.
[0041] In the example of Figure 12, the degree of "joy" is analyzed, but other emotional information can also be calculated in a similar manner, such as "sadness" with an emotional ID (123A1) of "2", "anger" with an emotional ID (123A1) of "3", and "fear" with an emotional ID (123A1) of "4".
[0042] <Step S110: Emotion Analysis of Audio Information> The emotion information analysis unit 115 analyzes the emotion information of the client using the utterance record 1000 generated in step S104 and the voice information master table 124A stored in the voice information DB (124). For example, as shown in FIG. 13, in the voice feature extraction process in step S106, the speech time for each line number in the utterance record 1000 may be used as a frame to extract voice features defined in the voice information master table 124A, and the emotions registered in the emotion information master table 121A may be analyzed using a general AI model (such as a voice emotion recognition model). In this case, the estimation accuracy or strength of each emotion may be used as the degree of each emotion.
[0043] <Step S111: Emotion Analysis of Facial Expression Information> The emotion information analysis unit 115 analyzes the emotion information of the client using the utterance record 1000 generated in step S104 and the emotion information master table 125A stored in the emotion information DB (125). For example, as shown in FIG. 14, in the emotion feature extraction process in step S107, the utterance time for each line number in the utterance record 1000 may be used as a frame to extract emotion features defined in the emotion information master table 125A, and the emotions registered in the emotion information master table 121A may be analyzed using a general AI model (e.g., an emotion recognition model). In this case, the estimation accuracy or strength of each emotion may be used as the degree of each emotion.
[0044] <Step S112: Emotion Analysis of Biometric Information> The emotion information analysis unit 115 analyzes the emotion information of the client using the utterance record 1000 generated in step S104 and the biometric information master table 126A stored in the biometric information DB (126). For example, as shown in Fig. 15, in the biometric feature extraction process in step S108, biometric features defined in the biometric information master table 126A may be extracted using the utterance time for each line number in the utterance record 1000 as a frame, and the emotions registered in the emotion information master table 121A may be analyzed using a general AI model (such as a biometric emotion recognition model). In this case, the estimation accuracy or strength of each emotion may be used as the degree of each emotion.
[0045] <Step S113: Caption Generation for Video Data> The caption generation unit 11A generates captions for video data included in input information. For example, as shown in Fig. 16, the utterance time for each line number in the speech record is defined as a frame, and a feature amount is extracted from each frame using a CNN (Convolution Neural Network), and a series of image data for each frame is translated into series of word data using an RNN (Recurrent Neural Network), thereby generating captions.
[0046] <Step S114: Identifying Important Emotional Information> The importance calculation unit 119 identifies important emotional information based on the emotional information analyzed from the language information, voice information, facial expression information, and biometric information in steps S109 to S112. For example, as shown in Fig. 17, the average value of the degree of emotion for each emotion ID calculated for each of the language information, voice information, facial expression information, and biometric information may be calculated for each statement content, and emotional information that exceeds a predetermined threshold may be used as important emotional information. The average value may be weighted by an arbitrary coefficient that is independent for each of the language information, voice information, facial expression information, and biometric information.
[0047] In Figure 17, the simple average of the degree of emotional information analyzed from the linguistic information, voice information, facial expression information, and biometric information for each utterance is used as the average value. In this case, for example, when the threshold value is set to "0.50," the average value of "joy" for line number "19" in the utterance record, "0.81," exceeds the threshold value "0.50," and therefore "joy" is identified as important emotional information for the speaker.
[0048] <Step S115: Identifying Important Situation Information> The importance calculation unit 119 also identifies important situation information based on the captions generated from the video data in step S113. For example, as shown in Fig. 18, for each utterance, the importance calculation unit 119 may calculate the similarity between the generated caption and each piece of situation information defined in the situation information master table 122A, calculate the average value, and use the caption whose average value exceeds a predetermined threshold as important situation information.
[0049] In Figure 18, when the threshold is set to "0.50," the average values of the captions on line numbers "7" and "13" in the speech record, "The woman suddenly looked away" and "The woman suddenly began crying and spoke while sobbing," are "0.81" and "0.89," respectively, which are both above the threshold of "0.50," and therefore these captions are identified as important situational information.
[0050] <Step S116: Adding to the Summary Document> Here, the important emotion information and important situation information extracted in steps S114 and S115 are added to the summary document generated in step S105. A detailed flow of this process is shown in the flowchart of Figure 19. This flowchart will be explained below.
[0051] <<Step ST1: Start>> After step S115 is completed, a series of processes from step ST1 to step ST11 are performed in step S116. <<Step ST2: Specifying line numbers in the summary document>> In the summary document created in step S105, the line number of the sentence in the summary document that identifies the sentence that serves as the basis for each summary sentence (summary basis sentence) is specified. Below, we will explain the processing when line number "n" in the summary document (n is a natural number that is 1 or more and is equal to or less than the maximum line number in the summary document) is specified as the line number of the sentence in the summary document that identifies the basis sentence.
[0052] <<Step ST3: Identifying the basis sentence>> The summary basis sentence of the sentence at line number "n" in the summary document is identified. A large-scale language model may be used to identify the summary basis sentence. For example, the summary document and the speech record may be provided as a prompt to the large-scale language model, and then the large-scale language model may be instructed to "extract which sentence in the speech record is summarized by the sentence at line number n in the summary document," thereby identifying the summary basis sentence from the speech record. There may be multiple identified summary basis sentences.
[0053] <<Step ST4: Specify the line number of the rationale statement>> Among the summary basis sentences identified in step ST3, the line number in the speech record linked to the summary basis sentence is specified as the summary basis sentence to which important emotional information and important situation information are assigned. Below, we will explain the processing when "m" (m is a natural number greater than or equal to 1 and less than or equal to the number of summary basis sentences identified for the sentence with line number n in the summary document) is specified as the line number of the speech record linked to the summary basis sentence.
[0054] <<Step ST5: Check whether important emotional information is included / Step ST6: Add emotional information>> In step ST5, it is confirmed whether important emotional information was identified in step S114 for the summary basis sentence at line number "m" in the speech record. If important emotional information was identified, the process proceeds to step ST6, where the identified important emotional information is added to the sentence at line number "n" in the summary document, and the process proceeds to step ST7. If important emotional information was not identified in step ST5, the process proceeds directly to step ST7.
[0055] <<Step ST7: Check whether important situation information is included / Step ST8: Add situation information>> In step ST7, it is confirmed whether important situation information was identified in step S115 for the summary basis sentence at line number "m" in the speech record. If important situation information was identified, the process proceeds to step ST8, where the identified important situation information is added to the sentence at line number "n" in the summary document, and the process proceeds to step ST9. If important situation information was not identified in step ST7, the process proceeds directly to step ST9.
[0056] <<Step ST9: Check whether there are any unspecified rationale statements>> If the summary basis sentence for the sentence with line number "n" in the summary document identified in step ST3 exists other than the sentence with line number "m" specified in step ST4 ("Yes" in step ST9), the processes from step ST4 to step ST9 are repeated for the sentences with other line numbers. When the processes from step ST4 to step ST9 are completed for all the sentences of the summary basis sentence ("No" in step ST9), the process proceeds to step ST10.
[0057] <<Step ST10: Check whether there are any unspecified summary document line numbers>> If there is a sentence with a line number other than the sentence with line number "n" in the summary document specified in step ST2 ("Yes" in step ST10), steps ST2 to ST10 are executed for all sentences in the summary document.
[0058] Fig. 20 shows an example of the processing results of steps ST1 to ST11. The summary basis sentences for each sentence in the summary document 1100 shown in Fig. 11 are identified from the speech record 1000 shown in Fig. 10, and the important emotional information and important situation information linked to the summary basis sentences shown in Fig. 17 and Fig. 18 are associated with each sentence in the summary document.
[0059] When the processing from step ST2 to step ST10 is completed for all sentences in the summary document ("No" at step ST10), the process proceeds to step ST11, the flow of FIG. 19 ends, and the process proceeds to step S117 in the sequence of FIG.
[0060] <Step S117: Sending the results / Step S118: Displaying the final consultation record document> The document summarization device 100 adds the important emotion information and the important situation information to the summary document and transmits the result (final consultation record document) to the user client terminal 200. The user client terminal 200 displays the result on a display device or the like. Figure 21 shows an example of the GUI (2100).
[0061] In FIG. 21, the GUI (2100) includes a speech record tab 2101 that displays speech records, a summary document tab 2102 that displays a summary document, and a result tab 2103 that displays the results of adding important emotional information and important situational information to the summary document.
[0062] The GUI (2100) also displays important emotional information in association with the sentences in the corresponding summary document. For example, in FIG. 21, the important emotional information "appears happy" 2110 is displayed with a bar 2112 to indicate that it corresponds to sentence 2111 in the summary document (the sentence at line 4 in the summary document in FIG. 20). Similarly, the important situation information "said in a serious manner" 2120 is displayed with a bar 2122 to indicate that it corresponds to sentence 2121 in the summary document (the sentence at line 1 in the summary document in FIG. 20). In this case, when annotating the important situation information "the woman spoke in a serious manner" at line 1 in the summary document in FIG. 20, a large-scale language model is used to modify and simplify the expression, such as by deleting the subject and ending the sentence with a noun, and the "said in a serious manner" is assigned.
[0063] Furthermore, the GUI (2100) may add the important situation information "The woman suddenly looked away," which is the important situation information in line number 2 of the summary document in Figure 20, immediately before or after the summary sentence "My husband drinks alcohol every afternoon, which is interfering with his family life." 2130, and display the result 2131. Here, when adding "The woman suddenly looked away," a sentence such as "She said, 'I drink alcohol every afternoon,' while looking away" is generated and added using a large-scale language model or the like as a quote from the sentence "I drink alcohol every afternoon" in the linked utterance record.
[0064] As described above, according to this embodiment, it is possible to create a summary document of the contents of a conversation to which emotional information of the speaker at the time of the consultation is added.
[0065] It should be noted that in this embodiment, various configurations can be changed, modified, replaced, omitted, etc. to the extent possible. For example, the user client terminal 200 need not be a personal computer, but may be a tablet terminal or the like, or may be integrated with the document summarization device 100.
[0066] Furthermore, although this embodiment has described the creation of summary documents at welfare consultation desks at government agencies, etc., the present invention is not limited to this and can also be applied to, for example, recording conversations at service counters at general stores, consultation services at call centers, and creating minutes during meetings. [Explanation of symbols]
[0067] 10: Document Summarization System 100: Document Summarizer 200: User client terminal 300:Consultant information acquisition device 400, 400a: communication network
Claims
1. A document summarization device that presents a document summarizing the content of a speaker's speech, A storage device and a computing device are included. the storage device stores emotion information; The computing device acquiring emotion information from one or more sentences among the plurality of sentences describing the utterance content; Identifying correspondence between sentences in a summary document that summarizes a text including a plurality of sentences describing the content of the utterance and sentences in the text that are the basis for the summary; based on the identified correspondence, the emotion information is added to the summary document and presented; A document summarization device comprising:
2. 2. The document summarization device according to claim 1, the computing device has a summary document creation unit that creates the summary document; A document summarization device comprising:
3. 3. The document summarization device according to claim 2, the summary document creation unit summarizes the text using a large-scale language model or a text summarization service provided by a web service or web API; A document summarization device comprising:
4. 2. The document summarization device according to claim 1, the storage device stores language information, voice information, facial expression information, and biometric information; the computing device calculates the user's emotional information and its degree from the language information, the voice information, the facial expression information, and the biometric information; A document summarization device comprising:
5. 5. A document summarizing apparatus according to claim 4, The computing device Calculating emotion information and its degree from the similarity between each sentence in the text and the linguistic feature; Calculating emotion information and its degree based on the voice information for each sentence in the text; calculating emotion information and its degree based on the facial expression information for each sentence in the text; calculating emotion information and its degree based on the biometric information for each sentence in the text; Calculate the average value of each emotion level. A document summarization device comprising:
6. 6. A document summarizing apparatus according to claim 5, The computing device identifying, as important emotional information, an average value of the degree of emotional information calculated for each sentence in the text that exceeds a predetermined threshold; A document summarization device comprising:
7. 7. A document summarizing apparatus according to claim 6, The computing device Identifying sentences from the text that serve as a basis for summarizing sentences in the summary document; Correlating the identified important emotion information for each sentence with a sentence in the summary document; A document summarization device comprising:
8. 8. A document summarizing apparatus according to claim 7, The computing device generating a sentence based on the text, the summary document, and the important emotional information, and assigning the important emotional information to the summary document based on the correspondence; A document summarization device comprising:
9. 9. A document summarizing apparatus according to claim 8, the storage device stores situation information indicating the situation of the speaker; A document summarization device comprising:
10. 10. The document summarizing device according to claim 9, The computing device generating a caption explaining the content of the video from video information of the speaker; calculating a similarity with the situation information stored in the storage device; A document summarization device comprising:
11. 11. The document summarizing device according to claim 10, The computing device identifying, as important situation information, information whose similarity between the generated caption and the situation information exceeds a predetermined threshold; A document summarization device comprising:
12. 12. The document summarizing device according to claim 11, The computing device Identifying sentences from the text that serve as a basis for summarizing sentences in the summary document; the important situation information identified for each of the sentences that form the basis for summarization is associated with a sentence in the summary document; A document summarization device comprising:
13. 13. The document summarizing device according to claim 12, The computing device generating a sentence based on the text, the summary document, and the important situation information, and adding the important situation information to the summary document based on the correspondence; A document summarization device comprising:
14. A document summarization device according to any one of claims 1 to 13; a speaker information acquisition device that acquires information about the speaker; a client terminal that transmits information from the speaker information acquisition device to the document summarization device and displays the summary document including the emotion information generated by the document summarization device on a GUI; A document summarization system comprising:
15. A document summarization method for presenting a document summarizing a speaker's speech content with emotional information of the speaker attached thereto, comprising: acquiring emotion information from one or more sentences among the plurality of sentences describing the utterance content; Identifying correspondence between sentences in a summary document that summarizes a text including a plurality of sentences describing the content of the utterance and sentences in the text that are the basis for the summary; based on the identified correspondence, the emotion information is added to the summary document and presented; A document summarization method comprising:
Citation Information
Patent Citations
Minutes generation device, method, computer program, and recording medium
JP2022020499A