Conference analysis device and program
The conference analysis apparatus addresses the challenge of analyzing meeting speeches by calculating and displaying various analysis values, offering a comprehensive tool for understanding participant engagement and communication dynamics.
Patent Information
- Application Number
- JP2023213389
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-30
AI Technical Summary
Existing technologies lack an efficient method to analyze the speeches of participants in meetings, particularly in converting audio and video communications into text records and providing detailed analysis of speech content.
A conference analysis apparatus and program that analyzes the speeches of participants by calculating predetermined analysis values such as freshness, divergence/convergence, affirmative/negative, sympathy, and filler values, and displays these results in tabular and graphical formats.
Enables comprehensive analysis of meeting speeches, providing insights into participant engagement, content similarity, and emotional reactions, facilitating more effective communication and review processes.
Smart Images

Figure 2025097221000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a meeting analysis apparatus and a program, and more particularly to a meeting analysis apparatus and a program capable of analyzing the speeches of participants in a meeting.
Background Art
[0002] As an apparatus for performing so-called character conversion to create a text record from voices recorded in meetings, discussions, and other communications using an information system, for example, the voice recognition apparatus of Patent Document 1 is known.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] These days, it is widely practiced to conduct communications such as meetings and discussions using an information system. By using an information system, participants in remote locations can participate in the communication, and the communication can be recorded by video and audio.
[0005] When participants communicate with each other or when they retrospectively confirm the content of the communication, they can efficiently proceed with the communication or efficiently review it by simultaneously viewing the information converted into characters by the prior art, the results of analyzing the content of the communication, and other additional information. On the other hand, technologies for analyzing Japanese and other natural languages have also improved these days, and there has been a demand for a technology that analyzes the information converted into characters of the speech content in a meeting to analyze the speeches of the participants in the meeting.
[0006] In view of the above problems, an object of the present invention is to provide a conference analysis apparatus and a program capable of analyzing the speeches of participants in a conference.
Means for Solving the Problems
[0007] A conference analysis apparatus according to the present invention, which is made to solve the above problems, is a conference analysis apparatus for analyzing the speeches of two or more participants in a conference, and for each of the participants, an analysis unit that calculates a predetermined analysis value based on the speeches in the conference in which the participant participated, and the two or more participants, and an analysis result display unit that displays some or all of the calculated predetermined analysis values for each participant in tabular form, and is characterized by comprising these.
[0008] A conference analysis apparatus according to the present invention, which is made to solve the above problems, is a conference analysis apparatus for analyzing the speeches of two or more participants in two or more conferences, and for each of the two or more participants, an analysis unit that calculates a predetermined analysis value based on the speeches in the conference in which the participant participated, and the two or more participants, and an analysis result display unit that displays some or all of the calculated predetermined analysis values for each participant in tabular form, and is characterized by comprising these.
[0009] In the conference analysis apparatus according to the present invention, the analysis unit may further calculate a predetermined analysis value in the conference based on the predetermined analysis values of the participants who participated in the conference for each of the two or more conferences, and the analysis result display unit may further show some or all of the calculated predetermined analysis values in the conference for the two or more conferences.
[0010] A conference analysis apparatus according to the present invention, which aims to solve the above problems, is a conference analysis apparatus that analyzes the speeches of two or more participants in a conference. For each of the conferences and / or each of the participants, an analysis unit calculates a predetermined analysis value for each of the phases obtained by dividing the speech based on a predetermined criterion or for each of the participants, and an analysis result display unit that displays some or all of the predetermined analysis values in a tabular form. The predetermined analysis value includes one or more of a freshness value indicating how many new words were spoken in each of the phases, a divergence / convergence value indicating the similarity between each of the phases and all the phases, an affirmative / negative value indicating the degree of affirmative and negative reactions of the participants, a sympathy value indicating the degree of reaction of the participants to sympathize with other participants, and a filler value indicating the degree of speech of the participants having no specific meaning.
[0011] In the conference analysis apparatus according to the present invention, the analysis result display unit may display an analysis result graph screen including a graph area in which the predetermined analysis value is displayed by a graph having the predetermined analysis value on the vertical axis and the elapsed time in the conference on the horizontal axis, a video area for reproducing the video of the conference, and a transcription area for displaying the information obtained by textifying the speech in the conference together with the information indicating the participant who made the speech.
[0012] In the conference analysis apparatus according to the present invention, the analysis result display unit may further display, for each phase, an analysis result detailed screen that lists the predetermined analysis value and the topic words extracted from the speech in the phase, together with or selectively with the analysis result graph screen.
[0013] In the conference analysis apparatus according to the present invention, when an arbitrary one of the topic words is selected, the analysis result display unit may highlight the phase in which the selected topic word was spoken in the graph area.
Advantages of the Invention
[0014] According to the configuration of the present invention, for each participant, an analysis value for calculating a predetermined analysis value based on the speech in the meeting in which the participant participated, and an analysis result display unit for displaying some or all of the predetermined analysis values calculated by the analysis value in tabular form are provided. Therefore, it is possible to analyze the speech of the participants in the meeting.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Modes for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a diagram showing the configuration of a conference analysis apparatus 1 according to an example of an embodiment of the present invention. The conference analysis apparatus 1 in an example of the present embodiment is an apparatus that provides a function of executing a conference by participants and a function of analyzing the executed conference. As shown in FIG. 1, it includes a conference execution unit 11, a speech recognition unit 12, an analysis unit 13, an analysis result display unit 14, and a database 15.
[0017] In an example of the present embodiment, the conference is a so-called online conference in which a plurality of participants use the participant terminal 2 described later. However, the target to be analyzed as a conference may be arbitrarily changed. For example, it may be configured to analyze group work, meetings, or other communications in which a plurality of participants gather as a conference.
[0018] The conference execution unit 11 executes a conference by two or more participants and records the video and audio in the conference. The participants in the conference participate in the conference using the participant terminal 2 described later. The participant terminal 2 transmits the video of the participant being photographed and the audio of the participant's speech to the conference analysis apparatus 1. The conference execution unit 11 of the conference analysis apparatus 1 transmits the transmitted video and audio to the participant terminal 2 used by other participants, and provides a function of holding a so-called online conference using the video and audio between two or more participant terminals 2.
[0019] The speech recognition unit 12 performs speech recognition on the speech of each participant in the conference executed by the aforementioned conference execution unit 11 and converts it into text information. The text information obtained by converting the speech is recorded in the database 15 described later, together with the information indicating the participant who made the speech, the information indicating the date and time of the speech, and the information indicating the time taken for the speech.
[0020] The analysis unit 13 calculates a predetermined analysis value based on the statements of the participants in the meeting. In an example of this embodiment, the analysis unit 13 calculates, as the predetermined analysis values, a freshness value indicating how much new language has been spoken, a divergence / convergence value indicating the degree of similarity to all phases, an affirmative / negative value indicating the degree of affirmative and negative reactions, a sympathy value indicating the degree of reactions that empathize with other participants, and a filler value indicating the degree of statements having no meaning.
[0021] The analysis result display unit 14 displays the result of the analysis by the aforementioned analysis unit 13. In an example of this embodiment, the viewer who views the analysis result uses a viewing terminal 3 described later, and the analysis result display unit 14 displays an analysis result screen for displaying the analysis result on a display device 31 provided in the viewing terminal 3.
[0022] The database 15 is a database that records information handled by the meeting analysis device 1. The database 15 in an example of this embodiment is an RDBMS (Relational Database Management System) constructed in the meeting analysis device 1, but how the database 15 is configured may be arbitrarily changed. For example, the meeting analysis device 1 may be configured to record the above information as a file in a predetermined format, using a part of the storage area provided in it as the database 15.
[0023] Note that, in an example of this embodiment, the conference analysis apparatus 1 uses a well-known server computer as its hardware configuration. By loading a program pre-recorded in an HDD (Hard Disk Drive), SSD (Solid State Drive), or other storage device of the server computer into the memory and having the CPU (Central Processing Unit) execute it, the server computer is configured to function as the conference analysis apparatus 1. Note that the hardware configuration of the conference analysis apparatus 1 may be arbitrarily changed, and instead of a server computer, a general desktop personal computer or other computer may be used according to performance requirements and the like. Alternatively, the conference analysis apparatus 1 may be configured using two or more computers. Also, the location where the conference analysis apparatus 1 is installed may be, for example, installed at the location where the participant terminal 2 and / or the browsing terminal 3 described later are installed, or may be installed in a data center or other remote location. Further, the functions provided by the conference analysis apparatus 1 may be configured to be provided to the participant terminal 2 and the browsing terminal 3 via the network 4 as a cloud service.
[0024] The participant terminal 2 is a terminal used by the participants of the conference to be analyzed by the conference analysis apparatus 1. The participant terminal 2 includes a display or other display device 21, a keyboard, a mouse, or other input device 22, a camera or other imaging device 23 for photographing the participants, and a microphone or other audio input device 24 for voice-inputting the participants' speeches. Note that the participant terminal 2 may be configured using a well-known computer. For example, a smartphone, a tablet computer, or other portable terminal possessed by the participant may be used, or a desktop personal computer may be used.
[0025] The viewing terminal 3 is a terminal for viewing the results of the analysis by the conference analysis device 1. The viewing terminal 3 includes a display and other display devices 31, and a mouse, keyboard, and other input devices 32. The viewing terminal 3 may be configured using a well-known computer, for example, a smartphone, tablet computer, or other portable terminal owned by a viewer who views the analysis results, or a stationary personal computer may also be used. Further, when the participant and the viewer are the same person, the participant terminal 2 may be configured to also serve as the viewing terminal 3.
[0026] The network 4 is a computer network that communicably connects the conference analysis device 1, the participant terminal 2, and the viewing terminal 3. The network 4 may use a well-known computer network. For example, the Internet or other wide area network may be used, or a LAN (Local Area Network) may be used, or a VPN (Virtual Private Network) may be configured on a wide area network to serve as the network 4. Further, the network 4 may be a wired communication computer network, a wireless communication computer network, or a computer network that combines wired communication and wireless communication.
[0027] The above is the configuration of the conference analysis device 1 in an example of this embodiment.
[0028] Next, the analysis process in an example of this embodiment will be described. The conference analysis device 1 in an example of this embodiment includes an analysis unit 13 that analyzes the speech of the participants in the conference, and the analysis unit 13 calculates the result of the above analysis as a predetermined analysis value. The analysis process executed by the analysis unit 13 includes a first analysis process of dividing the conference into one or two or more phases and calculating a predetermined analysis value for each of the divided phases, and a second analysis process of calculating a predetermined analysis value for each participant in the conference. Note that the specific configuration of the analysis process by the analysis unit 13 may be arbitrarily changed. For example, in addition to the process of analyzing the speech of the participants, an analysis process regarding the total time of the conference and an analysis process based on the captured video may be further performed.
[0029] FIG. 2 is a flowchart showing the flow of the first analysis process in an example of the present embodiment. In an example of the present embodiment, the first analysis process is a process of dividing a meeting into one or more phases and calculating a divergence / convergence value and a freshness value as predetermined analysis values for each of the phases.
[0030] In an example of the present embodiment, as described above, the speech of each participant in the meeting executed by the meeting execution unit 11 is converted into text information by the speech recognition unit 12 of the meeting analysis device 1 and recorded in the database 15. In the first analysis process in an example of the present embodiment, divergence / convergence analysis (step S1) and freshness analysis (step S2) are performed based on the text information recorded in the database 15.
[0031] In an example of the present embodiment, the divergence / convergence value is a value indicating the similarity between each of the phases obtained by dividing the speech in the meeting according to a predetermined criterion and all the phases, as described above. In step S1, first, the text information recorded in the database 15, that is, the speech of all the participants in the meeting, is divided into one or more phases according to a predetermined criterion (step S11).
[0032] The predetermined criterion may be arbitrarily selected. For example, a time such as 1 minute or 5 minutes may be used as the predetermined criterion. In this case, for example, if the predetermined criterion is 5 minutes and the total time of the meeting is 1 hour, the speech in the meeting is divided into 12 phases. Alternatively, the speech in the meeting may be divided into a predetermined number of phases so that the amount of speech is equal between each phase. In the analysis process in an example of the present embodiment, for each of the phases divided in step S1, as predetermined analysis values, a freshness value indicating how many new words were spoken and a divergence / convergence value indicating the similarity with all the phases are calculated.
[0033] Next, the utterances of each participant in each phase are subjected to morphological analysis (step S12). The morphological analysis in an example of the present embodiment is a process of dividing each of the utterances converted into text information by the speech recognition unit 12 into one or more morphemes, that is, words which are the smallest units having meanings.
[0034] When step S12 is completed, next, nouns are extracted as predetermined words from the results of the morphological analysis in step S12 (step S13). In the aforementioned step S12, the utterances of the participants are divided into one or more morphemes. In step S13, the part of speech of each of the divided morphemes is determined, and the morphemes whose determined part of speech is a noun are extracted. Note that in an example of the present embodiment, nouns are extracted, but other parts of speech may be extracted. For example, adjectives may be extracted as predetermined words. Also, not only a single type of part of speech but also a plurality of parts of speech, for example, nouns and adjectives, may be extracted as predetermined words. Further, in an example of the present embodiment, morphological analysis (step S12) and word extraction (step S13) are performed after phase division (step S11), but phase analysis may be executed after morphological analysis or word extraction. In this case, for example, the number of extracted words or the like may be used as a predetermined criterion in phase division.
[0035] When the word extraction by step S13 is completed, next, weighting of each of the extracted words is performed (step S14). The weighting method may be arbitrarily selected. In an example of the present embodiment, weighting is performed by TF-IDF (Term Frequency - Inverse Document Frequency) with the entire utterance in each phase regarded as one document.
[0036] When the word weighting by step S14 is completed, next, vectors for each phase are calculated (step S15). As described above, in an example of the present embodiment, nouns are extracted as words from the utterances in each phase, and weighting of the extracted words is performed. In step S15, vectors for each phase are calculated as vectors having the number of words that appear in all phases as the number of dimensions.
[0037] When the vector calculation in step S15 is completed, next, the similarity between each phase is calculated (step S16). In an example of the present embodiment, the vectors of each phase have been calculated in step S15 described above, and in step S16, the similarity between each phase is calculated using the calculated vectors of each phase. Note that in an example of the present embodiment, the cosine measure is used as the similarity between vectors.
[0038] When the calculation of the similarity between each phase in step S16 is completed, next, the divergence / convergence value of each phase is calculated based on the calculated similarity (step S17). In an example of the present embodiment, the divergence / convergence value is a value indicating the similarity between the speech content of each phase and the speech content of the entire meeting. In an example of the present embodiment, for each phase, the sum of the similarities with other phases is obtained and divided by the maximum value of the above sums for all phases to calculate the divergence / convergence value of each phase. The divergence / convergence value calculated by the above method is a value normalized in the range of 0 to 1, and the divergence / convergence value of the phase whose sum of similarities is the above maximum value is 1.
[0039] The above is the calculation process of the divergence / convergence value in an example of the present embodiment. In an example of the present embodiment, following the divergence / convergence analysis in step S1, the freshness analysis in step S2 is performed. The freshness analysis in an example of the present embodiment calculates an analysis value indicating how many new words are spoken in each phase, as described above.
[0040] As described above, in the first analysis process in an example of this embodiment, the similarity of each phase feeling is calculated in step S15. In step S2, first, cluster classification is performed for each phase based on this similarity (step S21). In an example of this embodiment, when the similarity between each phase exceeds a predetermined threshold, the two phases are regarded as belonging to the same cluster, a predetermined threshold at which the number of clusters is maximized is calculated, and each phase is classified into a cluster. Note that a phase whose similarity with other phases does not exceed the predetermined threshold is classified as a phase that does not belong to any cluster.
[0041] When the cluster classification is completed for all phases, next, the topic words of each cluster and each phase are extracted (steps S22, S23). A topic word is a word that characterizes the speech content of each cluster. In an example of this embodiment, the topic words of each cluster are extracted (step S22), and then, from the topic words of the cluster to which each phase belongs, the topic words that appear in that phase are extracted (step S23). For the extraction of topic words, the weights in step S14 described above are used for the phases classified into clusters. The maximum number of topic words extracted in an example of this embodiment is 20, and among the words that appear in all phases belonging to one cluster, that is, one cluster, the maximum 20 words are extracted in descending order of the weight values. After the topic words of each cluster are extracted, in each phase, the topic words that appear in that phase are extracted.
[0042] When the extraction of the topic words in each phase is completed, next, the number of topic words extracted in each phase is calculated as the freshness (step S24). As described above, in an example of this embodiment, the value indicating how many new words are spoken in each phase is the freshness, and the number of the above topic words is used as the freshness value of each phase.
[0043] The above is the flow of the first analysis process in an example of this embodiment.
[0044] Next is the flow of the second analysis process in an example of this embodiment. FIG. 3 is a flowchart showing the flow of the second analysis process in an example of this embodiment. Note that the second analysis process shown in FIG. 3 is a process of calculating an affirmation / negation value, a sympathy value, a filler value, and spoken words as predetermined analysis values.
[0045] In the second analysis process in an example of this embodiment, morphological analysis is performed on the utterances of the participants in the meeting (step S3), and affirmation / negation analysis (step S4), sympathy analysis (step S5), and filler analysis (step S6) performed using the results of the morphological analysis, and spoken word analysis (step S7) performed without using morphological analysis are executed. Note that the order of executing steps S4 to S7 may be arbitrarily selected. Steps S4 to 6 may be executed in parallel after step S3, and step S7 may be executed in parallel with steps S3 to S6, or steps S3 to S7 may be executed sequentially.
[0046] The morphological analysis in step S3 is a process of dividing each of the utterances of the participants in the meeting, that is, each of the text information converted by the speech recognition unit 12, into one or more morphemes. Note that the specific process of the morphological analysis in step S3 is the same as that of step S12 described above. Whether to execute the morphological analysis in both the first analysis process and the second analysis process, or to execute the morphological analysis first and then execute the first and second analysis processes may be arbitrarily selected.
[0047] Step S4 calculates a positive / negative value as a predetermined analysis value, that is, a value indicating the degree of positive / negative reactions of the participants. In step S4, first, the number of predetermined words indicating positive or negative reactions in the participants' statements is tabulated (step S41). The predetermined words can be arbitrarily selected. For example, as nouns indicating positive reactions, words such as "honest", "peaceful", "kind", etc., and as nouns indicating negative reactions, words such as "cowardly", "depressed", "disappointed", etc. are recorded in advance in database 15 or other dictionaries, and the number of words indicating positive reactions included in each statement and the number of words indicating negative reactions included in each statement are tabulated. Note that the tabulation process in step S41 may be performed on the result of the morphological analysis executed in step S3 described above, or may be performed on the text information obtained by converting the participants' statements by the speech recognition unit 12, that is, on the object of the process in step S3.
[0048] When the tabulation in step S41 is completed, next, the positive / negative value of the participants is calculated (step S42). The positive / negative value is composed of a positive value indicating the degree to which the participants showed positive reactions in the meeting and a negative value indicating the degree to which the participants showed negative reactions in the meeting. In step S5 described above, the participants' statements were morphologically analyzed and divided into one or two or more morphemes. In step S42, the value obtained by expressing, as a percentage, the number of predetermined words indicating positive / negative reactions tabulated in step S41 in the number of morphemes of all the participants' statements in the meeting is calculated as the above positive value and negative value.
[0049] Note that in an example of this embodiment, as described above, the participants' statements in the meeting are converted into text information by the speech recognition unit 12 for each statement and recorded in the database 15. How to execute step S4 on the text information can be arbitrarily selected. For example, step S4 may be executed on all the above text information in the meeting, or step S41 may be executed for each text information, and after step S41 is executed for all the text information, the execution results may be tabulated and step S42 may be executed.
[0050] In step 5, as a predetermined analysis value, a sympathy value, that is, a value indicating the degree of reaction with which the participants sympathize, is calculated. In step S5, first, for each statement of the participants, it is determined whether the statement indicates sympathy (step S51).
[0051] In an example of this embodiment, in step S3 described above, the statements of the participants are divided into one or more morphemes by morphological analysis. In step S51, the number of the above-mentioned one or more morphemes divided is tabulated for each part of speech, and when the part of speech with the most interjections in the tabulation result, it is determined that the statement indicates sympathy.
[0052] When the determination according to step S51 described above is made for all the statements of the participants in the meeting, then, a sympathy value is calculated (step S52). In step S51 described above, for each statement of the participants, it has been determined whether the statement indicates sympathy, and in step S52, a value obtained by expressing, as a percentage, the number of statements indicating the above-mentioned sympathy in the number of statements of the participants in the meeting is calculated as the above-mentioned sympathy value.
[0053] In step 6, as a predetermined analysis value, a filler value, that is, a value indicating the degree of statements having no specific meaning, is calculated. In step S6, first, the number of filler words in the statements of the participants is tabulated (step S61). According to step S3 described above, the statements of the participants are morphologically analyzed for each statement and divided into one or more morphemes. In step S61, the number of morphemes determined to be a part of speech of filler, that is, a word having no specific meaning, in the above-mentioned one or more morphemes divided is tabulated.
[0054] When the counting of the number of filler words for all the remarks of the participants in the meeting is completed by the above-described step S61, then the filler value is calculated (step S62). In the above-described step S81, the number of filler words is counted for each remark of the participants. In step S62, a value obtained by expressing, as a percentage, the number of filler words in the number of morphemes of all the remarks of the participants in the meeting is calculated as the filler value.
[0055] Step S7 calculates, as predetermined analysis values, the total number of characters, the average number of characters per remark, and the speaking speed of the remarks of the participants in the meeting. In step S7, first, the number of characters in the remarks of the participants in the meeting is measured (step S71). In an example of the present embodiment, as described above, each of the remarks of the participants in the meeting is converted into text information by the speech recognition unit 12. In step S71, the number of characters included in each remark is measured.
[0056] When the measurement of the number of characters per remark is completed, then the speaking speed in each remark of the participants in the meeting is calculated (step S72). The speech recognition unit 12 in an example of the present embodiment converts each remark of the participants in the meeting into text information and records it in the database 15 together with the information indicating the participant who made the remark, the timing of the remark in the meeting, and the time taken for the remark, as described above. In step S72, the number of characters measured in the above-described step S71 is divided by the value obtained by dividing the number of characters by the time taken for the above-described remark, and the number of characters spoken per unit time is calculated as the speaking speed.
[0057] Next, the total number of characters in the remarks of the participants is calculated (step S73). In the above-described step S71, the number of characters per remark of the participants in the meeting is measured, and step S73 calculates the total number of characters in the remarks of the participants in the meeting by totaling the measured number of characters for all the remarks.
[0058] Next, the average number of characters in the participants' statements is calculated (step S74). In the above-described step S71, the number of characters for each statement of the participants in the meeting has been measured, and in step S74, the average value of the measured number of characters in all the statements of the participants in the meeting is calculated.
[0059] In an example of the present embodiment, the first analysis process of steps S1 to S2 and the second analysis process of steps S3 to S6 may be configured to analyze the statements in a single meeting, or may be configured to analyze the statements in one or more meetings. In an example of the present embodiment, the first and second analysis processes are executed for each of one or more meetings to calculate a predetermined analysis value from the statements for each meeting, and when displaying the predetermined analysis values in two or more meetings, it is configured to use the value obtained by aggregating the predetermined analysis values calculated for each meeting. On the other hand, when a predetermined analysis value for a single meeting is not necessarily required, the first and second analysis processes may be executed for the statements in all of one or more meetings.
[0060] Also, in an example of the present embodiment, among the predetermined analysis values, the divergence / convergence value and the freshness value are calculated for each phase, and the positive / negative value, the empathy value, and the filler value are calculated for each participant. However, the calculation unit of the predetermined analysis value may be arbitrarily selected, and all of the predetermined analysis values may be calculated for each meeting or each phase obtained by dividing the meeting, or may be calculated for each participant. Further, a part or all of the predetermined analysis values may be calculated for each participant in each phase.
[0061] Next, the display of the analysis results in an example of the present embodiment will be described. As described above, in an example of the present embodiment, the conference analysis apparatus 1 includes an analysis result display unit 14, and the analysis result display unit 14 displays an analysis result screen and / or an analysis result details screen, which will be described later, on the display device 31 of the browsing terminal 3. Note that, as will be described later, the analysis result screen in an example of the present embodiment is configured to selectively switch between any of the following display formats: displaying analysis results for each participant in a single conference, displaying analysis results for each participant in two or more conferences, displaying analysis results for each conference, and graphically displaying the analysis results for a conference.
[0062] FIG. 4 is a diagram showing the configuration of an analysis result screen W1 that displays the analysis results of a single conference for each participant in an example of the present embodiment. The analysis result screen W1 shown in FIG. 4 is a screen displayed on the display device 31 of the browsing terminal 3 by the analysis result display unit 14 of the conference analysis apparatus 1, and is a screen that displays the results of the first and second analysis processes for each participant in a single conference. As shown in FIG. 4, the analysis result screen W1 is a screen that displays the number of characters W12a, the average number of characters W12b, the speaking speed W12c, the positive / negative value W12d, the empathy value W12e, and the filler value W12f in a tabular format, with the participants W11 as rows and the columns W12.
[0063] The participants W11 are the participants in the conference analyzed by the first and second analysis processes. In an example of the present embodiment, the names of the participants are displayed as the participants W11.
[0064] Column W12 displays the analysis results based on the statements of each participant in the meeting. In an example of this embodiment, the values displayed in analysis results W12a to W12f are, respectively, the total number of characters calculated in step S73 described above, the average number of characters calculated in step S74 described above, the speaking speed calculated in step S72 described above, the positive / negative value consisting of the positive value and the negative value calculated in step S4 described above, the empathy value calculated in step S5 described above, and the filler value calculated in step S6 described above. Also, each of columns W12 is configured to display a predetermined analysis value calculated in steps S4 to S7 described above and the predetermined analysis value in a bar graph. Note that the bar graph in an example of this embodiment is configured to display the relative magnitude of values for the purpose of facilitating comparison of the magnitude of values with other participants for each of columns W12. However, the specific configuration of the values, the graph, and other displays may be arbitrarily changed. For example, the bar graph may be configured to show the absolute magnitude of values instead of the relative magnitude of values. Also, the positive / negative value included in the predetermined analysis value is composed of the positive value and the negative value as described above, and the positive / negative value W12d for displaying this is configured to display the positive value on the left end side and the negative value on the right end side from the center of the bar graph, each showing the relative magnitude.
[0065] FIG. 5 is a diagram showing the configuration of an analysis result screen W2 that displays the analysis results of two or more meetings for each participant in an example of this embodiment. The above-described analysis result screen W1 displays the analysis results in a single meeting, while the analysis result screen W2 shown in FIG. 5 is a screen showing the analysis results of two or more meetings. As shown in FIG. 5, the analysis result screen W2 has participant W21 as a row and, as column W22, displays the number of meetings attended W22a, the number of characters W22b, the average number of characters W22c, the speaking speed W22d, the positive / negative value W22e, the empathy value W22f, and the filler value W22g in a tabular format. Note that the specific configuration of column W22 of the analysis result screen W2 and the method of expression may be arbitrarily changed. For example, in an example of this embodiment, the number of characters W22b and the average number of characters W22c are displayed as values indicating the number of characters as they are. However, instead of displaying the value indicating the number of characters, it may be expressed as a value in a predetermined unit by normalization, standardization, or other methods.
[0066] In the analysis result screen W2, the number of meetings attended W22a is the number of meetings each participant attended. Also, the number of characters W22b to the filler value W22g are values obtained by analyzing each of the predetermined analysis values W12a to W12f in the analysis result screen W1 for two or more meetings.
[0067] FIG. 6 is a diagram schematically showing the configuration of an analysis result screen W3 that displays the analysis results for each meeting in an example of the present embodiment. The above-described analysis result screens W1 and W2 are screens that display each predetermined analysis value in a tabular form with participants as rows, but as shown in FIG. 6, the analysis result screen W3 has meetings W31 as rows and, as columns W32, the date and time W32a, the meeting time W32b, the number of participants W32c, the number of characters W32d, the average number of characters W32e, the speaking speed W32f, the positive / negative value W32g, the empathy value W32h, and the filler value W32i are displayed in a tabular form.
[0068] In the column W32, the date and time W32a is the date and time when the meeting W31 was held, the meeting time W32b is the length of the meeting W31, and the number of participants W32c is the number of participants who participated in the meeting W31. Also, the number of characters W33d, the average number of characters W32e, the positive / negative value W32g, the empathy value W32h, and the filler value W32i are values obtained by analyzing each of the predetermined analysis values W12a to W12f in the above-described analysis result screen W1 for each meeting W31.
[0069] Note that the specific method of analysis may be arbitrarily selected. In an example of the present embodiment, values calculated by aggregating the predetermined analysis values calculated in the above-described steps S4 to 7 for each meeting are used. Also, the predetermined analysis values to be displayed on the analysis result screen W6 may be arbitrarily changed. For example, a value indicating the percentage of the meeting time W32b during which no participant spoke in the meeting W31 may be displayed.
[0070] FIG. 7 is a diagram showing the configuration of an analysis result graph screen W4 that graphically displays the analysis results of a meeting in an example of this embodiment. As shown in FIG. 7, the analysis result screen W4 includes a video area W41, a transcription area W42, and a graph area W43.
[0071] The video area W41 is an area for displaying and playing back the video taken during the meeting, that is, the video taken by the imaging device 23 of the participant terminal 2. The transcription area W42 is an area for displaying the text information converted from the speech by the speech recognition unit 12 in chronological order together with the information indicating the speaker.
[0072] The graph area W43 is an area for displaying a predetermined analysis value by a graph having the predetermined analysis value on the vertical axis and the elapsed time of the meeting on the horizontal axis. In an example of this embodiment, from the divergence / convergence value, freshness value, positive / negative value, empathy value, filler value, and number of spoken characters included in the predetermined analysis value, one analysis value desired by the viewer of the browsing terminal 3 is selectively displayed in the graph area W43. Note that how to display the predetermined analysis value in the graph area may be arbitrarily selected. For example, the positive / negative value, empathy value, filler value, and number of spoken characters are values calculated for each participant as described above. When displaying in the graph area W43, graphs for each participant in the meeting may be superimposed or represented as a stacked graph.
[0073] Also, in the analysis result screen W4 in an example of the present embodiment, as described above, in the transcription area W42, the participants' remarks are displayed in chronological order. When the participants' remarks do not fit within the transcription area W42, a scroll bar is displayed on the right end side or the left end side of the transcription area W42, and the participants' remarks at any timing can be viewed as text information. Further, when the video is played in the video area W42 while the scroll bar is displayed, the display content of the transcription area W42 is scrolled in accordance with the timing of the played video, and the contents of the video area W42 and the transcription area W43 are configured to be synchronized. When any one of the remarks displayed in the transcription area W42 is selected, the video in the video area W42 is played back from the time of the selected remark, and in the graph area W43, the time of the selected remark is cooperatively displayed by the vertical band W43a. Also, when any position is selected in the graph area W43, the video in the video area W42 is played back from the elapsed time corresponding to the selected position.
[0074] FIG. 8 is a diagram showing the configuration of an analysis result details screen W5 for displaying details of the analysis result in an example of the present embodiment. The analysis result details screen W5 is a diagram that is displayed together with or selectively with the above-described analysis result graph screen W4. As shown in FIG. 8, in the analysis result details screen W5, the column W51 is each of the phases obtained by dividing the meeting into one or two or more parts in the above-described step S11, and it is a tabular screen in which the analysis content W52 regarding the phase of each column W51 is in a row. In an example of the present embodiment, the analysis content W52 is composed of a start time W52a, an end time W52b, the number of characters W52c, the representativeness W52d, the freshness W52e, the topic number W52f, and the topic word W52g.
[0075] The start time W52a and the end time W52b are the start time and the end time of the corresponding phase, the number of characters W52c is the number of characters of the speeches of all participants in the corresponding phase. Also, the representativeness W52d is the divergence / convergence value of the corresponding phase. As described above, in an example of this embodiment, the divergence / convergence value is a value obtained by dividing the sum of the similarities of each phase by the maximum value of the sum. The divergence / convergence value of each phase can be used simultaneously as a value indicating how much the speech of each phase represents the content of the entire meeting. On the analysis result details screen W5, the divergence / convergence value is displayed as the representativeness W52d.
[0076] The freshness W52e is the freshness value of the corresponding phase, the topic number W52f is the number indicating the cluster to which the corresponding phase belongs, and the topic words W52g are a list of the topic words extracted in the above step S23.
[0077] In an example of this embodiment, on the analysis result details screen W5, when the viewer operates the viewing terminal 3 to select an arbitrary topic word W52g, in the graph area W43 of the above-described analysis result graph screen W4, the phase in which the selected topic word W52g was spoken is highlighted with a vertical band W43a. As described above, the analysis result details screen W5 is a screen that is displayed together with or selectively with respect to the above-described analysis result graph screen W4. When the analysis result graph screen W4 and the analysis result details screen W5 are selectively displayed, that is, when either one is selected and displayed, when the above topic word W52g is selected, first, a screen transition is made from the analysis result details screen W5 to the analysis result graph screen W4, and then, in the graph area W43 of the transitioned analysis result graph screen W4, the corresponding phase may be configured to be cooperatively displayed with the band W43a.
[0078] The description of an example of this embodiment is as above. Note that the embodiments of the present invention are not limited to the above. For example, in an example of this embodiment, the conference analysis apparatus 1 first holds a conference by the conference execution unit 11, and then the analysis unit 13 performs analysis based on the text information obtained by converting the speech of the participants by speech recognition by the speech recognition unit 12 of the result of the conference. However, in parallel with the conference by the conference execution unit 11, each time the speech recognition unit 12 performs conversion into text information, the analysis unit 13 may sequentially analyze this and perform analysis in real time.
[0079] Other specific configurations are not limited to this embodiment, and various changes can be made without departing from the spirit of the present invention.
Explanation of Reference Numerals
[0080] 1 Analysis apparatus 11 Conference execution unit 12 Speech recognition unit 13 Analysis unit 14 Analysis result display unit 15 Database 2 Participant terminal 21 Display device 22 Input device 23 Photographing device 24 Speech input device 3 Browsing terminal 31 Display device 32 Input device 4 Network
Claims
1. A conference analysis device for analyzing the speeches of two or more participants in a conference, an analysis unit that calculates a predetermined analysis value for each of the participants based on the speeches in the conference in which the participant participated; the two or more participants and an analysis result display unit that displays some or all of the calculated predetermined analysis values for each of the participants in a tabular format; A conference analysis device comprising:
2. A conference analysis device for analyzing the speeches of two or more participants in two or more conferences, an analysis unit that calculates a predetermined analysis value for each of the two or more participants based on the speeches in the conference in which the participant participated; the two or more participants and an analysis result display unit that displays some or all of the calculated predetermined analysis values for each of the participants in a tabular format; A conference analysis device comprising:
3. The analysis unit further calculates a predetermined analysis value in the conference based on the predetermined analysis values of the participants who participated in the conference for each of the two or more conferences, The analysis result display unit further shows the two or more conferences and some or all of the calculated predetermined analysis values in the conference, The conference analysis device according to Claim 2.
4. A conference analysis device for analyzing the speeches of two or more participants in a conference, an analysis unit that calculates a predetermined analysis value for each of the phases obtained by dividing the speech based on a predetermined criterion for each of the conferences and / or the participants or for each of the participants, an analysis result display unit that displays some or all of the predetermined analysis values in a tabular format; comprising The predetermined analysis value is a freshness value indicating how many new words are spoken in each of the phases, a divergence / convergence value indicating the similarity between each of the phases and all the phases, an affirmative / negative value indicating the degree of affirmative and negative reactions of the participant, a sympathy value indicating the degree of reaction in which the participant sympathizes with the other participants, a filler value indicating the degree of speech of the participant that has no specific meaning, A conference analysis device including one or more of the above.
5. The analysis result display unit a graph area that displays the predetermined analysis value by a graph having the predetermined analysis value on the vertical axis and the elapsed time in the conference on the horizontal axis; a video area that plays back the video of the conference; a transcription area that displays the information obtained by textifying the speech in the conference together with the information indicating the participant who made the speech; displays an analysis result graph screen including The conference analysis device according to Claim 4.
6. The analysis result display unit further displays, together with or selectively instead of the analysis result graph screen, a detailed analysis result screen that lists, for each phase, the predetermined analysis value and the topic words extracted from the speech in that phase. The conference analysis device according to claim 5.
7. The analysis result display unit when any of the topic words is selected, highlights the phase in which the selected topic word was spoken in the graph area. The conference analysis device according to claim 6.
8. A program for causing a computer to function as the conference analysis device according to any one of claims 1 to 7.
Citation Information
Patent Citations
Speech recognition program, speech recognition method, speech recognition device, and speech recognition system
JP2022121643A