Speech analysis device and speech analysis method
The voice analysis device addresses the challenge of analyzing discussions between two groups by generating and outputting section information on dominant speech trends, enhancing the understanding of speech patterns and transitions in discussions.
Patent Information
- Application Number
- JP2025178546
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-21
AI Technical Summary
Existing systems struggle to analyze speech trends in discussions between two groups, specifically in a way that existing technologies have not addressed or effectively solve the technical problem of the technical problem of addressing the technical problem of the technical problem of the technical problem of analyzing discussions between two groups, such as a learner and a teacher, or a discussion between two groups, where the system disclosed in Patent Document 1 displays what each participant has said, making it difficult to grasp the trends in the speech of the two groups, hindering effective analysis.
A voice analysis device that acquires time series information on speech situations of each group, generates section information indicating dominant speech trends, and outputs this information to facilitate analysis of discussions between two groups, including classification, section determination, and comparison with reference information.
Enables easier analysis of speech trends in discussions between two groups by providing real-time and post-output of section information, allowing analysts to understand dominant speech patterns and transitions, thereby improving the analysis of discussions.
Smart Images

Figure 2026010176000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech analysis device and a speech analysis method for analyzing speech uttered in a discussion. [Background technology]
[0002] Patent document 1 discloses a system that uses a microphone to capture the voices of participants during a meeting, identifies the participant currently speaking based on voiceprint data extracted from the voice, and displays the speech status of each of the multiple participants on a display. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-208482 [Non-patent literature]
[0004] [Non-Patent Document 1] Tanigawa, Wataru, Kato, Naoki, and Takano, Masaaki, "Analysis of changes in lesson content using ST analysis," Tokyo Gakugei University Educational Practice Research, Tokyo Gakugei University, 2021, Vol. 17, pp. 77-85 Summary of the Invention [Problem to be solved by the invention]
[0005] There are cases where a discussion between two groups, such as a discussion between a learner and a teacher, or a discussion between two groups, is analyzed. The system disclosed in Patent Document 1 displays what each of the multiple participants has said at each time, making it difficult for an analyst to grasp the trends in the speech of the two groups into which the multiple participants are divided, making it difficult to analyze the discussion between the two groups.
[0006] The present invention has been made in consideration of these points, and aims to make it easier to analyze speech trends in discussions between two groups. [Means for solving the problem]
[0007] A first aspect of the voice analysis device of the present invention includes an acquisition unit that acquires time series information indicating the speech situations of each of the first and second groups in the speech uttered by each of the participants belonging to a first group and the second group over time in a discussion; a generation unit that generates, based on the time series information, section information that associates each of a plurality of sections that make up the discussion with a section tendency that indicates which of the first and second groups is the dominant speech in that section, for all or part of the discussion; and an output unit that outputs the section information.
[0008] The section tendency may indicate which of the first and second groups is the dominant speech, or whether the first and second groups are competitive with each other.
[0009] The generation unit may determine each of the plurality of sections so that the section is equal to or longer than a predetermined time, and may determine the section tendency for the section by comparing the speech situations of the first group and the second group in the section.
[0010] The time-series information may be information indicating which of the first group and the second group has a larger volume of speech for each predetermined time frame during a period from the start point to the end point of the discussion.
[0011] The voice analysis device may further include a classification unit that classifies a plurality of participants into the first group and the second group.
[0012] The classification unit may change the participants belonging to the first group and the second group between multiple periods in the discussion.
[0013] The classification unit may generate a first parent group including the first group and the second group into which a portion of the plurality of participants are classified, and a second parent group including the first group and the second group into which a portion of the plurality of participants who do not belong to the first parent group are classified, and the output unit may simultaneously output the section information of the first parent group and the section information of the second parent group.
[0014] The classification unit may generate a first parent group including the first group and the second group into which a portion of the plurality of participants are classified, and a second parent group including the first group and the second group into which a portion of the plurality of participants who do not belong to the first parent group are classified, and further generate a third parent group including the first group and the second group into which participants belonging to the first parent group and the second parent group are classified, and the output unit may output the section information of at least one of the first parent group and the second parent group, and the section information of the third parent group.
[0015] The classification unit may change the participants belonging to the first group to which the specific participant belongs and the participants belonging to the second group based on the position of the specific participant.
[0016] The output unit may output words included in the utterance of each of the plurality of sections, which words are extracted by performing a speech recognition process on the speech, in association with the section.
[0017] The output unit may output characteristics of the entire discussion based on the section trends of the multiple sections that make up the discussion.
[0018] The speech analysis device may further include a selection unit that selects reference section information to be compared with the section information, and the output unit may output a result of the comparison between the section information and the reference section information.
[0019] The output unit may output, during the discussion, information corresponding to a difference between the section information and the reference section information as the comparison result.
[0020] After the discussion, the output unit may output the section trend of each of the multiple sections indicated by the section information in association with the section trend of each of the multiple sections indicated by the reference section information.
[0021] A second aspect of the speech analysis method of the present invention includes the steps of: acquiring time series information executed by a processor that indicates the speech situations of each of the first and second groups over time in the speech uttered by each of the participants belonging to a first group and the second group in a discussion; generating section information based on the time series information that associates, for all or part of the discussion, each of a plurality of sections that make up the discussion with a section tendency that indicates which of the first and second groups is the dominant speech in that section; and outputting the section information. [Effects of the Invention]
[0022] The present invention provides the advantage of making it easier to analyze speech trends in discussions between two groups. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a schematic diagram of a voice analysis system S according to an embodiment. [Figure 2] 1 is a block diagram of a voice analysis system S according to an embodiment. [Figure 3] 10 is a schematic diagram for explaining a method in which a selection unit selects reference interval information. FIG. [Figure 4] 10 is a schematic diagram for explaining a method in which an acquisition unit acquires time-series information. FIG. [Figure 5] 10 is a schematic diagram for explaining a method in which the generation unit determines a section tendency. FIG. [Figure 6]10 is a schematic diagram for explaining a method in which an output unit outputs section information in real time. FIG. [Figure 7] FIG. 10 is a schematic diagram for explaining a method in which an output unit outputs section information afterward. [Figure 8] FIG. 2 is a flowchart illustrating an exemplary voice analysis method executed by a voice analysis device according to an embodiment. [Figure 9] FIG. 10 is a diagram showing a flowchart of a section information generation process in an exemplary voice analysis method executed by a voice analysis device according to an embodiment. [Figure 10] FIG. 10 is a schematic diagram for explaining a method in which an output unit outputs section information in a modified example. [Figure 11] 10 is a schematic diagram for explaining a method in which a generating unit generates section information in a modified example. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0024] [Outline of Voice Analysis System S] 1 is a schematic diagram of a voice analysis system S according to this embodiment. The voice analysis system S includes a voice analysis device 1, a sound collection device 2, and an information terminal 3. The number of sound collection devices 2 and information terminals 3 included in the voice analysis system S is not limited. The voice analysis system S may also include other devices such as servers and terminals.
[0025] The voice analysis device 1 is a computer that analyzes voices uttered in a discussion involving multiple participants and provides the analysis results to an analyst. The analyst may be some of the multiple participants or may be a person different from the multiple participants. The voice analysis device 1 analyzes voices acquired by the sound collection device 2 and outputs the analysis results to the sound collection device 2 or the information terminal 3. The voice analysis device 1 is connected to the sound collection device 2 and the information terminal 3 via a network such as a local area network or the Internet, either wired or wirelessly.
[0026] The speech analysis device 1 analyzes speech from a discussion in which multiple participants are divided into at least two groups. The discussion to be analyzed may be, for example, a class, a group discussion, a debate, or a meeting. The multiple participants are classified into either a first group or a second group. Participants in the first group are, for example, instructors such as teachers or tutors. Participants in the second group are, for example, learners such as pupils or students. The multiple learners may also be classified into the first group and the second group. The multiple participants may also be classified based on other criteria.
[0027] A plurality of participants may be categorized into a plurality of parent groups, and within each of the parent groups, a plurality of participants may be categorized into a first group or a second group. In this case, each parent group may include a first group and a second group. For example, one parent group corresponds to a table surrounded by a plurality of participants, and a sound collection device 2 is placed at the table. The plurality of participants surrounding the table are categorized into a first group and a second group.
[0028] The voice analysis device 1 may also analyze the voice of a discussion (for example, a web conference) held over a network. In this case, a sound collection device 2 is placed in each space where multiple participants are present during the discussion, and each sound collection device 2 is associated with one of the multiple participants.
[0029] The sound collection device 2 is a device that acquires sounds uttered during a discussion. The sound collection device 2 is equipped with, for example, a microphone array including a sound collection unit such as a plurality of microphones arranged in different orientations. The microphone array includes, for example, a plurality of microphones (e.g., eight) arranged at equal intervals on the same circumference in a horizontal plane relative to the ground. By using such a microphone array, the voice analysis device 1 can identify which participant is the speaker (sound source) based on the sounds uttered by the plurality of participants surrounding the sound collection device 2. The sound collection device 2 transmits the sounds acquired using the microphone array to the voice analysis device 1 as voice data. The sound collection device 2 may also be equipped with an audio output unit such as a speaker.
[0030] The information terminal 3 is a computer that outputs information, and is, for example, a smartphone, a tablet terminal, or a personal computer. The information terminal 3 is used, for example, by at least some of the multiple participants. The information terminal 3 may also be used by an analyst different from the multiple participants. The information terminal 3 has, for example, a display unit such as a liquid crystal display. The information terminal 3 displays the information received from the voice analysis device 1 on the display unit.
[0031] Furthermore, the information terminal 3 may have a sound collection unit such as a microphone, and function as the sound collection device 2. In this case, the information terminal 3 used by each of the multiple participants transmits the voice acquired using the sound collection unit to the voice analysis device 1 as voice data.
[0032] An overview of the process of analyzing speech by the speech analysis system S according to this embodiment will be described below. The speech analysis device 1 classifies multiple participants into a first group or a second group. For example, the speech analysis device 1 receives a setting from an information terminal 3 as to whether multiple participants belong to the first group or the second group, or automatically classifies multiple participants into the first group or the second group based on the attributes of the multiple participants.
[0033] The voice analysis device 1 acquires voices uttered by multiple participants in a discussion from the sound collection device 2. The voice analysis device 1 identifies the speech periods of each of the multiple participants in the acquired voices, thereby acquiring time-series information indicating the speech situations of the first group and the second group over time. The time-series information is, for example, information indicating which of the first group and the second group speaks more loudly for each predetermined time frame.
[0034] Based on the acquired time-series information, the speech analysis device 1 generates section information that associates each of a plurality of sections constituting the discussion with a section tendency indicating which of the first and second groups is the dominant speech in that section. The section information may indicate whether the speech of the first and second groups is dominant, as well as whether the speech of the first and second groups is competitive. The speech analysis device 1 outputs the generated section information to at least one of the sound collection device 2 and the information terminal 3.
[0035] In this way, the speech analysis system S determines the section tendency indicating which of the two groups is the dominant speaker for each section of the discussion based on the audio of the discussion, and notifies the analyst of the section and the section tendency in association with each other. This makes it easier for the analyst to grasp the speech tendencies of the two groups and to analyze the speech tendencies in the discussion between the two groups.
[0036] [Configuration of the voice analysis system S] FIG. 2 is a block diagram of a speech analysis system S according to this embodiment. In FIG. 2, arrows indicate main data flows, and data flows other than those shown in FIG. 2 may also exist. In FIG. 2, each block indicates a functional configuration rather than a hardware (device) configuration. Therefore, the blocks shown in FIG. 2 may be implemented in a single device, or may be implemented separately in multiple devices. Data may be exchanged between blocks via any means, such as a data bus, a network, or a portable storage medium.
[0037] The speech analysis device 1 includes a storage unit 11 and a control unit 12. The speech analysis device 1 may be configured by two or more physically separate devices connected by wire or wirelessly. The speech analysis device 1 may also be configured by a cloud, which is a collection of computer resources.
[0038] The storage unit 11 is a storage medium including a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk drive, etc. The storage unit 11 stores in advance programs to be executed by the control unit 12. The storage unit 11 may be provided outside the voice analysis device 1, in which case data may be exchanged between the storage unit 11 and the control unit 12 via a network.
[0039] The control unit 12 includes a selection unit 121, a classification unit 122, an acquisition unit 123, a generation unit 124, and an output unit 125. The control unit 12 is a processor such as a CPU (Central Processing Unit), and functions as the selection unit 121, the classification unit 122, the acquisition unit 123, the generation unit 124, and the output unit 125 by executing a program stored in the storage unit 11. At least some of the functions of the control unit 12 may be performed by an electric circuit. Furthermore, at least some of the functions of the control unit 12 may be realized by the control unit 12 executing a program executed via a network.
[0040] The processing executed by the speech analysis device 1 will be described in detail below. The selection unit 121 selects reference section information to be compared with the section information. The section information is information that associates each of multiple sections constituting a discussion with the section tendency of that section. The section tendency indicates which of the utterances from the first group or the second group is dominant, or whether the utterances from the first group and the second group are in balance. The section information is generated by the generation unit 124, described below, based on the speech of the discussion to be analyzed. The reference section information is section information generated in advance and is stored in advance in the storage unit 11.
[0041] The selection unit 121 selects the reference section information based on, for example, the content specified by an analyst on the information terminal 3. The analyst is one of the multiple participants taking part in the discussion to be analyzed, or a person different from the multiple participants.
[0042] 3(a) and 3(b) are schematic diagrams illustrating a method by which the selection unit 121 selects reference section information. In the example of FIG. 3(a), the storage unit 11 pre-stores discussion information indicating at least one of the attributes of past discussions (subject, type, time, topic, format, etc.), the date and time the discussion was held, and the attributes of the participants in the discussion (teacher, grade, etc.), in association with section information generated by the generation unit 124 using a method described below. The selection unit 121 accepts, for example, on the information terminal 3, designation of search conditions for past discussions by an analyst. The search conditions are, for example, at least one of the attributes of the discussion, the date and time the discussion was held, and the attributes of the participants. The selection unit 121 extracts discussion information that matches the designated search conditions from the storage unit 11 and displays the extracted discussion information on the information terminal 3 together with section information 31 associated with each of the extracted one or more pieces of discussion information.
[0043] The selection unit 121 then accepts the analyst's designation of any one of the section information 31 on the information terminal 3, and selects the designated section information 31 as the reference section information. This allows the speech analysis device 1 to compare the discussion to be analyzed with discussions that match the search criteria designated by the analyst.
[0044] In the example of FIG. 3(b), the storage unit 11 stores in advance a template of section information. The template of section information is information indicating the order of, for example, a section in which the speech of the first group is dominant, a section in which the speech of the second group is dominant, and a section in which the speech of the first group and the second group are in balance. The selection unit 121 causes the information terminal 3 to display a plurality of templates 32 stored in the storage unit 11. FIG. 3(b) shows an example in which a plurality of participants are classified into a first group, a T (instructor) group, and a second group, an S (learner) group, but the plurality of participants may also be classified into two groups based on other criteria.
[0045] The selection unit 121 then accepts the analyst's designation of one of the templates 32 at the information terminal 3, and selects the order of the sections indicated by the designated template 32 as the reference section information. The selection unit 121 may also accept input at the information terminal 3 of the order of the section where the first group's speech is dominant, the section where the second group's speech is dominant, and the section where the first and second groups' speeches are in balance, and select the input order of the sections as the reference section information. This allows the speech analysis device 1 to compare the discussion to be analyzed with the order of the sections designated by the analyst.
[0046] The classification unit 122 classifies multiple participants taking part in the discussion to be analyzed into a first group or a second group. The classification unit 122 may, for example, accept a setting by an analyst on the information terminal 3 as to whether each of the multiple participants belongs to the first group or the second group. In this case, the classification unit 122 stores information indicating whether each of the multiple participants belongs to the first group or the second group in the storage unit 11, according to the content set on the information terminal 3.
[0047] Furthermore, the classification unit 122 may automatically classify the multiple participants into a first group or a second group based on, for example, the attributes of each of the multiple participants. In this case, the storage unit 11 stores information indicating the attributes of each of the multiple participants in advance. The attribute used for classification is, for example, the role of the participant (instructor, learner, etc.). For example, if the attributes of the participant satisfy a predetermined condition, the classification unit 122 classifies the participant into the first group, and if not, classifies the participant into the second group. The classification unit 122 stores information indicating whether each of the multiple participants belongs to the first group or the second group in the storage unit 11 according to the classification result.
[0048] The acquisition unit 123 acquires the voices uttered by multiple participants in a discussion from the sound collection device 2. The acquisition unit 123 acquires the voices of a part of the discussion at predetermined time intervals during the discussion, or acquires the voices of the entire discussion after the discussion has ended.
[0049] The acquisition unit 123 identifies the speech period of each of the multiple participants based on the sound acquired from the sound collection device 2. In the case of a discussion taking place around the sound collection device 2 equipped with a microphone array, the acquisition unit 123 performs known sound source localization on the multiple channel sounds received from the sound collection device 2, for example. Sound source localization is a process of estimating the direction of the sound source included in the sound acquired by the acquisition unit 123 for each time period (for example, every 10 to 100 milliseconds). The acquisition unit 123 associates the direction of the sound source estimated for each time period with the direction of each of the multiple participants that is preset in the information terminal 3.
[0050] The acquisition unit 123 can use other sound source localization methods such as the MUSIC (Multiple Signal Classification) method and the beamforming method, as long as it is possible to identify the direction of the sound source based on the acquired sound.
[0051] Next, the acquisition unit 123 determines which participant spoke (made a statement) during the discussion at predetermined time intervals (for example, every 10 to 100 milliseconds) based on the acquired voice and the estimated direction of the sound source. The acquisition unit 123 identifies a continuous period from when one participant starts speaking to when he or she finishes speaking as a speech period. When multiple participants speak at the same time, at least a portion of the speech periods of the multiple participants may overlap.
[0052] In the case of a discussion taking place over a network, the acquisition unit 123, for example, estimates a participant associated with the sound collection device 2 that is the sender of the acquired audio as the sound source, and identifies the speech period of each of the multiple participants based on the acquired audio and the estimated sound source.
[0053] The acquisition unit 123 is not limited to using the specific method shown here, and may identify the speech period of each of the multiple participants using other methods.
[0054] The acquiring unit 123 acquires time-series information indicating the speech situations of the first group and the second group by time based on the speech periods of each of the identified participants. Fig. 4 is a schematic diagram for explaining a method for the acquiring unit 123 to acquire the time-series information.
[0055] The acquiring unit 123 calculates the amount of speech of each of the multiple participants based on the identified speech period. For example, the acquiring unit 123 calculates, for each predetermined time frame (5 seconds, 10 seconds, 30 seconds, etc.), a value corresponding to the length of the speech period of the participant within the time frame as the amount of speech. Instead of or in addition to the length of the speech period, the acquiring unit 123 may calculate, as the amount of speech, a value corresponding to the number of speeches or the speech volume. The acquiring unit 123 calculates, for each of the multiple participants, the amount of speech for each time frame during the period from the start point (start time) to the end point (end time) of the discussion.
[0056] The acquiring unit 123 classifies the amount of speech of each of the multiple participants for each time frame into a first group or a second group according to the classification result of the multiple participants by the classifying unit 122. In the example of Fig. 4, the multiple participants are classified into a T (instructor) group, which is the first group, and an S (learner) group, which is the second group. The acquiring unit 123 calculates statistical values (average, median, etc.) of the amount of speech of the participants belonging to each of the first and second groups for each time frame.
[0057] The acquisition unit 123 determines, for each time frame, which of the first group and the second group has a larger statistical value of the amount of speech. According to the determination result, the acquisition unit 123 acquires, as time-series information, information indicating which of the first group and the second group has a larger amount of speech for each predetermined time frame during the period from the start point to the end point of the discussion. If, in the time-series information, there are multiple consecutive time frames in which the first group and the second group have the same larger amount of speech, the acquisition unit 123 may integrate the multiple time frames.
[0058] The generation unit 124 determines multiple sections that constitute the discussion to be analyzed based on the time-series information acquired by the acquisition unit 123, and determines a section tendency indicating which of the utterances from the first group or the second group is dominant for each section. The generation unit 124 may determine the sections and section tendency for a part of the discussion during the discussion, or may determine the sections and section tendency for the entire discussion after the discussion has ended.
[0059] 5 is a schematic diagram for explaining a method for determining a section tendency by the generation unit 124. The generation unit 124 generates a transition graph showing the transition of whether the speech volume of the first group or the second group is larger, based on the time-series information acquired by the acquisition unit 123. FIG. 5 shows an example in which the speech analysis system S according to this embodiment is applied to the generation of a graph of ST analysis described in Non-Patent Document 1.
[0060] In the transition graph illustrated in Fig. 5, the horizontal axis represents the time of the first group (T group), and the vertical axis represents the time of the second group (S group). The acquisition unit 123 sets the origin as the starting point (starting point of the discussion), and draws a line to the right along the horizontal axis for periods in the time-series information when the amount of speech of the first group is greater, and draws a line up along the vertical axis for periods in the time-series information when the amount of speech of the second group is greater. The acquisition unit 123 repeats this process from the start to the end of the time-series information, that is, from the start to the end of the discussion, to generate a transition graph showing the transition of whether the amount of speech of the first group or the second group is greater.
[0061] The generation unit 124 uses the generated transition graph to divide the discussion to be analyzed into multiple sections, and determines a section tendency indicating which of the first and second groups is the dominant utterance for each section. First, the generation unit 124 sets the origin (the start point of the transition graph) as the start point of the section. The generation unit 124 extracts one predetermined period (e.g., 5 seconds) in chronological order from the transition graph as a unit of interest.
[0062] The generation unit 124 determines whether the elapsed time from the start point of the interval to the end point of the unit of interest is equal to or greater than a predetermined time. The predetermined time is a value set in advance as the minimum duration of the interval, such as 5 minutes or 10 minutes. The predetermined time may also be determined according to the reference interval information selected by the selection unit 121. If the elapsed time from the start point of the interval to the end point of the unit of interest is not equal to or greater than the predetermined time, the generation unit 124 extracts the next predetermined period as the unit of interest and repeats the determination of whether the elapsed time from the start point of the interval to the end point of the unit of interest is equal to or greater than the predetermined time.
[0063] If the elapsed time from the start point of the section to the end point of the unit of interest is equal to or longer than a predetermined time, the generation unit 124 determines the section tendency of the section by comparing the speech situations of the first and second groups. To compare the speech situations, the generation unit 124 calculates, for example, the slope (dashed line in FIG. 5) between the coordinates of the start point of the section and the coordinates of the end point of the unit of interest on the transition graph.
[0064] The generation unit 124 determines the section tendency based on the calculated slope. For example, if the slope is equal to or less than a first reference value, the generation unit 124 determines that the speech of the first group is dominant. For example, if the slope is greater than the first reference value and equal to or less than a second reference value, the generation unit 124 determines that the speech of the first group and the second group are in balance. For example, if the slope is greater than the second reference value, the generation unit 124 determines that the speech of the second group is dominant. The first reference value and the second reference value are stored in advance in the storage unit 11 or set in the information terminal 3. The generation unit 124 determines the determination result as the section tendency of the section.
[0065] If the section trend of the previous section is the same as the section trend of the current section, the generation unit 124 combines the unit of interest with the previous section, extracts the next specified period as the unit of interest, and repeatedly determines whether the elapsed time from the start of the section to the end of the unit of interest is greater than or equal to a specified time.
[0066] If the section trend of the previous section differs from the section trend of the current section, the generation unit 124 determines the previous section and section trend. The generation unit 124 sets the start point of the unit of interest as the start point of the section, and repeats the above-described process to determine the section and section trend up to the end point of the transition graph.
[0067] The generation unit 124 is not limited to the specific method shown here, and may use other methods to determine multiple sections that make up the discussion based on time series information, and may determine a section trend for each section that indicates which of the first and second groups' utterances is dominant, or that the utterances of the first and second groups are in balance.
[0068] The generation unit 124 generates section information that associates each of the multiple sections that make up the discussion with a section tendency indicating which of the first and second groups is the dominant speech in that section, and stores the information in the storage unit 11. In the section information, the section tendency in a part of the discussion may indicate that the speech of the first and second groups is in a balanced relationship (i.e., neither the first nor second group is dominant).
[0069] In this way, the generation unit 124 divides the discussion to be analyzed into multiple sections and determines a section trend indicating which of the first and second groups is the dominant utterance for each section. Because the time-series information precisely represents the relative magnitude of the utterances of the two groups at each time point, it is difficult for an analyst to analyze the utterance trends of the two groups simply by looking at the time-series information. In contrast, by dividing a period in which the same utterance trend continues in a discussion into a single section, the generation unit 124 makes it easier for the analyst to grasp the transition of the utterance trend throughout the discussion, making it easier to analyze the discussion between the two groups.
[0070] The output unit 125 outputs the section information generated by the generation unit 124 at least either during the discussion or after the discussion has ended. Hereinafter, the output of section information by the output unit 125 during the discussion will be referred to as real-time output, and the output of section information by the output unit 125 after the discussion has ended will be referred to as post-output.
[0071] 6(a) and 6(b) are schematic diagrams for explaining a method for real-time output of section information by the output unit 125. The output unit 125 controls information corresponding to the section information generated by the generation unit 124 to be displayed on a display unit included in the information terminal 3 as shown in Fig. 6(a) or to be output from a sound output unit included in the sound collection device 2 as shown in Fig. 6(b).
[0072] The output unit 125 transmits, for example, information corresponding to the time-series information acquired by the acquisition unit 123 and the section information generated by the generation unit 124 to the information terminal 3. In the example of Fig. 6(a), the output unit 125 causes the information terminal 3 to display a transition graph 33 indicating the transition of which of the first group and the second group has a larger speech volume, corresponding to the time-series information acquired by the acquisition unit 123. The output unit 125 may display the section trends on the transition graph 33 by generating the transition graph 33 using lines of colors according to the section trends for each section.
[0073] Furthermore, the output unit 125 causes the information terminal 3 to display a bar graph 34 indicating the length of the section and the section tendency of the section, which corresponds to the section information generated by the generation unit 124. This enables the speech analysis device 1 to make it easier for the analyst to grasp the speech tendency of the two groups during the discussion, and to analyze the speech tendency in the discussion between the two groups.
[0074] Furthermore, the output unit 125 displays on the information terminal 3 a transition graph 33 and a bar graph 34 corresponding to the reference section information selected by the selection unit 121. This allows the speech analysis device 1 to easily compare the section information of the discussion to be analyzed with the reference section information specified by the analyst.
[0075] Furthermore, the output unit 125 transmits, for example, information corresponding to the difference between the section information generated by the generation unit 124 and the reference section information selected by the selection unit 121 as the comparison result to the information terminal 3 or the sound collection device 2. The output unit 125 outputs, as the difference between the section information and the reference section information, whether or not the difference in the length, number, order, etc. of the segments satisfies a predetermined condition.
[0076] In the example of Fig. 6(a), the output unit 125 causes the information terminal 3 to display a message 35 indicating that the length of the segment of the first group (T group) in the section information is longer than the length of the segment of the first group in the reference section information. In the example of Fig. 6(b), the output unit 125 causes the sound collection device 2 to output a sound indicating that the length of the segment of the second group (S group) in the section information is longer than the length of the segment of the second group in the reference section information. This allows the sound analysis device 1 to successively notify the analyst of the difference between the section information of the discussion to be analyzed and the reference section information specified by the analyst, making it easier to reflect this in the ongoing discussion.
[0077] 7 is a schematic diagram for explaining a method for post-outputting section information by the output unit 125. The output unit 125 performs control to display information corresponding to the section information generated by the generation unit 124 on a display unit provided in the information terminal 3 as shown in FIG.
[0078] The output unit 125 transmits, for example, information corresponding to the time-series information acquired by the acquisition unit 123 and the section information generated by the generation unit 124 to the information terminal 3. In the example of Fig. 7, the output unit 125 causes the information terminal 3 to display a transition graph 36 indicating the transition of which of the first group and the second group has a larger speech volume, corresponding to the time-series information acquired by the acquisition unit 123. The output unit 125 may display the section trends on the transition graph 36 by generating the transition graph 36 using lines of colors according to the section trends for each section.
[0079] Furthermore, the output unit 125 causes the information terminal 3 to display a bar graph 37 indicating the length of the section and the section tendency of the section, which corresponds to the section information generated by the generation unit 124. This enables the speech analysis device 1 to make it easier for the analyst to grasp the speech tendency of the two groups in the entire discussion, and to analyze the speech tendency in the discussion between the two groups.
[0080] Furthermore, the output unit 125 displays on the information terminal 3 a transition graph 36 and a bar graph 37 corresponding to the reference section information selected by the selection unit 121. This allows the speech analysis device 1 to easily compare the section information of the discussion to be analyzed with the reference section information specified by the analyst.
[0081] In addition, the output unit 125 transmits to the information terminal 3, for example, information that associates the section trends of each of the multiple sections indicated by the section information generated by the generation unit 124 with the section trends of each of the multiple sections indicated by the reference section information selected by the selection unit 121.
[0082] 7, the output unit 125 displays, on the information terminal 3, section trend transitions 38 of multiple sections for each of the section information generated by the generation unit 124 and the reference section information selected by the selection unit 121. The output unit 125 detects, for example, by dynamic programming, the correspondence between the sections indicated by the section information generated by the generation unit 124 and the sections indicated by the reference section information selected by the selection unit 121, and displays increases, decreases, differences in order, etc. of the sections as section trend transitions 38. This allows the speech analysis device 1 to easily distinguish the differences between the section trend transitions in the section information of the discussion to be analyzed and the section trend transitions in the reference section information specified by the analyst.
[0083] Furthermore, the output unit 125 may output the words (phrases) uttered in each of the multiple sections determined by the generation unit 124 in association with the section. In this case, the output unit 125 extracts words included in the speech in each of the multiple sections, for example, by performing a known speech recognition process on the speech in each of the multiple sections. The output unit 125, for example, associates each of the multiple sections with some or all of the words extracted for that section and displays them on the information terminal 3. This allows the speech analysis device 1 to make it easier for the analyst to understand the content of each of the multiple sections.
[0084] Furthermore, the output unit 125 may output the characteristics of the entire discussion based on the section information generated by the generation unit 124. In this case, the output unit 125 determines the characteristics of the entire discussion based on the section trends of the multiple sections indicated by the section information generated by the generation unit 124. The output unit 125 determines the characteristics of the entire discussion based on, for example, the proportion of sections in which the utterance of the first group is dominant, sections in which the utterance of the second group is dominant, and sections in which the utterances of the first group and the second group are in balance among the multiple sections that make up the discussion.
[0085] For example, the output unit 125 determines that the discussion is lecture-based when the proportion of sections where the speech of group T (instructor group) is dominant is equal to or greater than a predetermined value, and determines that the discussion is exercise-based when the proportion of sections where the speech of group S (learner group) is dominant is equal to or greater than a predetermined value. The output unit 125 displays a message (report) indicating the determined characteristics of the entire discussion on the information terminal 3. This allows the speech analysis device 1 to make it easier for the analyst to grasp the overall trend of the discussion determined based on the section trends of multiple sections.
[0086] [Flowchart of voice analysis method] 8 is a flowchart of an exemplary speech analysis method executed by the speech analysis device 1 according to this embodiment. The selection unit 121 selects reference section information to be compared with section information (S11). The classification unit 122 classifies multiple participants taking part in the discussion to be analyzed into a first group or a second group (S12). The classification unit 122, for example, accepts a setting on the information terminal 3 as to whether the multiple participants belong to the first group or the second group, or automatically classifies the multiple participants into the first group or the second group based on the attributes of the multiple participants.
[0087] The subsequent processes are performed sequentially during the discussion or after the discussion has ended. The acquisition unit 123 acquires the voices uttered by the multiple participants in the discussion from the sound collection device 2. The acquisition unit 123 identifies the speech period of each of the multiple participants based on the voices acquired from the sound collection device 2 (S13). The acquisition unit 123 acquires time-series information indicating the speech status of each of the first and second groups by time based on the speech period of each of the identified multiple participants (S14).
[0088] The generation unit 124 performs a section information generation process (S2) to generate section information that associates each of the multiple sections that make up the discussion with a section tendency that indicates which of the utterances from the first group and the second group is dominant in that section, based on the time-series information acquired by the acquisition unit 123. The section information generation process of step S2 will be described later with reference to FIG.
[0089] The output unit 125 outputs the section information generated by the generation unit 124 to at least one of the sound collection device 2 and the information terminal 3 (S15).
[0090] 9 is a diagram showing a flowchart of a section information generation process in an exemplary speech analysis method executed by the speech analysis device 1 according to this embodiment. The generation unit 124 generates a transition graph showing the transition of which of the first and second groups has a larger speech volume, based on the time-series information acquired by the acquisition unit 123 (S21). The generation unit 124 sets the origin (starting point of the transition graph) to the starting point of the section (S22). The generation unit 124 extracts one predetermined period (e.g., 5 seconds) in chronological order from the transition graph as a unit of interest (S23).
[0091] The generation unit 124 determines whether the elapsed time from the start point of the interval to the end point of the unit of interest is equal to or greater than a predetermined time (S24). If the elapsed time from the start point of the interval to the end point of the unit of interest is not equal to or greater than the predetermined time (NO in S25), the generation unit 124 returns to step S23 and repeats the process for the next unit of interest.
[0092] If the elapsed time from the start point of the interval to the end point of the unit of interest is equal to or longer than the predetermined time (YES in S25), the generation unit 124 calculates the gradient between the coordinates of the start point of the interval and the coordinates of the end point of the unit of interest on the transition graph (S26).
[0093] The generation unit 124 determines the section tendency based on the calculated slope (S27). For example, if the slope is equal to or less than a first reference value, the generation unit 124 determines that the speech of the first group is dominant. For example, if the slope is greater than the first reference value and equal to or less than a second reference value, the generation unit 124 determines that the speech of the first group and the second group are in balance. For example, if the slope is greater than the second reference value, the generation unit 124 determines that the speech of the second group is dominant. The generation unit 124 determines the determination result as the section tendency of the section.
[0094] If the section trend of the previous section and the section trend of the current section are the same (YES in S28), the generation unit 124 combines the unit of interest with the previous section (S29). The generation unit 124 returns to step S23 and repeats the process for the next unit of interest.
[0095] If the section trend of the previous section and the section trend of the current section are different (NO in S28), the generation unit 124 determines the previous section and section trend (S30). If the time-series information has not ended (NO in S31), the generation unit 124 sets the start point of the unit of interest to the start point of the section (S32). The generation unit 124 returns to step S23 and repeats the process for the next unit of interest.
[0096] When the time series information has ended (YES in S31), the generation unit 124 generates section information that associates each of the multiple sections that make up the discussion with a section tendency indicating which of the first and second groups is the dominant utterance in that section, for all or part of the discussion, and stores the information in the storage unit 11. In the section information, the section tendency for part of the discussion may indicate that the utterances of the first and second groups are in opposition to each other.
[0097] [Effects of this embodiment] According to the speech analysis system S of this embodiment, the speech analysis device 1 determines a section tendency indicating which of the two groups is the dominant speaker for each section of the discussion based on the audio of the discussion, and notifies the analyst of the section and the section tendency in association with each other. This makes it easier for the analyst to grasp the speech tendencies of the two groups and to analyze the speech tendencies in the discussion between the two groups.
[0098] [First Modification] In some cases, the grouping of multiple participants may be changed during a discussion. In this modification, the speech analysis device 1 changes the participants belonging to the first group and the second group between multiple periods in the discussion, and generates section information based on the changed grouping.
[0099] The classification unit 122 changes the participants belonging to the first group and the second group between multiple periods in the discussion. For example, the classification unit 122 may receive a setting of the discussion structure (explanation period, practice period, etc.) in advance in the information terminal 3, and change the participants belonging to the first group and the second group when the structure changes. For example, during the explanation period, the classification unit 122 places the teacher in the first group and the students in the second group, while during the practice period, the classification unit 122 places some of the students in the first group and some of the other students in the second group.
[0100] Furthermore, the classification unit 122 may detect the arrangement of each of the multiple participants by performing a known image recognition process on a captured image acquired by a camera or the like, and change the participants belonging to the first group and the second group when the arrangement changes. For example, the classification unit 122 may classify participants sitting in a specific seat as the first group and participants sitting in other seats as the second group.
[0101] Furthermore, the classification unit 122 may change the parent group including the first group and the second group according to the grouping that differs for each period. That is, for each of multiple periods in the discussion, the classification unit 122 generates a first parent group including the first group and the second group into which some of the multiple participants are classified, and a second parent group including the first group and the second group into which some of the multiple participants who do not belong to the first parent group are classified. The classification unit 122 may generate three or more parent groups.
[0102] Furthermore, the classification unit 122 may generate a parent group including all of the participants (for example, the entire classroom) during at least one period of the discussion. That is, the classification unit 122 may generate a first parent group including a first group and a second group into which some of the participants are classified, and a second parent group including the first group and the second group into which some of the participants who do not belong to the first parent group are classified, and may further generate a third parent group including the first group and the second group into which the participants who belong to the first parent group and the second parent group are classified.
[0103] The generation unit 124 generates section information for each of the multiple parent groups. The output unit 125 outputs the section information for each of the multiple parent groups. Fig. 10 is a schematic diagram for explaining a method for the output unit 125 to output the section information in this modification. The output unit 125 performs control to display information corresponding to the section information for each of the multiple parent groups generated by the generation unit 124 on a display unit provided in the information terminal 3, as shown in Fig. 10.
[0104] The output unit 125 outputs, for example, the section information of the first parent group and the section information of the second parent group simultaneously. In the example of Fig. 10, the output unit 125 displays section information of three parent groups including the first parent group and the second parent group side by side for a period from 10 minutes to 50 minutes. This allows the voice analysis device 1 to allow the analyst to get an overview of the section information of multiple parent groups (multiple tables, etc.).
[0105] The output unit 125 also outputs, for example, section information for at least one of the first and second parent groups, and section information for the third parent group. In the example of FIG. 10, the output unit 125 displays section information for three parent groups, including the first and second parent groups, side by side for the period from 10 to 50 minutes, and further displays section information for the third parent group corresponding to all participants for the period from 0 to 10 minutes and the period from 50 to 60 minutes. This allows the speech analysis device 1 to provide the analyst with section information corresponding to different groupings for each period when the grouping of multiple participants is changed during a discussion. The speech analysis device 1 also facilitates hierarchical analysis of a parent group corresponding to all participants (such as an entire classroom) and multiple parent groups corresponding to groups into which multiple participants are divided (such as multiple tables).
[0106] [Second Modification] During a discussion, a specific participant, such as an instructor, may move between multiple tables and join a different group. In this modification, the speech analysis device 1 changes the grouping based on the location of the specific participant, and generates section information based on the changed grouping.
[0107] FIG. 11 is a schematic diagram illustrating a method in which the generation unit 124 generates section information in this modification. The classification unit 122 estimates the position of the instructor, who is a specific participant, during the discussion. The classification unit 122 may estimate which sound collection device 2 the instructor is closer to by, for example, comparing the voice acquired by the acquisition unit 123 from multiple sound collection devices 2 with pre-registered voice features (such as a voiceprint) of the instructor. The classification unit 122 may estimate which sound collection device 2 the instructor is closer to based on the strength of short-range wireless communication performed between a communication device (such as a smartphone) held by the instructor and multiple sound collection devices 2.
[0108] The classification unit 122 is not limited to the specific method shown here, and may estimate the position of the instructor during the discussion using other methods. The example in Figure 11 shows that the instructor moved to table 1, table 2, and table 3 in that order.
[0109] The classification unit 122 changes the participants belonging to the first group to which the instructor belongs and the participants belonging to the second group between multiple periods in the discussion based on the estimated position of the instructor. In the example of Fig. 11, the classification unit 122 generates a first group including the instructor and a second group including the students at table 1 during a period when the instructor is located at table 1. Furthermore, the classification unit 122 generates a first group including the instructor and a second group including the students at table 2 during a period when the instructor is located at table 2. Furthermore, the classification unit 122 generates a first group including the instructor and a second group including the students at table 3 during a period when the instructor is located at table 3.
[0110] The generation unit 124 generates section information for each of the multiple periods in which the grouping was changed, and combines the generated section information. As a result, when a specific participant such as an instructor moves during a discussion, the speech analysis device 1 generates section information according to the grouping that was changed according to the position of the specific participant, making it easier to analyze the tendency of speech centered on the specific participant.
[0111] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments.
[0112] The processor of the speech analysis device 1 is responsible for each step (process) included in the speech analysis method shown in Figures 8 and 9. That is, the processor of the speech analysis device 1 reads a program for executing the speech analysis method shown in Figures 8 and 9 from the storage unit 11, and executes the program to control each unit of the speech analysis device 1, thereby executing the speech analysis method shown in Figures 8 and 9. Some of the steps included in the speech analysis method shown in Figures 8 and 9 may be omitted, the order of the steps may be changed, or multiple steps may be performed in parallel. [Explanation of symbols]
[0113] S Voice Analysis System 1. Voice analysis device 11 Storage section 12 Control Unit 121 Selection Section 122 Classification Department 123 Acquisition Department 124 Generation part 125 Output section 2 Sound collection device 3. Information terminals
Claims
1. a classification unit that classifies a plurality of participants into a first group and a second group; an acquisition unit that acquires time-series information indicating the speech situations of the first group and the second group in the speeches uttered by the participants belonging to the first group and the participants belonging to the second group in the discussion; A generation unit that generates section information that associates each of a plurality of sections that constitute the discussion with a section tendency that indicates which of the utterances of the first group and the second group is dominant in the section based on the time series information, for all or part of the discussion; an output unit that outputs the section information; and the classification unit changes, between the first period and the second period, the participants belonging to the first group to which the specific participant belongs and the participants belonging to the second group to which the specific participant does not belong, based on a position of the specific participant in a first period and a position of the specific participant in a second period different from the first period. Voice analysis device.
2. The section tendency indicates which of the first group and the second group is dominant, or indicates that the first group and the second group are competitive with each other. The speech analysis device according to claim 1 .
3. the generation unit determines each of the plurality of sections so that the section has a predetermined time or more, and determines the section tendency for the section by comparing the speech situations of the first group and the second group in the section. The voice analysis device according to claim 1 or 2.
4. The time series information is information indicating which of the first group and the second group has a larger amount of speech for each predetermined time frame during a period from the start point to the end point of the discussion. The speech analysis device according to claim 1 .
5. the classification unit generates a first parent group including the first group and the second group into which a portion of the plurality of participants are classified, and a second parent group including the first group and the second group into which a portion of the plurality of participants who do not belong to the first parent group are classified; the output unit simultaneously outputs the section information of the first parent group and the section information of the second parent group. The speech analysis device according to claim 1 .
6. the classification unit generates a first parent group including the first group and the second group into which a portion of the plurality of participants are classified, and a second parent group including the first group and the second group into which a portion of the plurality of participants who do not belong to the first parent group are classified, and further generates a third parent group including the first group and the second group into which participants belonging to the first parent group and the second parent group are classified; the output unit outputs the section information of at least one of the first parent group and the second parent group, and the section information of the third parent group. The speech analysis device according to claim 1 .
7. the output unit outputs words included in the utterance of each of the plurality of sections, which are extracted by performing a speech recognition process on the speech, in association with the section. The speech analysis device according to claim 1 .
8. The output unit outputs a feature of the entire discussion based on the section trends of the plurality of sections constituting the discussion. The speech analysis device according to any one of claims 1 to 7.
9. a selection unit for selecting reference section information to be compared with the section information; the output unit outputs a comparison result between the section information and the reference section information. The speech analysis device according to any one of claims 1 to 8.
10. the output unit outputs, during the discussion, information corresponding to a difference between the section information and the reference section information as the comparison result. The speech analysis device according to claim 9 .
11. the output unit, after the discussion, outputs the section trend of each of the plurality of sections indicated by the section information in association with the section trend of each of the plurality of sections indicated by the reference section information. The speech analysis device according to claim 9 or 10.
12. The plurality of participants conduct the discussion in a plurality of groups; the classification unit generates, for each of the first period and the second period, the first group including the specific participant and the second group including participants of a group that is closest to the specific participant among the plurality of groups. The speech analysis device according to any one of claims 1 to 11.
13. The processor executes classifying a plurality of participants into a first group and a second group; acquiring time-series information indicating the speech situations of the first group and the second group over time in the speeches uttered by the participants belonging to the first group and the participants belonging to the second group in the discussion; generating section information that associates, based on the time-series information, each of a plurality of sections constituting the discussion with a section tendency indicating which of the first group and the second group is the dominant utterance in the section, for all or part of the discussion; outputting the section information; and In the classifying step, between the first period and the second period, participants belonging to the first group to which the specific participant belongs and participants belonging to the second group to which the specific participant does not belong are changed based on a position of the specific participant in a first period and a position of the specific participant in a second period different from the first period. Voice analysis methods.
Citation Information
Patent Citations
Device, method, and program for assisting activation of conference, and recording medium
JP2006208482A