Information transmission device, information transmission method, and program
The information transmission device addresses the challenge of identifying and summarizing content of interest in online seminars by receiving learner reactions and generating summary data, enhancing engagement and feedback for instructors.
Patent Information
- Application Number
- JP2024001254
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing systems fail to provide real-time insights into the content that seminar attendees find interesting during online seminars, limiting the ability of individuals other than the lecturer to grasp the content discussed when attendees show interest.
An information transmission device that distributes audio content, receives reaction information from learners, specifies time regions of interest, generates summary data, and outputs this data to relevant parties, utilizing a large language model to summarize the content.
Facilitates easier understanding of the content that learners find interesting during audio distribution, enabling better engagement and feedback for instructors and administrators.
Smart Images

Figure 2025107804000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information transmission device, an information transmission method, and a program for distributing content via a network.
Background Art
[0002] Distribution of audio content such as online seminars is widely performed. Patent Document 1 describes a web server that acquires reaction information indicating the reaction of a seminar attendee to a seminar and provides the acquired reaction information to the lecturer's terminal at any time.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the web server described in Patent Document 1, since the reaction information is provided to the lecturer at any time, the lecturer can know the reaction of the attendee during the seminar and provide a lecture of content corresponding to the reaction of the attendee. The content that the lecturer was talking about when the attendee showed interest is considered important, but seminar-related persons other than the lecturer could not know the content of the lecture at the time when the attendee showed interest after the seminar ended.
[0005] Therefore, the present invention has been made in view of these points, and an object thereof is to provide an information transmission device that can make it easy for a person related to the distribution of audio content to grasp the content that the lecturer was talking about when the attendee of the audio content showed interest.
Means for Solving the Problems
[0006] The information transmission device according to the first aspect of the present invention includes a distribution unit that distributes audio content to a learner terminal, a reception unit that receives reaction information indicating that a learner has operated a reaction button while the distribution unit is distributing the audio content, from the learner terminal, a specifying unit that specifies, in the audio content, a time region that includes the timing at which the reaction information was received and is shorter than the total length of the audio content, a generation unit that generates summary data obtained by summarizing the audio content corresponding to the time region specified by the specifying unit, and an output unit that outputs the summary data.
[0007] The generation unit may generate the summary data by converting the audio content into text data, specifying the text data corresponding to the time region, and transmitting the specified text data to an external device having a large language model that creates a summary sentence of the input sentence.
[0008] The distribution unit distributes the audio content to a plurality of the learner terminals, the specifying unit specifies, in the audio content, one or more time regions in which the number of the learner terminals that the reception unit has received the reaction information from is relatively large, and the generation unit may generate the summary data obtained by summarizing one or more pieces of the audio content corresponding to the one or more time regions.
[0009] The information transmission device further includes an acquisition unit that acquires the attributes of the plurality of learners. The distribution unit distributes the voice content to the learner terminals of the plurality of learners. The specifying unit specifies, in the voice content, one or more first attribute time regions in which the number of learner terminals that have received the reaction information from the learner terminals of the learners with the first attribute is relatively large, and specifies, in the voice content, one or more second attribute time regions in which the number of learner terminals that have received the reaction information from the learner terminals of the learners with the second attribute is relatively large. The generation unit generates one or more first summary data obtained by summarizing the voice content corresponding to the one or more first attribute time regions, and generates one or more second summary data obtained by summarizing the voice content corresponding to the one or more second attribute time regions. The output unit may output the one or more first summary data in association with the first attribute, and output the one or more second summary data in association with the second attribute.
[0010] The reception unit receives, as the reaction information, positive reaction information indicating that a learner has operated a positive reaction button for indicating a positive reaction to the voice content, and reconfirmation reaction information indicating that a learner has operated a reconfirmation reaction button for indicating that the voice content is difficult to hear. The output unit may transmit the summary data corresponding to the positive reaction information to the learner terminal on which the positive reaction button has been operated, and transmit the summary data corresponding to the reconfirmation reaction information to the learner terminal on which the reconfirmation reaction button has been operated.
[0011] When the specific unit receives the positive reaction information, it specifies a first reaction time region including the timing when the positive reaction information was received in the voice content. When the reconfirmation reaction information is received, it specifies a second reaction time region that includes the timing when the reconfirmation reaction information was received and starts at a timing earlier than the first reaction time region in the voice content. When the positive reaction information is received, the output unit may transmit the summary data corresponding to the first reaction time region to the learner terminal on which the positive reaction button was operated. When the reconfirmation reaction information is received, the output unit may transmit the summary data corresponding to the second reaction time region to the learner terminal on which the reconfirmation reaction button was operated.
[0012] The distribution unit distributes the voice content to a plurality of the learner terminals. The specific unit specifies one or more first majority reaction time regions in the voice content where the number of learner terminals that the reception unit has received the positive reaction information from is relatively large. Regardless of the number of learner terminals that the reception unit has received the reconfirmation reaction information from, the specific unit specifies one or more second reaction time regions in the voice content where the reception unit has received the reconfirmation reaction information. The output unit may transmit the summary data obtained by summarizing the voice content corresponding to the one or more first majority reaction time regions to the distribution-side terminal managed by the instructor or administrator of the voice content, and transmit the summary data obtained by summarizing the voice content corresponding to the one or more second reaction time regions to the learner terminal that is the transmission source of the reconfirmation reaction information.
[0013] When the reception unit receives the reaction information exceeding a predetermined number of times from the same learner terminal within a reference time, the reception unit may invalidate the reaction information received exceeding the predetermined number of times. The output unit may transmit the summary data to the learner terminal on which the reaction button was operated.
[0014] The information transmission method according to the second aspect of the present invention includes steps that a computer executes: delivering voice content to a learner terminal; receiving reaction information indicating that a learner has operated a reaction button during the delivery of the voice content from the learner terminal; specifying a time region that includes the timing when the reaction information is received and is shorter than the entire length of the voice content; generating summary data obtained by summarizing the voice content corresponding to the specified time region; and outputting the summary data.
[0015] The program according to the third aspect of the present invention causes a computer to execute steps of: delivering voice content to a learner terminal; receiving reaction information indicating that a learner has operated a reaction button during the delivery of the voice content from the learner terminal; specifying a time region that includes the timing when the reaction information is received and is shorter than the entire length of the voice content; generating summary data obtained by summarizing the voice content corresponding to the specified time region; and outputting the summary data.
Advantages of the Invention
[0016] According to the present invention, there is an effect that it becomes easier for the parties related to the delivery of voice content to grasp the content that the lecturer was speaking at the time when the learner of the voice content showed interest.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0018] [Outline of Content Delivery System S] FIG. 1 is a diagram showing the configuration of a content delivery system S according to the present embodiment. The content delivery system S includes a plurality of trainee terminals 1a to 1c, a delivery-side terminal 2, an information transmission device 3, and an external device 4.
[0019] The plurality of trainee terminals 1a, 1b, and 1c (hereinafter also referred to as "trainee terminal 1") are computers used by trainees who receive voice content. The trainee terminal 1 communicates with the information transmission device 3 via a network.
[0020] The trainee terminal 1 receives the voice content distributed from the information transmission device 3. The voice content is content including voice data provided, for example, in a webinar or an online meeting in which a lecture by a lecturer is distributed online. The trainee terminal 1 may receive video content together with the voice content.
[0021] The recipient terminal 1 accepts operations by the recipient on the reaction buttons while the voice content is being distributed. For example, when a touch panel is provided on the display of the recipient terminal 1, the recipient terminal 1 accepts the recipient's operations on the reaction buttons displayed on the display. The reaction buttons are a user interface that the recipient operates when the recipient is interested in the topic indicated by the voice content or when the recipient feels that the voice content is difficult to hear.
[0022] The distribution-side terminal 2 is a computer managed by the lecturer or administrator of the voice content. The distribution-side terminal 2 communicates with the information transmission device 3 via a network. The distribution-side terminal 2 creates voice content based on the voice of the lecturer during the lecture and transmits the generated voice content to the information transmission device 3. The distribution-side terminal 2 may create video content based on the image of the lecturer taken during the lecture and transmit the generated video content to the information transmission device 3. As an example, the distribution-side terminal 2 transmits voice content based on the voice of the lecturer in real time, but may also transmit voice content generated based on recorded voice.
[0023] The information transmission device 3 is, for example, a server for distributing voice content. The information transmission device 3 communicates with the recipient terminals 1a to 1c, the external device 4, and the distribution-side terminal 2 via a network.
[0024] The information transmission device 3 acquires the voice content from the distribution-side terminal 2. The information transmission device 3 distributes the acquired voice content to a plurality of recipient terminals 1a, 1b, and 1c.
[0025] The external device 4 is, for example, a server having a large language model that creates a summary of the input text. As an example, the large language model is ChatGPT (Generative Pretrained Transformer, registered trademark) developed by OpenAI (registered trademark), but may be a large language model other than ChatGPT. The external device 4 creates summary data by summarizing the received text data.
[0026] The processing flow of the content distribution system S in FIG. 1 will be described below. First, when the learner terminal 1 operates the reaction button while the information transmission device 3 is distributing audio content, the learner terminal 1 generates reaction information indicating that the reaction button has been operated. The learner terminal 1 transmits the generated reaction information to the information transmission device 3 ((1) in FIG. 1).
[0027] The information transmission device 3 identifies a large number of reaction time regions where the number of learner terminals 1 that have received reaction information in the distributed audio content is relatively large. The information transmission device 3 identifies large number of reaction text data corresponding to the identified large number of reaction time regions among the text data obtained by converting the content spoken by the instructor.
[0028] The information transmission device 3 transmits the identified large number of reaction text data to the external device 4, and acquires summary data obtained by summarizing this large number of reaction text data from the external device 4 ((2), (3) in FIG. 1). The information transmission device 3 transmits a content report including the acquired summary data to the distribution-side terminal 2 that is the transmission source of the reaction information ((4) in FIG. 1).
[0029] In this way, the information transmission device 3 outputs summary data obtained by summarizing the content of the audio content corresponding to the large number of reaction time regions where the number of learner terminals 1 that have received reaction information is relatively large to the distribution-side terminal 2 and the like. For this reason, the information transmission device 3 can make it easier for the persons related to the distribution of the audio content to grasp the content that the instructor was speaking at the timing when the learners of the audio content were interested.
[0030] The information transmission device 3 may identify a reaction time region including the timing when the reaction information is received from the learner terminal 1 in the distributed audio content, and transmit the reaction text data corresponding to the reaction time region to the external device 4, thereby acquiring summary data corresponding to the reaction time region. In this case, the information transmission device 3 may transmit the summary data corresponding to the reaction information to the learner terminal 1 that has received the reaction information ((5) in FIG. 1).
[0031] [Configuration of Information Transmission Device 3] Figure 2 shows the configuration of the information transmission device 3. The information transmission device 3 includes a communication unit 31, a storage unit 32, and a control unit 33. The control unit 33 includes a distribution unit 331, a reception unit 332, an acquisition unit 333, a specification unit 334, a generation unit 335, and an output unit 336.
[0032] The communication unit 31 is an interface for communicating with the learner terminal 1, the distribution-side terminal 2, and the external device 4. The storage unit 32 includes a storage medium such as a ROM (Read Only Memory) or a RAM (Random Access Memory). The storage unit 32 stores the programs executed by the control unit 33.
[0033] The control unit 33 is, for example, a CPU (Central Processing Unit). By executing the programs stored in the storage unit 32, the control unit 33 functions as the distribution unit 331, the reception unit 332, the acquisition unit 333, the specification unit 334, the generation unit 335, and the output unit 336.
[0034] The distribution unit 331 distributes the content to the learner terminal 1 and the distribution-side terminal 2 via the communication unit 31. The distribution unit 331 acquires the audio content from the distribution-side terminal 2. The distribution unit 331 may acquire the audio content from an external recording device (not shown) or the like. The distribution unit 331 distributes the acquired audio content to one or more learner terminals 1. The distribution unit 331 distributes the audio content to the learner terminal 1 in real time, but may also distribute the audio content generated based on the recorded audio to the learner terminal 1.
[0035] The reception unit 332 receives reaction information indicating that the learner has operated the reaction button from the learner terminal 1 while the distribution unit 331 is distributing the audio content. Figure 3 is a diagram for explaining the operation of the reaction button. Figure 3(a) shows an example of the learner terminal 1 in a state where the reaction button 11 is displayed. Figure 3(b) is a time chart showing the timing when the reaction button 11 is operated.
[0036] In the example of Fig. 3(a), a touch panel is provided on the display of the attendee terminal 1. When the attendee terminal 1 receives an operation by the attendee on the display area of the reaction button 11 displayed on the display, reaction information indicating that the reaction button 11 has been operated is transmitted from the attendee terminal 1 to the information transmission device 3. The reception unit 332 receives this reaction information.
[0037] In Fig. 3(b), the timing at which the distribution unit 331 starts distributing the audio content is shown as "0:00". In the example of Fig. 3(b), the timing at which the reception unit 332 receives the reaction information is indicated by a star mark. Fig. 3(b) shows that the reception unit 332 received the reaction information at the timings of 13 minutes 33 seconds, 29 minutes 10 seconds, and 42 minutes 31 seconds after the start of the distribution by the distribution unit 331.
[0038] When the reception unit 332 receives reaction information exceeding a predetermined number of times from the same attendee terminal 1 within a reference time, the reception unit 332 may invalidate the reaction information received exceeding the predetermined number of times. The reference time is, for example, several tens of seconds or several minutes. The predetermined number of times is, for example, once during the time when one topic is being talked about. The specifying unit 334, which will be described later, totals the time regions where the number of times the reaction information has been received is large in order to specify topics that a plurality of attendees are interested in in the audio content distributed by the distribution unit 331. If an attendee operates the reaction button 11 many times for one topic, the number of times the reaction information is received becomes too large. By operating in this way, the accuracy with which the specifying unit 334 specifies which topic the attendee is interested in is increased.
[0039] As an example, when the predetermined number of times is once, when the reception unit 332 receives reaction information exceeding the predetermined number of times (once) from the same attendee terminal 1 within the reference time, the reception unit 332 receives only the first reaction information. At this time, the reception unit 332 does not receive the reaction information from the same attendee terminal 1 after the second time. The reception unit 332 outputs the received reaction information to the specifying unit 334.
[0040] The acquisition unit 333 acquires the attributes of a plurality of learners. The attributes of the learners are represented by, for example, the name of the industry in which the learner works, the name of the organization, or the name of the department of the organization in which the learner works. The acquisition unit 333 acquires, from the learner terminal 1, the attributes of the learner input, for example, in the learner terminal 1. The acquisition unit 333 outputs information indicating the acquired attributes of the learner to the specifying unit 334. The acquisition unit 333 may notify the specifying unit 334 of the attributes by storing the attributes in the storage unit 32 in association with the learner.
[0041] The specifying unit 334 specifies a reaction time region including the timing at which the reception unit 332 received the reaction information in the voice content being distributed to the learner terminal 1. This reaction time region is shorter than the overall length of the voice content. As an example, the specifying unit 334 specifies a reaction time region that starts a predetermined time before the timing at which the reception unit 332 received the reaction information and ends a predetermined time after the timing at which the reception unit 332 received the reaction information. The predetermined time is, for example, several tens of seconds to several minutes.
[0042] The specifying unit 334 specifies, in the voice content, one or more majority reaction time regions where the number of learner terminals 1 at which the reception unit 332 received the reaction information is relatively large. FIG. 4 shows an example of a method by which the specifying unit 334 specifies a majority reaction time region where the number of learner terminals 1 at which the reaction information was received is relatively large. The horizontal axis in FIG. 4 indicates the time since the start of distribution of the voice content. The vertical axis in FIG. 4 indicates the number of learner terminals 1 at which the reception unit 332 received the reaction information.
[0043] In the example of FIG. 4, the specifying unit 334 specifies the timing at which the number of the learner terminals 1 that have received the reaction information reaches a peak. In FIG. 4, the peak specified by the specifying unit 334 is shown with hatching. The specifying unit 334 specifies a large number reaction time region from a predetermined time before this peak to a predetermined time after this peak as a large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is relatively large. In FIG. 4, the large number reaction time region specified by the specifying unit 334 is indicated by double arrows. Although FIG. 4 shows an example in which there is one peak, the specifying unit 334 may specify a plurality of large number reaction time regions corresponding to a plurality of peaks, respectively.
[0044] Further, in order to specify a large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is relatively large, the specifying unit 334 may specify a large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is equal to or greater than a threshold value. FIG. 5 shows another example of a method in which the specifying unit 334 specifies a large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is relatively large. In the example of FIG. 5, the specifying unit 334 specifies a large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is greater than the threshold value. In FIG. 5, the large number reaction time region specified by the specifying unit 334 is indicated by double arrows. For example, the threshold value is determined such that at least one large number reaction time region in which the number of the learner terminals 1 that have received the reaction information is relatively large is specified by the specifying unit 334 for each voice content.
[0045] The specifying unit 334 may specify, for each attribute of the learner obtained by the acquisition unit 333, one or more large number reaction time regions in which the number of the learner terminals 1 that have received the reaction information is relatively large. More specifically, the specifying unit 334 specifies, in the voice content, one or more large number reaction time regions (hereinafter, also referred to as first attribute time regions) in which the number of the learner terminals 1 of the first attribute that have received the reaction information is relatively large. The specifying unit 334 specifies, in the voice content, one or more large number reaction time regions (hereinafter, also referred to as second attribute time regions) in which the number of the learner terminals 1 of the second attribute that have received the reaction information is relatively large. The specifying unit 334 outputs information indicating the specified large number reaction time region or reaction time region to the generation unit 335.
[0046] [Generation of summary data] The generation unit 335 communicates with the external device 4 via the communication unit 31. The generation unit 335 generates summary data in which the content of the voice content corresponding to the multiple reaction time regions or the reaction time region specified by the specifying unit 334 is summarized. More specifically, the generation unit 335 converts the voice content distributed by the distribution unit 331 into text data. The generation unit 335 specifies text data corresponding to the multiple reaction time regions or the reaction time region specified by the specifying unit 334 among the text data converted from the voice content. The generation unit 335 transmits the specified text data to the external device 4 having a large language model for creating a summary sentence of the input sentence. The generation unit 335 acquires the summary data of the text data generated by the external device 4 from the external device 4.
[0047] FIG. 6 shows an example of creating summary data by the generation unit 335. FIG. 6(a) shows an example of text data converted by the generation unit 335 from voice content. FIG. 6(b) shows an example of summary data generated by the generation unit 335. In the example of FIG. 6(a), this text data is generated by the generation unit 335 in association with the time from the start of distribution of the voice content. For example, in the first paragraph from the top of FIG. 6(a), the time "32:30" from the start of distribution of the voice content is associated with the paragraph "The famous Schrödinger's cat... has not been determined yet...".
[0048] The generation unit 335 identifies a large number of reaction text data corresponding to the large number of reaction time regions identified by the identification unit 334. In the example of FIG. 6(a), when the identification unit 334 identifies a large number of reaction time regions that start 34 minutes after the start of distribution and end immediately before 37 minutes after the start of distribution, as shown by the thick rounded rectangle in FIG. 6, the generation unit 335 corresponds to the time period from 34 minutes after the start of distribution to 35 minutes and 30 seconds, the paragraph of "Next, we will talk about quantum computers. Innovative technical fields such as quantum computing are constantly exploring new knowledge. ···", and the paragraph of "···· Quantum computers are said to be able to quickly process complex calculations by applying quantum mechanics. For example, ···" corresponding to the time period from 35 minutes and 30 seconds after the start of distribution to 37 minutes are identified as a large number of reaction text data.
[0049] The generation unit 335 generates summary data that is less than the number of characters of the identified large number of reaction text data. For example, the generation unit 335 generates 100-character summary data based on 1000-character large number of reaction text data. FIG. 6(b) shows the summary data obtained by summarizing the large number of reaction text data identified by the generation unit 335. The generation unit 335 transmits the identified large number of reaction text data to the external device 4, and as shown in FIG. 6(b), obtains from the external device 4 the summary data "Quantum computers applying quantum mechanics can quickly process complex calculations." that summarizes the large number of reaction text data.
[0050] In this way, the generation unit 335 generates summary data corresponding to the large number of reaction time regions identified by the identification unit 334. When the identification unit 334 identifies two or more large number of reaction time regions, the generation unit 335 generates summary data obtained by summarizing the large number of reaction text data of two or more voice contents corresponding to the two or more identified large number of reaction time regions respectively. Similarly, the generation unit 335 may generate summary data corresponding to the reaction time region identified by the identification unit 334.
[0051] The method by which the generation unit 335 generates summary data is not limited to the method of transmitting the text data corresponding to the time region specified by the specifying unit 334 to the external device 4. The generation unit 335 may transmit the text data converted by the generation unit 335 from the audio content and the time region data indicating the time region specified by the specifying unit 334 to the external device 4, and instruct the external device 4 to summarize the portion of the transmitted text data corresponding to the time region indicated by the time region data.
[0052] When the specifying unit 334 specifies one or more multiple reaction time regions where the number of learner terminals 1 that have received the reaction information is relatively large for each attribute of the learners, the generation unit 335 may generate summary data obtained by summarizing the one or more multiple reaction time regions for each attribute of the learners. For example, when the specifying unit 334 specifies one or more first attribute time regions where the number of learner terminals 1 of the first attribute that have received the reaction information is relatively large, the generation unit 335 generates one or more first summary data obtained by summarizing the content of the audio content corresponding to the first attribute time region.
[0053] When the specifying unit 334 specifies one or more second attribute time regions where the number of learner terminals 1 of the second attribute that have received the reaction information is relatively large, the generation unit 335 generates one or more second summary data obtained by summarizing the content of the audio content corresponding to the second attribute time region. The generation unit 335 outputs the generated summary data and the summary data to the output unit 336.
[0054] [Output of Summary Data] The output unit 336 communicates with the learner terminal 1 and the distribution-side terminal 2 via the communication unit 31. The output unit 336 outputs the summary data generated by the generation unit 335. The output unit 336 transmits to the distribution-side terminal 2 the summary data obtained by summarizing one or more voice contents corresponding to one or more multiple reaction time regions where the number of learner terminals 1 that the reception unit 332 has received reaction information from is relatively large. The output unit 336 outputs, for example, a content report for the instructor of the voice content to the distribution-side terminal 2. In this way, the output unit 336 can enable the instructor or the administrator of the distribution of the voice content to grasp which topics the learners are interested in in the voice content.
[0055] FIG. 7 shows an example of the content report output by the output unit 336. FIG. 7 shows the content report output by the output unit 336 for the instructor of the voice content. Above the content report in FIG. 7, a graph showing the time change in the number of learner terminals 1 that have received reaction information is shown. The hatched regions (A) and (B) in this graph indicate multiple reaction time regions where the number of learner terminals 1 that have received reaction information is relatively large. Below the content report in FIG. 7, the text summarizing the voice content corresponding to the multiple reaction time region including peak (A) and the text summarizing the voice content corresponding to the multiple reaction time region including peak (B) are shown.
[0056] When the generation unit 335 generates summary data obtained by summarizing one or more multiple reaction time regions where the number of learner terminals 1 that have received reaction information is relatively large for each attribute of the learners, the output unit 336 may output the generated summary data in association with the attributes of the learners. For example, the output unit 336 outputs, in association with the first attribute, one or more first summary data corresponding to one or more first attribute time regions specified by the specifying unit 334. The output unit 336 outputs, in association with the second attribute, one or more second summary data corresponding to one or more second attribute time regions specified by the specifying unit 334.
[0057] FIG. 8 shows another example of the output of the summary data by the output unit 336. In the example of FIG. 8, the output unit 336 outputs, in association with the attributes of the attendees, summary data summarizing the content of the voice content corresponding to the multiple reaction time regions in which the reaction button 11 was relatively frequently operated by the attendees with such attributes. In the example of the first row in FIG. 8, the first attribute “manufacturing industry” is associated with the summary data “There is a saying that ‘failure is the mother of success’. Learn from failure...”. In the example of the second row from the top in FIG. 8, the second attribute “medical service industry” is associated with the summary data “‘Continuous effort makes strength’ means that even little by little every day...”. Since the output unit 336 outputs the summary data for each attribute of the attendees, the lecturer or the administrator of the distribution of the voice content can grasp which topic in the voice content the attendees with which attributes were interested in.
[0058] The output unit 336 may transmit the summary data corresponding to the reaction time region including the timing when the reaction information was received to the attendee terminal 1 where the reaction button 11 was operated. For example, when the reception unit 332 receives reaction information indicating that the reaction button 11 was operated in the attendee terminal 1c, the output unit 336 outputs to the attendee terminal 1c the summary data summarizing the content of the voice content corresponding to the reaction time region including the timing when this reaction information was received.
[0059] Since the output unit 336 outputs the summary data summarizing the content of the voice content before and after the timing when the reaction button was operated to the attendee terminal 1 where the reaction button was operated, the attendees who want to obtain the summary data are motivated to actively operate the reaction button during the distribution of the voice content.
[0060] [Example in which a plurality of reaction buttons are provided] By the way, the reasons why the learner tries to obtain the summary data include the case where the learner is interested in the topic at that time in the audio content and the case where the audio content is difficult to understand. Even though the learner presses the reaction button because the audio content is difficult to understand, if the summary sentence at the location where the reaction button is pressed is described in the content report, the instructor or administrator cannot appropriately grasp which topic the learner is interested in. Therefore, for the purpose of clarifying whether the learner operates the reaction button with which intention, the learner terminal 1 may be provided with a positive reaction button for indicating a positive reaction to the audio content and a reconfirmation reaction button for indicating that the audio content is difficult to understand.
[0061] In this case, the reception unit 332 receives, as reaction information, positive reaction information indicating that the learner has operated the positive reaction button and reconfirmation reaction information indicating that the learner has operated the reconfirmation reaction button, respectively. FIG. 9 is a diagram for explaining the operation when a plurality of reaction buttons are provided. In FIG. 9(a), a plurality of reaction buttons including a positive reaction button 101 and a reconfirmation reaction button 102 are displayed on the learner terminal 1. FIG. 9(b) is a time chart showing an example of the timing when the positive reaction button 101 is operated. FIG. 9(c) is a time chart showing an example of the timing when the reconfirmation reaction button 102 is operated.
[0062] When the learner terminal 1 receives an operation by the learner on the display area of the positive reaction button 101, positive reaction information is transmitted from the learner terminal 1 to the information transmission device 3. When the learner terminal 1 receives an operation by the learner on the display area of the reconfirmation reaction button 102, reconfirmation reaction information is transmitted from the learner terminal 1 to the information transmission device 3.
[0063] In FIGS. 9(b) and 9(c), the time when the distribution of the audio content starts is indicated as "0:00". The star marks in FIG. 9(b) indicate the timings at which the reception unit 332 receives positive reaction information. In the example of FIG. 9(b), it shows that the reception unit 332 received positive reaction information at the timings of 12 minutes 37 seconds, 26 minutes 19 seconds, and 42 minutes 31 seconds from the start of the distribution by the distribution unit 331. The star marks in FIG. 9(c) indicate the timings at which the reception unit 332 receives reconfirmation reaction information. In the example of FIG. 9(c), it shows that the reception unit 332 received reconfirmation reaction information at the timings of 20 minutes 12 seconds and 38 minutes 32 seconds from the start of the distribution by the distribution unit 331.
[0064] The specifying unit 334 specifies in the audio content one or more first majority reaction time regions where the number of the learner terminals 1 at which the reception unit 332 receives positive reaction information is relatively large. The method by which the specifying unit 334 specifies one or more first majority reaction time regions where the number of the learner terminals 1 is relatively large is the same as the method by which the specifying unit 334 specifies one or more majority reaction time regions where the number of the learner terminals 1 at which the reception unit 332 receives reaction information is relatively large in FIG. 2.
[0065] The specifying unit 334 respectively specifies in the audio content one or more time regions (hereinafter also referred to as first reaction time regions) at which the reception unit 332 receives positive reaction information. For example, when the specifying unit 334 receives positive reaction information from the reception unit 332 from the learner terminal 1c, it specifies in the audio content a first reaction time region including the timing at which this positive reaction information is received regardless of whether positive reaction information is received from other learner terminals 1a or 1b.
[0066] The generation unit 335 generates one or more first majority reaction summary data obtained by summarizing the content of the audio content corresponding to one or more first majority reaction time regions where the number of the learner terminals 1 at which the specifying unit 334 specifies positive reaction information is relatively large. Since the method by which the generation unit 335 generates the first majority reaction summary data is the same as the method by which the generation unit 335 in FIG. 2 generates summary data, the description thereof is omitted.
[0067] The generation unit 335 generates one or more first reaction summary data obtained by summarizing the content of the voice content corresponding to the one or more first reaction time regions specified by the specifying unit 334. Since the method by which the generation unit 335 generates the first reaction summary data is the same as the method by which the generation unit 335 in FIG. 2 generates the summary data, the description thereof is omitted.
[0068] The output unit 336 transmits the first multiple reaction summary data obtained by summarizing the content of the voice content corresponding to the one or more first multiple reaction time regions to the distribution-side terminal 2 managed by the instructor or administrator of the voice content. The output unit 336 may transmit the first reaction summary data corresponding to the positive reaction information to the learner terminal 1 on which the positive reaction button 101 has been operated.
[0069] Incidentally, from the viewpoint of specifying a topic in which a large number of learners are interested in the voice content, there is little need to specify a multiple reaction time region in which the number of learner terminals 1 that the reception unit 332 has received reconfirmation reaction information is relatively large. For this reason, the specifying unit 334 does not necessarily need to specify one or more multiple reaction time regions in which the number of learner terminals 1 that the reception unit 332 has received reconfirmation reaction information is relatively large. That is, for the purpose of transmitting summary data to the learner terminal 1 that has operated the reconfirmation reaction button 102, regardless of the number of learner terminals 1 that the reception unit 332 has received reconfirmation reaction information, the reception unit 332 may respectively specify one or more reaction time regions (hereinafter also referred to as second reaction time regions) in which the reception unit 332 has received reconfirmation reaction information in the voice content.
[0070] In this case, the generation unit 335 generates one or more second reaction summary data obtained by summarizing the content of the voice content corresponding to the one or more second reaction time regions specified by the specifying unit 334. The output unit 336 transmits the second reaction summary data corresponding to the one or more reconfirmation reaction information to the learner terminal 1 on which the reconfirmation reaction button 102 has been operated.
[0071] In addition, for the purpose of providing feedback to the instructor of the audio content or the administrator of the distribution of the audio content regarding parts of the audio content that were difficult to understand, the specifying unit 334 may specify a large-response time region where the number of learner terminals 1 that the reception unit 332 has received reconfirmation response information from is relatively large. In this case, the output unit 336 may output summary data corresponding to the large-response time region where the reconfirmation response information is relatively large to the distribution-side terminal 2.
[0072] In this way, since the output unit 336 transmits the first large-response summary data corresponding to the first large-response time region where there are relatively many learner terminals 1 that have received positive response information to the distribution-side terminal 2, it is possible for the instructor of the audio content or the administrator of the distribution of the audio content to easily grasp the topics that the learners were interested in in the audio content. At this time, regarding the topics that the learners felt were difficult to understand in the audio content, the learners operate the reconfirmation response button 102, so the output unit 336 can prevent the instructor or the administrator of the distribution of the audio content from confusing the topics that the learners felt were difficult to understand with the topics that the learners were interested in in the audio content.
[0073] When the reception unit 332 receives reconfirmation response information, the specifying unit 334 may specify a second response time region that is different from the case when the reception unit 332 receives positive response information. When the reception unit 332 receives positive response information, the specifying unit 334 specifies, in the audio content, a first response time region that includes the timing at which the positive response information was received. When the reception unit 332 receives reconfirmation response information, the specifying unit 334 may specify, in the audio content, a second response time region that includes the timing at which the reconfirmation response information was received and starts at a timing earlier than the first response time region.
[0074] When the listener operates the reconfirmation response button 102 when the audio content is difficult to hear. For this reason, when the specific unit 334 receives the reconfirmation response information, by specifying the second reaction time region for creating summary data going back to an earlier point in time compared to the first reaction time region, the convenience for the listener can be improved.
[0075] FIG. 10 shows an example of the first reaction time region and the second reaction time region specified by the specific unit 334. FIG. 10(a) shows an example of the first reaction time region including the timing when the positive reaction information is received. FIG. 10(b) shows an example of the second reaction time region including the timing when the reconfirmation reaction information is received. The thick right arrow in FIG. 10(a) indicates the time during the distribution of the audio content. The star mark in FIG. 10(a) indicates the timing when the reception unit 332 receives the positive reaction information. In FIG. 10(a), the first reaction time region is the hatched region. In the example of FIG. 10(a), when the reception unit 332 receives the positive reaction information, the specific unit 334 specifies the first reaction time region that starts 90 seconds before the timing when the positive reaction information is received and ends 90 seconds after the timing when the positive reaction information is received.
[0076] The star mark in FIG. 10(b) indicates the timing when the second reaction information is received. In FIG. 10(b), the second reaction time region is the hatched region. As shown in FIG. 10(b), when the specific unit 334 receives the reconfirmation reaction information, the specific unit 334 specifies the second reaction time region that starts 150 seconds before the timing when the reconfirmation reaction information is received and ends 30 seconds after the timing when the reconfirmation reaction information is received.
[0077] When the reception unit 332 receives the positive reaction information, the output unit 336 transmits the summary data corresponding to the first reaction time region to the listener terminal 1 where the positive reaction button 101 is operated. When the output unit 336 receives the reconfirmation reaction information, the output unit 336 transmits the summary data corresponding to the second reaction time region to the listener terminal 1 where the reconfirmation reaction button 102 is operated.
[0078] [Processing Procedure for Output of Summary Data by Output Unit 336 to Audience Terminal 1] FIG. 11 is a flowchart showing the processing procedure for output of summary data by output unit 336. This processing procedure starts, for example, during communication between information transmission device 3 and a plurality of audience terminals 1. First, distribution unit 331 starts distributing audio content to a plurality of audience terminals 1 (S101). Reception unit 332 receives reaction information indicating that the audience has operated reaction button 11 from audience terminal 1 while distribution unit 331 is distributing the audio content (S102). Identification unit 334 identifies a reaction time region including the timing at which reception unit 332 received the reaction information in the audio content being distributed to audience terminal 1 (S103).
[0079] Distribution unit 331 determines whether or not the distribution of the audio content has ended (S104). When distribution unit 331 determines that the distribution of the audio content has ended (YES in S104), generation unit 335 identifies one or more majority reaction time regions in the audio content where the number of audience terminals 1 that reception unit 332 received reaction information from is relatively large (S105). Generation unit 335 generates summary data in which the content of the audio content corresponding to the majority reaction time region identified by identification unit 334 is summarized (S106).
[0080] Output unit 336 creates a content report including the summary data corresponding to all the majority reaction time regions identified by identification unit 334 (S107). Output unit 336 outputs the created content report to distribution-side terminal 2 (S108). When distribution unit 331 determines that the distribution of the audio content has not ended (NO in S104), generation unit 335 returns to the process of S102.
[0081] [Effects of Information Transmission Device 3 of the Present Embodiment] According to the information transmission device 3 of the present embodiment, the output unit 336 outputs summary data obtained by summarizing the content of the audio content corresponding to the time region specified by the specifying unit 334 to the distribution-side terminal 2 or the like. Therefore, it is possible for the lecturer of the audio content or the administrator of the distribution of the audio content to easily grasp the content that the lecturer was speaking at the time when the listener of the audio content showed interest.
[0082] As described above, the present invention has been described using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist. For example, all or part of the device can be configured by functionally or physically distributing and integrating it in any unit. Also, new embodiments resulting from any combination of multiple embodiments are included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination have the effects of the original embodiments combined.
Explanation of Reference Numerals
[0083] 1 Audience Terminal 1a Audience Terminal 1b Audience Terminal 1c Audience Terminal 2 Distribution-Side Terminal 3 Information Transmission Device 4 External Device 11 Reaction Button 31 Communication Unit 32 Storage Unit 33 Control Unit 101 Positive Reaction Button 102 Reconfirmation Reaction Button 331 Distribution Unit 332 Reception Unit 333 Acquisition Unit 334 Specifying Unit 335 Generation Unit 336 Output Unit S Content Distribution System
Claims
1. A distribution unit that distributes audio content to a recipient terminal; A reception unit that receives reaction information indicating that the recipient has operated a reaction button from the recipient terminal while the distribution unit is distributing the audio content; A specifying unit that specifies, in the audio content, a time region that includes the timing at which the reaction information was received and is shorter than the entire length of the audio content; A generation unit that generates summary data obtained by summarizing the audio content corresponding to the time region specified by the specifying unit; An output unit that outputs the summary data; An information transmission device comprising the above.
2. The generation unit converts the audio content into text data, specifies the text data corresponding to the time region, and transmits the specified text data to an external device having a large language model that creates a summary sentence of the input sentence, thereby generating the summary data. The information transmission device according to Claim 1.
3. The distribution unit distributes the audio content to a plurality of the recipient terminals; The specifying unit specifies, in the audio content, one or more time regions in which the number of recipient terminals from which the reception unit has received the reaction information is relatively large; The generation unit generates the summary data obtained by summarizing one or more pieces of the audio content corresponding to the one or more time regions. The information transmission device according to Claim 2.
4. Further comprising an acquisition unit that acquires attributes of the plurality of recipients; The distribution unit distributes the audio content to the recipient terminals of the plurality of recipients; The specifying unit specifies, in the audio content, one or more first attribute time regions in which the number of recipient terminals from which the reception unit has received the reaction information from the recipient terminals of the recipients of the first attribute is relatively large, and specifies, in the audio content, one or more second attribute time regions in which the number of recipient terminals from which the reception unit has received the reaction information from the recipient terminals of the recipients of the second attribute is relatively large; The generation unit generates one or more first summary data obtained by summarizing the audio content corresponding to the one or more first attribute time regions, and generates one or more second summary data obtained by summarizing the audio content corresponding to the one or more second attribute time regions; The output unit outputs the one or more first summary data in association with the first attribute, and outputs the one or more second summary data in association with the second attribute. The information transmission device according to Claim 3.
5. The reception unit receives, as the reaction information, positive reaction information indicating that a participant has operated a positive reaction button for indicating a positive reaction to the voice content, and reconfirmation reaction information indicating that the participant has operated a reconfirmation reaction button for indicating that the voice content was difficult to hear. The output unit transmits the summary data corresponding to the positive reaction information to the participant terminal at which the positive reaction button was operated, and transmits the summary data corresponding to the reconfirmation reaction information to the participant terminal at which the reconfirmation reaction button was operated. The information transmission device according to any one of claims 1 to 4.
6. When receiving the positive reaction information, the specifying unit specifies, in the voice content, a first reaction time region including the timing at which the positive reaction information was received. When receiving the reconfirmation reaction information, the specifying unit specifies, in the voice content, a second reaction time region that includes the timing at which the reconfirmation reaction information was received and starts at a timing earlier than the first reaction time region. When receiving the positive reaction information, the output unit transmits the summary data corresponding to the first reaction time region to the participant terminal at which the positive reaction button was operated. When receiving the reconfirmation reaction information, the output unit transmits the summary data corresponding to the second reaction time region to the participant terminal at which the reconfirmation reaction button was operated. The information transmission device according to claim 5.
7. The distribution unit distributes the voice content to a plurality of the participant terminals. The specifying unit specifies, in the voice content, one or more first majority reaction time regions in which the number of participant terminals that received the positive reaction information by the reception unit is relatively large, and specifies, regardless of the number of participant terminals that received the reconfirmation reaction information by the reception unit, one or more second reaction time regions in which the reception unit received the reconfirmation reaction information. The output unit transmits the summary data obtained by summarizing the voice content corresponding to the one or more first majority reaction time regions to a distribution-side terminal managed by an instructor or administrator of the voice content, and transmits the summary data obtained by summarizing the voice content corresponding to the one or more second reaction time regions to the participant terminal that is the transmission source of the reconfirmation reaction information. The information transmission device according to claim 5.
8. When the reception unit receives the reaction information exceeding a predetermined number of times from the same learner terminal within a reference time, the reaction information received exceeding the predetermined number of times is invalidated. The information transmission device according to any one of claims 1 to 4.
9. The output unit transmits the summary data to the learner terminal where the reaction button has been operated. The information transmission device according to any one of claims 1 to 4.
10. Executed by a computer, distributing audio content to a learner terminal; receiving reaction information indicating that a learner has operated a reaction button from the learner terminal while distributing the audio content; identifying a time region that includes the timing at which the reaction information was received and is shorter than the total length of the audio content; generating summary data obtained by summarizing the audio content corresponding to the identified time region; outputting the summary data; An information transmission method comprising:
11. Causing a computer to distribute audio content to a learner terminal; receive reaction information indicating that a learner has operated a reaction button from the learner terminal while distributing the audio content; identifying a time region that includes the timing at which the reaction information was received and is shorter than the total length of the audio content; generating summary data obtained by summarizing the audio content corresponding to the identified time region; outputting the summary data; A program for causing the above to be executed.
Citation Information
Patent Citations
Communication system, communication apparatus and program
JP2009200935A
Seminar distribution system, terminal device, seminar distribution method, and seminar distribution program
JP2019033463A
Video information providing system
JP2019213038A
Information processing device, information processing method, and program
JP2021022085A
Program, information processing device, and method
JP2022044558A