Content distribution device, receiving device and program
The system synchronizes live subtitles with live broadcast programs over the internet by applying multiple speech recognition processes and correcting subtitle time information, addressing display delays and enhancing accessibility.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON HOSO KYOKAI
- Filing Date
- 2022-01-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing content distribution systems for live broadcast programs over the internet struggle to accurately synchronize live subtitles with the audio, leading to significant display delays, especially when speech recognition performance is low.
A content distribution system that applies multiple speech recognition processes to audio data, calculates text matching rates, and corrects subtitle time information using the speech recognition data with the highest matching rate to synchronize subtitles with the video content.
This approach enables high-accuracy synchronization of live subtitles with the program content, reducing display delays and improving accessibility for hearing-impaired viewers.
Smart Images

Figure 0007867231000001 
Figure 0007867231000002 
Figure 0007867231000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a content distribution device, a receiving device, and a program for live streaming video, including subtitle data, over the internet. [Background technology]
[0002] Traditionally, television broadcasting has offered subtitled broadcasting, a service for the hearing impaired, which displays the audio of broadcast programs as text on the screen. Subtitles broadcast during live programs (hereinafter referred to as "live subtitles") are produced by manually transcribing the audio of the live program. As a result, live subtitles are delayed by the time it takes to transcribe them, and are displayed on the screen with a delay compared to the audio of the live program.
[0003] To mitigate the display delay of these live subtitles, efforts are being made to improve the efficiency of live subtitle production through manual transcription, including the use of speech recognition technology or high-speed keyboards. Generally, there are several methods for producing subtitles, including methods that directly produce them from the audio of the broadcast program, and methods that involve re-speaking the audio in a quiet room to improve the accuracy of speech recognition. Currently, these differences in methods result in variations in subtitle production delays and the accuracy of subtitle reproduction relative to the audio of the broadcast program.
[0004] On the other hand, with the recent proliferation of smartphones and video streaming technologies, there is a growing demand for broadcast programs to be provided not only on television but also simultaneously on the internet.
[0005] Some broadcasters abroad already broadcast programs on the internet simultaneously with their television broadcasts, and it is expected that similar services will be rolled out in Japan in the future. In order to provide the same service in Japan, it will be necessary to achieve the same level of service on the internet as on television, and the same level of service for subtitling will also be required.
[0006] In addition, as a technology widely used in recent video distribution, there is adaptive streaming. Adaptive streaming is a technology that realizes seamless video distribution by changing the video quality of distributing multi-bitrate content according to the communication speed of the receiving device.
[0007] Specifically, on the distribution side, the content is encoded at multiple bitrates, and files divided in units of seconds are generated. The receiving side that receives the streaming sequentially acquires files with a bitrate suitable for the communication speed of the receiving device itself from the distribution side, and plays them back by connecting the files together. Thereby, even in a receiving device where the communication speed fluctuates, the playback of the content can be continued, and seamless video distribution can be realized (for example, refer to Non-Patent Document 1).
[0008] However, in adaptive streaming, on the distribution side, the content of the input video and audio data is temporarily held in a buffer, and files are generated every few seconds, so at least a few seconds of delay occurs.
[0009] On the other hand, in a live broadcast program, when directly performing a file generation process for adaptive streaming (hereinafter referred to as "encoding") using the same signal as the broadcast and distributing the generated file as distribution data via the Internet, the display of live subtitles will be delayed in the same way as the broadcast. In this case, for hearing-impaired people, it is easier to understand the program content when the display delay of the live subtitles is smaller.
[0010] As a technology for suppressing this delay, a content distribution device has been proposed that changes the time correction process of live subtitles according to the degree of display delay of live subtitles (for example, refer to Patent Document 1).
Prior Art Documents
Patent Documents
[0011]
Patent Document 1
[0012] [Non-Patent Document 1] A. Zambelli, “IIS Smooth Streaming Technical Overview”, Mar.2009 [Overview of the project] [Problems that the invention aims to solve]
[0013] The content distribution device described in Patent Document 1 above achieves synchronization of live subtitles in live distribution of video content of live broadcast programs without requiring a distribution delay.
[0014] Specifically, this content distribution device determines the subtitle delay elapsed time from the delay time between the raw subtitle data extracted from the broadcast transmission signal and the speech recognition data generated by applying speech recognition processing to the audio contained in the broadcast transmission signal. The content distribution device then compares the subtitle delay elapsed time with the encoding completion time when the encoding of the broadcast transmission signal is completed, and corrects the subtitle time information regarding the time when the raw subtitle data is displayed on the screen according to the comparison result.
[0015] Thus, since the content distribution device described in Patent Document 1 is based on the premise of performing speech recognition processing, it is only possible to suppress the display delay of live subtitles for program content if the recognition performance of the speech recognition processing is high.
[0016] However, if the speech recognition processing performance is low, it may not be possible to perform correct time correction on the subtitle time information in the raw subtitle data. As a result, it may become impossible to suppress the display delay of the raw subtitles relative to the program content with high accuracy.
[0017] Therefore, the present invention was made to solve the above-mentioned problems, and its objective is to provide a content distribution device, a receiving device, and a program that can suppress with high accuracy the display delay of live subtitles relative to the program content in a system for distributing video content of live broadcast programs over the internet. [Means for solving the problem]
[0019] In order to solve the aforementioned problem, Claim 1 The content distribution device is A content distribution device that, when distributing video content of a live broadcast program over the internet, receives a broadcast transmission signal containing the video content, generates distribution data based on the broadcast transmission signal, and corrects the subtitle time information of the live subtitle data included in the broadcast transmission signal, comprises an encoder that encodes the broadcast transmission signal and generates the distribution data, and a subtitle processing unit that extracts the live subtitle data from the broadcast transmission signal and generates new live subtitle data by correcting the subtitle time information indicating the time when the subtitles of the live subtitle data are displayed on the screen, wherein the subtitle processing unit comprises a subtitle extraction unit that extracts the live subtitle data from the broadcast transmission signal, a speech recognition unit that applies a predetermined number of different speech recognition processes to the audio included in the broadcast transmission signal and generates a number of different speech recognition data, and the audio generated by the speech recognition unit For each of the aforementioned multiple different speech recognition data, the speech recognition data is, Extracted by the subtitle extraction unit The raw subtitle data is divided into units of the same number of characters as the raw subtitle data to generate multiple different divided speech recognition data, the raw subtitle data is used as the ground truth data, the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated, the divided speech recognition data with the highest similarity is determined to be the matching target, the text matching rate is calculated for each of the multiple different matching targets corresponding to the multiple different speech recognition data with the raw subtitle data, the matching target with the highest text matching rate is determined, and the matching target Indicates the time when the audio is output. Using the audio time information, the subtitle time information of the raw subtitle data is corrected, and the new raw subtitle data is generated. A matching section, It is characterized by the following:
[0021] moreover, Claim 2 The receiving device is A receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back the video, audio, and subtitles contained in the broadcast signal, comprising: a decoder that decodes the IP content and generates the broadcast signal; a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data, wherein the subtitle processing unit comprises: a subtitle extraction unit that extracts the raw subtitle data from the broadcast signal; a speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal to generate a number of different speech recognition data; and the speech recognition unit generates For each of the aforementioned multiple different speech recognition data, the speech recognition data is, Extracted by the subtitle extraction unitThe raw subtitle data is divided into units of the same number of characters as the raw subtitle data to generate multiple different divided speech recognition data, the raw subtitle data is used as the ground truth data, the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated, the divided speech recognition data with the highest similarity is determined to be the matching target, the text matching rate is calculated for each of the multiple different matching targets corresponding to the multiple different speech recognition data with the raw subtitle data, the matching target with the highest text matching rate is determined, and the matching target Indicates the time when the audio is output. Using the audio time information, the subtitle time information of the raw subtitle data is corrected, and the new raw subtitle data is generated. A matching section, It is characterized by the following:
[0022] Furthermore, claims 3 The receiving device is A receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back the video, audio, and subtitles contained in the broadcast signal, comprising: a decoder that decodes the IP content and generates the broadcast signal; a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data, wherein the subtitle processing unit comprises: a subtitle extraction unit that extracts the raw subtitle data from the broadcast signal; a speech recognition unit that applies a predetermined set of different speech recognition processes to the audio contained in the broadcast signal to generate a set of different speech recognition data; and a matching unit that calculates a text matching rate between each of the set of different speech recognition data generated by the speech recognition unit and the raw subtitle data extracted by the subtitle extraction unit, determines the speech recognition data with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using the audio time information indicating the time when the audio of the speech recognition data is output, and generates new raw subtitle data, Furthermore, the decoder is characterized by having a delay unit that delays the broadcast signal generated by the decoder by the time it takes for the subtitle processing unit to receive the broadcast signal and output the new raw subtitle data. Furthermore, the receiving device according to claim 4 is a receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back video, audio, and subtitles included in the broadcast signal, comprising: a decoder that decodes the IP content and generates the broadcast signal; a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data, wherein the subtitle processing unit comprises: a subtitle extraction unit that extracts the raw subtitle data from the broadcast signal; a speech recognition unit that applies a predetermined number of different speech recognition processes to the audio included in the broadcast signal to generate a number of different speech recognition data; and for each of the number of different speech recognition data generated by the speech recognition unit, the speech recognition data is processed by the subtitle extraction unit The matching unit comprises: a unit that divides the extracted raw subtitle data into units of the same number of characters as the original data to generate multiple different divided speech recognition data; the raw subtitle data is used as the ground truth data to calculate the similarity between the ground truth data and each of the multiple different divided speech recognition data; the divided speech recognition data with the highest similarity is determined to be the matching target; for each of the multiple different matching targets corresponding to the multiple different speech recognition data, a text matching rate is calculated between the raw subtitle data and the original data; the matching target with the highest text matching rate is determined; and the matching unit corrects the subtitle time information of the raw subtitle data using audio time information indicating the time when the audio of the matching target is output, thereby generating new raw subtitle data; and further comprises a delay unit that delays the broadcast signal generated by the decoder by the time from when the subtitle processing unit receives the broadcast signal until it outputs the new raw subtitle data.
[0023] Furthermore, claims 5The program provides a computer that constitutes a content distribution device for distributing video content of a live broadcast program over the internet. This computer receives a broadcast transmission signal containing the video content, generates distribution data based on the broadcast transmission signal, and corrects the subtitle time information of the raw subtitle data included in the broadcast transmission signal. The computer functions as an encoder that encodes the broadcast transmission signal and generates the distribution data, and as a subtitle processing unit that extracts the raw subtitle data from the broadcast transmission signal and corrects the subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen, thereby generating new raw subtitle data. The subtitle processing unit comprises a subtitle extraction unit that extracts the raw subtitle data from the broadcast transmission signal, a speech recognition unit that applies a predetermined set of different speech recognition processes to the audio included in the broadcast transmission signal and generates a set of different speech recognition data, and for each of the set of different speech recognition data generated by the speech recognition unit, The speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the matching target. For each of the multiple different matching targets corresponding to the multiple different speech recognition data, The text matching rate is calculated between the raw subtitle data and the above, and the highest text matching rate is calculated. Matching target Determine the Matching target The system is characterized by comprising: a matching unit that uses audio time information indicating the time when the audio is output to correct the subtitle time information of the raw subtitle data and generate new raw subtitle data.
[0024] Furthermore, claims 6The program causes a computer that constitutes a receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back the video, audio, and subtitles contained in the broadcast signal to function as a subtitle processing unit that inputs the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data, wherein the subtitle processing unit comprises a subtitle extraction unit that extracts the raw subtitle data from the broadcast signal, a speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal to generate a number of different speech recognition data, and for each of the number of different speech recognition data generated by the speech recognition unit, The speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the matching target. For each of the multiple different matching targets corresponding to the multiple different speech recognition data, The text matching rate is calculated between the raw subtitle data and the above, and the highest text matching rate is calculated. Matching target Determine the Matching target The system is characterized by comprising: a matching unit that uses audio time information indicating the time when the audio is output to correct the subtitle time information of the raw subtitle data and generate new raw subtitle data. [Effects of the Invention]
[0025] As described above, according to the present invention, in a system for distributing video content of live broadcast programs over the internet, it is possible to suppress the display delay of live subtitles relative to the program content with high accuracy. [Brief explanation of the drawing]
[0026] [Figure 1] This diagram shows a schematic representation of the overall configuration of a content distribution system including a content distribution device according to an embodiment of the present invention, and a block diagram showing an example of the configuration of the content distribution device. [Figure 2] This block diagram shows an example configuration of a subtitle processing unit included in a content distribution device. [Figure 3]This is a block diagram showing an example of the matching unit configuration. [Figure 4] This flowchart shows an example of the processing performed by the speech recognition judgment unit. [Figure 5] This figure illustrates an example of calculating the text matching rate between raw subtitle data a and speech recognition data b1. [Figure 6] This flowchart shows an example of the matching processing unit's operation. [Figure 7] This diagram illustrates a specific example of processing performed by the matching processing unit. [Figure 8] This diagram shows a schematic example of the overall configuration of a content distribution system including a receiving device according to an embodiment of the present invention, and a block diagram showing an example of the configuration of the receiving device. [Figure 9] This is a block diagram showing an example configuration of a subtitle processing unit provided in a receiving device. [Figure 10] This flowchart shows an example of how the delay section is handled. [Modes for carrying out the invention]
[0027] Hereinafter, embodiments for carrying out the present invention will be described in detail with reference to the drawings. However, descriptions that are unnecessarily detailed may be omitted. For example, detailed descriptions of already well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid the following description becoming unnecessarily verbose and to facilitate understanding for those skilled in the art. The accompanying drawings and the following description are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter of the claims.
[0028] The present invention relates to a content distribution system for distributing video content of live broadcast programs over the internet, characterized in that it matches raw subtitle data included in a broadcast transmission signal with speech recognition data obtained by multiple different speech recognition processes applied to the audio included in the broadcast transmission signal, and corrects the time of the raw subtitle data using the time of the speech recognition data with the highest matching rate.
[0029] This ensures that the speech recognition process with the highest recognition performance is selected from among several different speech recognition processes. Furthermore, because speech recognition data obtained from the high-performance speech recognition process can be used, correct time correction processing can be performed on the raw subtitle data. By performing time correction processing on the raw subtitle data in this way, the timing of speech in the video is matched with the corresponding raw subtitle data. Therefore, the display delay of the raw subtitles relative to the program content can be suppressed with high accuracy.
[0030] [Content distribution system] First, a content distribution system including a content distribution device according to an embodiment of the present invention will be described. Figure 1 is a schematic diagram showing an example of the overall configuration of a content distribution system including a content distribution device according to an embodiment of the present invention, and a block diagram showing an example of the configuration of the content distribution device.
[0031] This content distribution system is a system that distributes video content of live broadcast programs over the internet via an IP network, that is, a system that performs live video streaming, and is composed of a content distribution device 1, a distribution server 2, and a receiving device 100.
[0032] The content distribution device 1 receives a broadcast transmission signal containing video content from an external source, encodes the broadcast transmission signal, divides it into multiple files, and generates distribution data D of multiple files. For example, an SDI (Serial Digital Interface) signal is used as the broadcast transmission signal.
[0033] The content distribution device 1 matches the raw subtitle data included in the broadcast transmission signal with the speech recognition data obtained from multiple different speech recognition processes applied to the audio included in the broadcast transmission signal, and calculates the character matching rate between the raw subtitle data and each of the speech recognition data. Then, the content distribution device 1 corrects the time of the raw subtitle data using the time included in the speech recognition data with the highest matching rate, and synchronizes the raw subtitle data with the program content of the video content in the distribution data D. The content distribution device 1 transmits the distribution data D and the synchronized (corrected) raw subtitle data a' to the distribution server 2.
[0034] The broadcast transmission signal input to content distribution device 1 consists of video, audio, and live subtitle data. Each of the video, audio, and live subtitle data contains time information based on a common time. As mentioned above, the live subtitle data is produced by manual transcription from the audio of a live broadcast program, and therefore is delayed compared to the video and audio program content. In other words, the time contained in the live subtitle data is delayed compared to the time contained in the speech recognition data obtained by speech recognition processing. The delay time of the live subtitle data relative to the program content varies depending on the operator producing it and the live subtitle data itself.
[0035] To explain in detail using an example, if a person in the video included in the broadcast transmission signal says "Good morning" at the time 0:00-0:02 in the video, the same "Good morning" in the raw subtitle data included in the broadcast transmission signal will be stored at a different time, such as 0:07-0:09. This is because raw subtitle data is generally created by manual transcription, and the video advances by the time it takes to generate the raw subtitle data. Therefore, when the raw subtitle data is added to the broadcast transmission signal, a time lag occurs between the video and the subtitle data. Thus, it can be said that it is common for there to be a discrepancy between the video included in the broadcast transmission signal and the raw subtitle data.
[0036] The distribution server 2 receives the video content distribution data D and raw subtitle data a' from the content distribution device 1 and stores them in memory.
[0037] The receiving device 100 is a conventional device, such as a video player like a smartphone. The receiving device 100 obtains a playlist (not shown) from the content distribution device 1 via the distribution server 2 and the IP network, and understands the file structure based on the playlist. Then, based on the playlist, the receiving device 100 obtains the IP content, including the distribution data D and raw subtitle data a', via the IP network using HTTP (Hypertext Transfer Protocol).
[0038] The receiving device 100 plays the content by combining the distribution data D and raw subtitle data a' contained in the IP content according to the time in the playlist, displaying the video and subtitles on the screen, and outputting audio.
[0039] As a result, the receiving device 100 can play video content with minimal delay in subtitle display relative to the video and audio. The smaller this subtitle display delay, the easier it is for the user to understand the program content. This is especially beneficial for people with hearing impairments, as live subtitles play a significant role in helping them understand the program content.
[0040] [Content distribution device 1] Next, a content distribution device 1 according to an embodiment of the present invention will be described. In Figure 1, the content distribution device 1 includes a distribution unit 10, an encoder 11, and a subtitle processing unit 12. The distribution unit 10 receives a broadcast transmission signal, distributes the broadcast transmission signal, and outputs the distributed broadcast transmission signal to the encoder 11 and the subtitle processing unit 12.
[0041] The encoder 11 receives the broadcast transmission signal from the distribution unit 10, encodes the broadcast transmission signal to divide it into files of several seconds each, and generates distribution data D. The encoder 11 sends the distribution data D to the distribution server 2.
[0042] The subtitle processing unit 12 receives a broadcast transmission signal from the distribution unit 10, extracts raw subtitle data from the broadcast transmission signal, and performs multiple different speech recognition processes on the audio contained in the broadcast transmission signal to generate multiple speech recognition data.
[0043] The subtitle processing unit 12 performs text matching between the raw subtitle data and each of the multiple speech recognition data. The subtitle processing unit 12 then determines the speech recognition data with the highest text matching rate among the multiple speech recognition data, and uses the time of that speech recognition data to correct the time of the raw subtitle data, thereby generating new raw subtitle data a'. The subtitle processing unit 12 then sends the raw subtitle data a' to the distribution server 2.
[0044] (Subtitle processing unit 12) Next, the subtitle processing unit 12 shown in Figure 1 will be described in detail. Figure 2 is a block diagram showing an example configuration of the subtitle processing unit 12 provided in the content distribution device 1. This subtitle processing unit 12 includes a subtitle extraction unit 20, speech recognition units 21-1, ..., 21-N, and a matching unit 22. N is an integer of 2 or more.
[0045] The subtitle extraction unit 20 receives the broadcast transmission signal from the distribution unit 10, extracts the raw subtitle data a from the broadcast transmission signal, and outputs the raw subtitle data a to the matching unit 22. The raw subtitle data a contains the time t when the raw subtitle is displayed on the screen. a This includes time information (subtitle time information).
[0046] The speech recognition units 21-1, ..., 21-N perform different speech recognition processing. For example, the speech recognition units 21-1, ..., 21-N may use different speech recognition libraries or perform different audio waveform processing. The speech recognition unit 21-1 receives a broadcast transmission signal from the distribution unit 10, applies known speech recognition processing to the audio contained in the broadcast transmission signal, generates speech recognition data b1, and outputs the speech recognition data b1 to the matching unit 22. The speech recognition data b1 contains the time t when the audio is output. b1This includes time information (audio time information) related to the event.
[0047] The speech recognition unit 21-N applies a known speech recognition process, different from that of other speech recognition units 21-1, etc., to the audio contained in the broadcast transmission signal input from the distribution unit 10, generates speech recognition data bN, and outputs the speech recognition data bN to the matching unit 22. The speech recognition data bN contains the time t when the audio is output. bN This includes time information related to this.
[0048] The matching unit 22 receives raw subtitle data a from the subtitle extraction unit 20, and also receives speech recognition data b1,...,bN from the speech recognition units 21-1,...,21-N.
[0049] The matching unit 22 matches the speech recognition data b1 with the raw subtitle data a and identifies the portion of the raw subtitle data a that is determined to be identical through the matching. The matching unit 22 then calculates the text matching rate between the identified raw subtitle data a and the speech recognition data b1.
[0050] The matching unit 22 performs the same processing on the speech recognition data b2, ..., bN as it did on the speech recognition data b1 to determine the text matching rate.
[0051] The matching unit 22 determines the speech recognition data with the highest text matching rate among the speech recognition data b1, ..., bN, and the time t of that speech recognition data b Using the time t of the raw subtitle data a By correcting this, new raw subtitle data a' is generated and output.
[0052] Figure 3 is a block diagram showing an example configuration of the matching unit 22. This matching unit 22 includes an input unit 30, a speech recognition determination unit 31, and a matching processing unit 32.
[0053] The input unit 30 receives raw subtitle data a from the subtitle extraction unit 20, and also receives speech recognition data b1,...,bN from the speech recognition units 21-1,...,21-N, and outputs this data to the speech recognition determination unit 31.
[0054] The granularity of the raw subtitle data a and speech recognition data b1,...,bN output from the input unit 30 to the speech recognition determination unit 31 shall be at the sentence level. However, the granularity may also be at the character level, word level, or multiple sentence level.
[0055] Figure 4 is a flowchart showing an example of processing by the speech recognition determination unit 31. The speech recognition determination unit 31 receives raw subtitle data a and speech recognition data b1,...,bN from the input unit 30 (step S401).
[0056] Specifically, the speech recognition determination unit 31 first inputs each of the speech recognition data b1, ..., bN, and then inputs the raw subtitle data a corresponding to each of the speech recognition data b1, ..., bN. The speech recognition determination unit 31 then performs a matching operation between each of the speech recognition data b1, ..., bN and the raw subtitle data a, and identifies the portion of the raw subtitle data a that is determined to be identical through the matching operation.
[0057] The speech recognition determination unit 31 identifies the raw subtitle data a for each of the speech recognition data b1, ..., bN as the correct data. Then, the speech recognition determination unit 31 performs text matching between the identified raw subtitle data a and each of the speech recognition data b1, ..., bN and calculates the text matching rate for each (step S402).
[0058] Figure 5 illustrates an example of calculating the text matching rate between raw subtitle data a and speech recognition data b1. The identified raw subtitle data a is "Today in Tokyo it will be sunny," and the speech recognition data b1 is "Today in Tokyo it will be sunny..." Raw subtitle data a is treated as the correct answer data and has 10 characters.
[0059] In the example shown in Figure 5, the speech recognition determination unit 31 divides the speech recognition data b1 into units of 10 characters, which is the number of characters in the raw subtitle data a, and generates 10-character speech recognition data (divided speech recognition data) b1-1, b1-2, b1-3, ... For example, "Today's Tokyo Island is sunny" is generated as speech recognition data b1-1, "Today's Tokyo Island will be sunny" is generated as speech recognition data b1-2, and "Today's Tokyo Island will be sunny" is generated as speech recognition data b1-3.
[0060] The speech recognition determination unit 31 calculates the similarity between the ground truth data and each of the speech recognition data b1-1, b1-2, b1-3, ... for example, by N-gram search. Then, the speech recognition determination unit 31 determines that the speech recognition data with the highest similarity among b1-1, b1-2, b1-3, ... is the target for matching. For example, suppose that speech recognition data b1-2 "Today Tokyo Island will be sunny" is determined to be the target for matching. Note that the process of calculating the similarity between the ground truth data and the speech recognition data is known, so a detailed explanation is omitted here.
[0061] The speech recognition judgment unit 31 calculates the text matching rate between the 10-character correct data "Today in Tokyo it will be sunny" and the 10-character speech recognition data b1-2 "Today on Tokyo Island it will be sunny". For example, the speech recognition judgment unit 31 scores both data based on matching of the first character, matching of each character, matching of consecutive characters, matching of the last character, etc., and calculates the total score (total score of the correct data, total score of the speech recognition data b1-2). Then, the speech recognition judgment unit 31 calculates the text matching rate by dividing the total score of the speech recognition data b1-2 by the total score of the correct data. Note that the method for calculating the text matching rate is known, so a detailed explanation is omitted here.
[0062] Furthermore, the speech recognition determination unit 31 sets the 10-character speech recognition data b1-2, which is the target of matching, as speech recognition data b1. As a result, speech recognition data b1-2, "Today's Tokyo Island will be sunny," is used as speech recognition data b1 instead of speech recognition data b1, "Today's Tokyo Island will be sunny..." in the processing described later.
[0063] The new speech recognition data b1, "Today, Tokyo Island will be sunny," is output at time t. b1 This refers to the time t when the voice recognition data b1 "It will be sunny on Tokyo Island today..." is output. b1 This will be different.
[0064] In this way, the text matching rate between the raw subtitle data a and the speech recognition data b1 is calculated using the raw subtitle data a as the ground truth data.
[0065] Furthermore, in step S402, the speech recognition determination unit 31 may perform text matching with the program script input from an external source as the correct answer data, instead of the raw subtitle data a, between each of the speech recognition data b1,...,bN.
[0066] In this case, as shown in Figure 3, the matching unit 22 includes an input unit 30, a speech recognition determination unit 31, and a matching processing unit 32, as well as a communication unit 33. The communication unit 33 receives program information, including program scripts, and outputs the program script to the speech recognition determination unit 31. The frequency at which the communication unit 33 receives program information and outputs the program script is arbitrary and may be in units of a few seconds, on a program-by-program basis, or on a daily basis.
[0067] Returning to Figures 3 and 4, the speech recognition determination unit 31 uses the text matching rates of each of the speech recognition data b1, ..., bN to determine which of the speech recognition data b1, ..., bN has the highest text matching rate (step S403).
[0068] The voice recognition determination unit 31 outputs the raw subtitle data a, the voice recognition data b, and the text matching rate of the voice recognition data b to the matching processing unit 32 (step S404).
[0069] FIG. 6 is a flowchart showing an example of the processing of the matching processing unit 32. The matching processing unit 32 inputs the raw subtitle data a, the voice recognition data b, and the text matching rate of the voice recognition data b from the voice recognition determination unit 31 (step S601).
[0070] The matching processing unit 32 compares the text matching rate with a preset threshold value (step S602).
[0071] If the matching processing unit 32 determines in step S602 that the text matching rate is greater than or equal to the threshold value (step S602: ≧), it determines that the matching between the raw subtitle data a and the voice recognition data b has succeeded.
[0072] Then, the matching processing unit 32 overwrites the time t a (the time t when the raw subtitle data a is displayed on the screen a ) with the time t b (the time t when the voice of the voice recognition data b is output b )(t a ←t b ) and generates new raw subtitle data a' (step S603). In the raw subtitle data a', the time t b when the raw subtitle data a is displayed on the screen a is included as the time t
[0073] If the matching processing unit 32 determines in step S602 that the text matching rate is less than the threshold value (step S602: <), it determines that the matching between the raw subtitle data a and the voice recognition data b has failed.
[0074] Then, the matching processing unit 32 sets the time t aSubtract a predetermined value P from the time t contained in the raw subtitle data a. a Then, the subtraction result is overwritten (t a ←t a -P), generate new raw subtitle data a' (step S604). Raw subtitle data a' contains the time t when raw subtitle data a is displayed on the screen. a -P, new time t a It will be included as such.
[0075] The predetermined value P may be a fixed value set in advance, or it may be a moving average of the actual value at the most recent successful matching. In the latter case, the matching processing unit 32 takes into account the most recent predetermined number of time t in the processing of step S603 when the matching is successful. a ,t b Keep this, at time t a From time t b The average value obtained by subtracting is calculated, and this average value is set to a predetermined value P. The value P set in this way is used in the process of step S604.
[0076] The matching processing unit 32 moves from steps S603 and S604 to output the raw subtitle data a' (step S605).
[0077] Furthermore, in the case of step S602(<), i.e., when the matching processing unit 32 determines that the matching between the raw subtitle data a and the speech recognition data b has failed, instead of processing step S604 described above, it processes the time t included in the raw subtitle data a, similar to the processing in step S603. a The time t included in the speech recognition data b b Alternatively, you could overwrite the existing data and generate new raw subtitle data a'.
[0078] Furthermore, if step S602(<) is selected, the matching processing unit 32 may choose not to perform the processing in step S604 described above. In this case, the matching processing unit 32 does not output the raw subtitle data a'.
[0079] Figure 7 illustrates a specific example of processing by the matching processing unit 32. The time t of the raw subtitle data a "Tokyo is sunny" a The time t is "10:00:10", and the time t corresponds to the voice recognition data b "Tokyo Island is sunny". b Let's assume that the time is "10:00:00". Also, in step S602 of Figure 6, the text matching rate in this case is above the threshold (step S602:≧), and the matching between the raw subtitle data a and the speech recognition data b is considered successful.
[0080] Then, according to step S603 in Figure 6, the time t included in the raw subtitle data a a At "10:00:10", the time t included in the speech recognition data b b "10:00:00" is overwritten, and raw subtitle data a' synchronized with the broadcast content is generated. As a result, the time t of raw subtitle data a "Tokyo is sunny" is generated. a "10:00:10" is corrected to "10:00:00", and new raw subtitle data a' is generated.
[0081] Also, the time t of the raw subtitle data a "It's raining in Kanagawa Prefecture" a The time is "10:00:17", and the time of the corresponding voice recognition data b "Kanagawa Prefecture is candy" is t b If the time is "10:00:06", then the matching is considered successful. In this case, the time t of the raw subtitle data a is... a "10:00:17" is corrected to "10:00:06", and live subtitle data a' synchronized with the broadcast content is generated.
[0082] Also, the time t of the raw subtitle data a "Saitama Prefecture is cloudy" a The time is "10:00:26", and the time t of the corresponding voice recognition data b "Saitama Prefecture is a pharmacy" b If the time is "10:00:15", then the matching is considered successful. In this case, the time t of the raw subtitle data a is... a "10:00:26" is corrected to "10:00:15", and live subtitle data a' synchronized with the broadcast content is generated.
[0083] Thus, if text matching is successful and it is determined that the content of raw subtitle data a and the content of speech recognition data b are the same, the time t of the corresponding raw subtitle data a is determined. a The time t of the speech recognition data b b It is overwritten. This generates raw subtitle data a' that is synchronized with the broadcast content.
[0084] Furthermore, when the matching processing unit 32 generates the raw subtitle data a' in steps S603 and S604 of Figure 6, it may also change the display time of the subtitle in the raw subtitle data a' according to the number of characters that make up the raw subtitle data a'. The display time of the subtitle is the time period between the time when the subtitle display starts and the time when the subtitle display ends.
[0085] Specifically, the matching processing unit 32 multiplies the number of characters that make up the raw subtitle data a' by a predetermined display time per character to determine the display time of the subtitle in the raw subtitle data a', and reflects this in the display time of the subtitle included in the raw subtitle data a'.
[0086] As described above, according to the content distribution device 1 of the embodiment of the present invention, the encoder 11 encodes the broadcast transmission signal to generate distribution data D. The subtitle extraction unit 20 of the subtitle processing unit 12 extracts raw subtitle data a from the broadcast transmission signal. In addition, the speech recognition units 21-1, ..., 21-N apply a known speech recognition process different from that of the other components to the audio contained in the broadcast transmission signal to generate speech recognition data b1, ..., bN.
[0087] The matching unit 22 calculates the text matching rate for each of the speech recognition data b1, ..., bN with the raw subtitle data a. Then, the matching unit 22 determines the speech recognition data with the highest text matching rate among the speech recognition data b1, ..., bN, and the time t of that speech recognition data. b Using the time t of the raw subtitle data a a By correcting this, new raw subtitle data a' is generated and output.
[0088] Distribution data D and raw subtitle data a' are sent to distribution server 2, and the IP content including distribution data D and raw subtitle data a' is sent to receiving device 100 via the IP network.
[0089] Thus, the time t of the raw subtitle data a a This is the time t of the speech recognition data obtained by the speech recognition processing with the highest recognition performance. b The data is corrected using this method, and new raw subtitle data a' is generated. This makes it possible to suppress the display delay of raw subtitles relative to the program content with high accuracy in a content distribution system that distributes video content of live broadcast programs over the internet, enabling the provision of programs that are easier to understand. In addition, the content distribution device 1 can suppress the display delay of subtitles by utilizing the time spent on encoding, by performing the processing of the subtitle processing unit 12 in parallel with the processing of the encoder 11.
[0090] Here, the recognition performance of the speech recognition processing by the speech recognition units 21-1, ..., 21-N generally differs depending on the type of video content of the live broadcast program (news, sports, variety, etc.). As mentioned above, the time t of the live subtitle data a a The time t of the speech recognition data obtained by the speech recognition process with the highest recognition performance among the speech recognition processes performed by the speech recognition units 21-1, ..., 21-N is the time t of the speech recognition data obtained by the speech recognition process with the highest recognition performance. b It is corrected using this. Therefore, in the embodiment of the present invention, it is possible to absorb differences in the recognition performance of speech recognition processing depending on the type of video content of the live broadcast program. In other words, the speech recognition processing with the highest recognition performance is used depending on the type of video content of the live broadcast program, so the time t of the speech recognition data obtained therefrom b This is the time t of the raw subtitle data a. a This results in high accuracy when used in this way. As a result, it is possible to suppress the display delay of live subtitles relative to the program content with high precision.
[0091] In the content distribution device 1 shown in Figure 1, the subtitle processing unit 12 generates raw subtitle data a' assuming the generation of subtitles for internet distribution, and sends the raw subtitle data a' to the distribution server 2. Alternatively, the subtitle processing unit 12 of the content distribution device 1 may perform processing for other applications, such as re-multiplexing the raw subtitle data a' with a signal for a broadcast system (e.g., an SDI signal).
[0092] [Other content distribution systems] Next, a content distribution system including a receiving device according to an embodiment of the present invention will be described. Figure 8 is a schematic diagram showing an example of the overall configuration of a content distribution system including a receiving device according to an embodiment of the present invention, and a block diagram showing an example of the configuration of the receiving device.
[0093] Similar to Figure 1, this content distribution system is a system that distributes video content of live broadcast programs over the internet via an IP network, that is, a system that performs live video streaming, and is comprised of a content distribution device 101, a distribution server 102, and a receiving device 3.
[0094] Comparing the content distribution system shown in Figure 1 with the content distribution system shown in Figure 8, in Figure 1, the content distribution device 1 matches the raw subtitle data a with multiple speech recognition data b1,...,bN, according to the matching result, and the time t of the raw subtitle data a a The raw subtitle data a' is corrected and raw subtitle data a' is generated. In contrast, in Figure 8, the receiving device 3, according to the matching result between the raw subtitle data a and the multiple speech recognition data b1,...,bN, determines the time t of the raw subtitle data a. a Correct the data and generate raw subtitle data a'.
[0095] The content distribution device 101 is a conventional content distribution device. The content distribution device 101 receives a broadcast transmission signal of video content from an external source, encodes the broadcast transmission signal, divides it into multiple files, and generates distribution data D of multiple files. The content distribution device 101 transmits the distribution data D to the distribution server 102.
[0096] The distribution server 102 is a conventional distribution server. The distribution server 102 receives video content distribution data D from the content distribution device 101 and stores it in memory. Here, in the distribution data D stored in memory, the time t of the raw subtitle data a included in the distribution data D a The subtitles are delayed relative to the time of the corresponding video and audio (video and audio included in the distribution data D). In other words, when the distribution data D stored on the distribution server 102 is viewed, the subtitles will be displayed with a delay relative to the video and audio.
[0097] The receiving device 3 is, for example, a video player such as a smartphone, television, or recorder. It obtains a playlist (not shown) from the content distribution device 101 via the distribution server 102 and the IP network, and understands the file structure based on the playlist. The receiving device 3 then obtains the IP content, including the distribution data D, via the IP network using HTTP (Hypertext Transfer Protocol) based on the playlist. However, the receiving device 3 is not limited to the playlist format; for example, it may prepare information about the IP content, including the distribution data D, required for each program or time slot, and obtain the target IP content based on that information.
[0098] The receiving device 3 decodes the distribution data D contained in the IP content according to the time in the playlist, and matches the raw subtitle data a generated by the decoding with each of the speech recognition data b1, ..., bN obtained by multiple different speech recognition processes on the audio. Then, the receiving device 3 matches the time t contained in the speech recognition data with the highest matching rate. b Using the time t of the raw subtitle data a a This corrects the error. In addition, the receiving device 3 delays the video and audio generated by decoding by the time required for speech recognition processing, etc. This allows the raw subtitle data a' to be synchronized with the program content of the video content in the distribution data D.
[0099] The receiving device 3 plays the content by displaying the video and the subtitles of the raw subtitle data a' on the screen, and also outputting audio.
[0100] As a result, the receiving device 3 can play video content with minimal delay in subtitle display relative to the video and audio. The smaller this subtitle display delay, the easier it is for the user to understand the program content. This is especially effective for people with hearing impairments, as live subtitles play a significant role in helping them understand the program content.
[0101] [Receiving device 3] Next, a receiving device 3 according to an embodiment of the present invention will be described. In Figure 8, the receiving device 3 includes a receiving unit 40, a decoder 41, a subtitle processing unit 42, a delay unit 43, and a display unit 44.
[0102] The receiving unit 40 receives IP content, including distribution data D, from the distribution server 102 via the IP network, performs reception processing, and outputs the distribution data D to the decoder 41.
[0103] The decoder 41 receives the distribution data D from the receiver 40, decodes the distribution data D, combines it, and generates a broadcast signal. The decoder 41 then extracts the video and audio signals from the broadcast signal, as well as the audio subtitle signal, outputs the video and audio signals to the delay unit 43, and outputs the audio subtitle signal to the subtitle processing unit 42. Here, the raw subtitle data a included in the audio subtitle signal is delayed relative to the corresponding audio. That is, the time t included in the raw subtitle data a a However, the time of the corresponding audio t b It is lagging behind. Therefore, if viewing continues in this state, the subtitles will appear with a delay compared to the audio of the video.
[0104] The subtitle processing unit 42 corresponds to the subtitle processing unit 12 shown in Figure 1. The subtitle processing unit 42 receives the audio subtitle signal from the decoder 41, extracts raw subtitle data a from the audio subtitle signal, and performs multiple different speech recognition processes on the audio contained in the audio subtitle signal to generate speech recognition data b1,...,bN.
[0105] The subtitle processing unit 42 performs text matching between the raw subtitle data a and each of the speech recognition data b1, ..., bN. Then, the subtitle processing unit 42 determines the speech recognition data with the highest text matching rate among the speech recognition data b1, ..., bN, and the time t of that speech recognition data. b Using the time t of the raw subtitle data a a By correcting this, new raw subtitle data a' is generated. The subtitle processing unit 42 outputs the raw subtitle data a' to the display unit 44.
[0106] The subtitle processing unit 42 outputs a completion of generation to the delay unit 43 when the generation of raw subtitle data a' is complete. The completion of generation is used in the delay unit 43 to delay the video and audio signals it receives by the time between the time the subtitle processing unit 42 receives the audio subtitle signal and the time it outputs the raw subtitle data a'. Details of the subtitle processing unit 42 will be described later.
[0107] Furthermore, the subtitle processing unit 42 may count the time from when it receives the audio subtitle signal until the generation of the raw subtitle data a' is completed, and output the counted time as a delay time to the delay unit 43 when the generation of the raw subtitle data a' is completed.
[0108] The delay unit 43 receives video and audio signals from the decoder 41 and holds them in a buffer. When the delay unit 43 receives a signal from the subtitle processing unit 42 indicating that generation is complete, it reads the video and audio signals corresponding to the generated raw subtitle data a' from the buffer and outputs the video and audio signals to the display unit 44.
[0109] Furthermore, the delay unit 43 calculates the time from when it holds the video and audio signal corresponding to the completion of generation in the buffer until it reads it out as the delay time. Then, the delay unit 43 holds the video and audio signal in the buffer until it receives input from the subtitle processing unit 42 that the next generation completion has been received, and when the delay time has elapsed, it reads the video and audio signal from the buffer and outputs it to the display unit 44. Details of the delay unit 43 will be described later.
[0110] Furthermore, when the delay unit 43 receives a delay time input from the subtitle processing unit 42, it reads and outputs the video and audio signals already held in the buffer after the delay time has elapsed since they were held in the buffer. The delay unit 43 also reads and outputs any new video and audio signals held in the buffer after the delay time has elapsed since they were held in the buffer. When the delay unit 43 receives a new delay time input from the subtitle processing unit 42, it performs the same processing as described above using the new delay time.
[0111] The display unit 44 receives raw subtitle data a' from the subtitle processing unit 42 and video and audio signals from the delay unit 43, and plays back and displays the video and audio signals and the raw subtitle data a'. Note that the display unit 44 may be a separate device (display device) from the receiving device 3. In this case, the receiving device 3 will output the video and audio signals and the raw subtitle data a' to the display device.
[0112] (Subtitle processing 42) Next, the subtitle processing unit 42 shown in Figure 8 will be described in detail. Figure 9 is a block diagram showing an example configuration of the subtitle processing unit 42 provided in the receiving device 3. This subtitle processing unit 42 includes a subtitle extraction unit 50, speech recognition units 51-1, ..., 51-N, and a matching unit 52. The subtitle processing unit 42 performs the same processing as the subtitle processing unit 12 shown in Figure 2, and further outputs a completion of generation to the delay unit 43 when the generation of raw subtitle data a' is completed.
[0113] The subtitle extraction unit 50 receives the audio subtitle signal from the decoder 41, performs the same processing as the subtitle extraction unit 20 shown in Figure 2, and outputs the raw subtitle data a to the matching unit 52. A detailed explanation of the processing of the subtitle extraction unit 50 is omitted.
[0114] The speech recognition units 51-1, ..., 51-N receive the speech subtitle signal from the decoder 41, perform the same processing as the speech recognition units 21-1, ..., 21-N shown in Figure 2, and output the speech recognition data b1, ..., bN to the matching unit 52. The explanation of the processing of the speech recognition units 51-1, ..., 51-N is omitted.
[0115] The matching unit 52 receives raw subtitle data a from the subtitle extraction unit 50 and speech recognition data b1,...,bN from the speech recognition units 51-1,...,51-N, performs the same processing as the matching unit 22 shown in Figure 2, and outputs the raw subtitle data a' to the display unit 44. The explanation of the processing of the matching unit 52 is omitted.
[0116] The matching unit 52 further outputs a completion of generation to the delay unit 43 when the generation of the raw subtitle data a' is complete.
[0117] Furthermore, the matching unit 52 may, similar to the matching unit 22 shown in Figures 2 and 3, use a program script input from an external source as the correct answer data instead of the raw subtitle data a, and perform text matching with each of the speech recognition data b1,...,bN.
[0118] (Delay section 43) Next, the delay unit 43 shown in Figure 8 will be described in detail. Figure 10 is a flowchart showing an example of the processing of the delay unit 43.
[0119] The delay unit 43 receives the video and audio signals from the decoder 41 (step S1001) and holds the video and audio signals in a buffer (step S1002).
[0120] The delay unit 43 determines whether or not it has received input from the subtitle processing unit 42 indicating that generation is complete (step S1003). If the delay unit 43 determines in step S1003 that it has not received input indicating that generation is complete (step S1003:N), it proceeds to step S1001 and performs the processing in steps S1001 and S1002.
[0121] If the delay unit 43 determines in step S1003 that generation is complete (step S1003:Y), it reads the video and audio signal corresponding to the generated raw subtitle data a' from the buffer and outputs it to the display unit 44 (step S1004).
[0122] As described above, according to the receiving device 3 of the embodiment of the present invention, the decoder 41 decodes the distribution data D to generate a broadcast signal and extracts the video audio signal and the audio subtitle signal from the broadcast signal. The subtitle extraction unit 50 of the subtitle processing unit 42 extracts the raw subtitle data a from the audio subtitle signal. In addition, the speech recognition units 51-1, ..., 51-N apply a known speech recognition process different from that of the other components to the audio included in the audio subtitle signal and generate speech recognition data b1, ..., bN.
[0123] The matching unit 52 calculates the text matching rate for each of the speech recognition data b1, ..., bN with the raw subtitle data a. Then, the matching unit 52 determines the speech recognition data with the highest text matching rate among the speech recognition data b1, ..., bN, and the time t of that speech recognition data. b Using the time t of the raw subtitle data a a By correcting this, new raw subtitle data a' is generated and output to the display unit 44.
[0124] Furthermore, the matching unit 52 outputs a completion message to the delay unit 43 when the generation of the raw subtitle data a' is complete.
[0125] The delay unit 43 holds the video and audio signals in a buffer, and when it receives input from the subtitle processing unit 42 that the generation is complete, it reads the video and audio signals corresponding to the completed raw subtitle data a' from the buffer and outputs them to the display unit 44.
[0126] Thus, the time t of the raw subtitle data a This is the time t of the speech recognition data obtained by the speech recognition processing with the highest recognition performance. b This is used to correct the data, and new raw subtitle data a' is generated. This makes it possible to suppress the display delay of raw subtitles relative to the program content with high accuracy in content distribution systems that distribute video content of live broadcast programs over the internet, enabling the provision of programs that are easier to understand.
[0127] Furthermore, similar to the content distribution device 1 shown in Figure 1, the display delay of live subtitles relative to the program content can be suppressed with high precision depending on the type of video content of the live broadcast program.
[0128] Furthermore, the video and audio signals are delayed in the delay unit 43 by the time it takes for the subtitle processing unit 42 to generate the raw subtitle data a'. As a result, the video and audio signals and the raw subtitle data a' can be synchronized, and the display unit 44 can play back the synchronized video, audio, and subtitles.
[0129] Furthermore, while the content distribution system shown in Figure 8 is a system that distributes IP content via an IP network, it can also be applied to systems that transmit IP content via broadcast waves. In this case, the receiving unit 40 of the receiving device 3 receives broadcast waves containing IP content transmitted from the broadcasting station and performs reception processing such as decoding.
[0130] While embodiments of the present invention have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, or equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments of the present invention described above can be combined in any way without departing from the spirit of the invention.
[0131] Furthermore, a standard computer can be used as the hardware configuration for the content distribution device 1 and the receiving device 3 according to the embodiments of the present invention. The content distribution device 1 and the receiving device 3 are composed of a computer equipped with a CPU, a volatile storage medium such as RAM, a non-volatile storage medium such as ROM, and an interface.
[0132] The functions of the distribution unit 10, encoder 11, and subtitle processing unit 12 provided in the content distribution device 1 are realized by having the CPU execute a program that describes these functions.
[0133] Furthermore, the functions of the receiving unit 40, decoder 41, subtitle processing unit 42, delay unit 43, and display unit 44 provided in the receiving device 3 are each realized by having the CPU execute a program that describes these functions.
[0134] These programs are stored in the aforementioned storage medium and are read and executed by the CPU. These programs can also be stored and distributed on storage media such as magnetic disks (floppy disks, hard disks, etc.), optical disks (CD-ROMs, DVDs, etc.), and semiconductor memory, and can be transmitted and received via a network. [Explanation of symbols]
[0135] 1,101 Content distribution device 2,102 distribution servers 3. Receiving device 10 Distribution section 11 encoders 12,42 Subtitle Processing Unit 20,50 Subtitle extraction part 21-1,···,21-N,51-1,···,51-N Voice Recognition Unit 22,52 Matching Department 30 Input section 31. Voice Recognition Judgment Unit 32 Matching Processing Unit 33 Communications Department 40 Receiver 41 Decoder 43 Delay section 44 Display section 100 Receiver a,a' Raw subtitle data b1,···,bN speech recognition data
Claims
1. A content distribution device that, when distributing video content of a live broadcast program over the internet, receives a broadcast transmission signal including the video content, generates distribution data based on the broadcast transmission signal, and corrects the subtitle time information of the live subtitle data included in the broadcast transmission signal, An encoder that encodes the broadcast transmission signal and generates the distribution data, The system includes a subtitle processing unit that extracts the raw subtitle data from the broadcast transmission signal and corrects the subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen, thereby generating new raw subtitle data. The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast transmission signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast transmission signal and generates a number of different speech recognition data, For each of the multiple different speech recognition data generated by the speech recognition unit, the speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the target for matching. A matching unit that calculates a text matching rate with the raw subtitle data for each of the multiple different matching targets corresponding to the multiple different speech recognition data, determines the matching target with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using the speech time information indicating the time when the audio of the matching target is output, and generates new raw subtitle data. A content distribution device characterized by having the following features.
2. A receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back video, audio, and subtitles included in the broadcast signal, A decoder that decodes the aforementioned IP content and generates the aforementioned broadcast signal, The system includes a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data. The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal and generates a number of different speech recognition data, For each of the multiple different speech recognition data generated by the speech recognition unit, the speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the target for matching. A matching unit that calculates a text matching rate with the raw subtitle data for each of the multiple different matching targets corresponding to the multiple different speech recognition data, determines the matching target with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using the speech time information indicating the time when the audio of the matching target is output, and generates new raw subtitle data. A receiving device characterized by having the following features.
3. A receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back video, audio, and subtitles included in the broadcast signal, A decoder that decodes the aforementioned IP content and generates the aforementioned broadcast signal, The system includes a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data. The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal and generates a number of different speech recognition data, The matching unit includes: for each of the plurality of different speech recognition data generated by the speech recognition unit, it calculates a text matching rate between the raw subtitle data extracted by the subtitle extraction unit, determines the speech recognition data with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using audio time information indicating the time when the audio of the speech recognition data is output, and generates new raw subtitle data. Furthermore, the receiving device is characterized by comprising a delay unit that delays the broadcast signal generated by the decoder by the time from when the subtitle processing unit receives the broadcast signal until it outputs the new raw subtitle data.
4. A receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back video, audio, and subtitles included in the broadcast signal, A decoder that decodes the aforementioned IP content and generates the aforementioned broadcast signal, The system includes a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen to generate new raw subtitle data, and outputs the new raw subtitle data. The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal and generates a number of different speech recognition data, For each of the multiple different speech recognition data generated by the speech recognition unit, the speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the target for matching. The matching unit includes: for each of the multiple different matching targets corresponding to the multiple different speech recognition data, it calculates a text matching rate with the raw subtitle data; it determines the matching target with the highest text matching rate; it corrects the subtitle time information of the raw subtitle data using speech time information indicating the time when the audio of the matching target is output; and it generates new raw subtitle data. Furthermore, the receiving device is characterized by comprising a delay unit that delays the broadcast signal generated by the decoder by the time from when the subtitle processing unit receives the broadcast signal until it outputs the new raw subtitle data.
5. A computer comprising a content distribution device that, when distributing video content of a live broadcast program over the internet, receives a broadcast transmission signal including the video content, generates distribution data based on the broadcast transmission signal, and corrects the subtitle time information of the live subtitle data included in the broadcast transmission signal, An encoder that encodes the broadcast transmission signal and generates the distribution data, and A program that functions as a subtitle processing unit to extract the raw subtitle data from the broadcast transmission signal and correct the subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen, thereby generating new raw subtitle data. The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast transmission signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast transmission signal and generates a number of different speech recognition data, For each of the multiple different speech recognition data generated by the speech recognition unit, the speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the target for matching. A matching unit that calculates a text matching rate with the raw subtitle data for each of the multiple different matching targets corresponding to the multiple different speech recognition data, determines the matching target with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using the speech time information indicating the time when the audio of the matching target is output, and generates new raw subtitle data. A program characterized by having the following features.
6. A computer that constitutes a receiving device that receives IP content including video content of a live broadcast program, decodes the IP content to generate a broadcast signal, and plays back the video, audio, and subtitles contained in the broadcast signal, A decoder that decodes the aforementioned IP content and generates the aforementioned broadcast signal, and A program that functions as a subtitle processing unit that receives the broadcast signal generated by the decoder, extracts raw subtitle data from the broadcast signal, corrects subtitle time information indicating the time when the subtitles of the raw subtitle data are displayed on the screen, generates new raw subtitle data, and outputs the new raw subtitle data, The subtitle processing unit, A subtitle extraction unit that extracts the raw subtitle data from the broadcast signal, A speech recognition unit that applies a predetermined number of different speech recognition processes to the audio contained in the broadcast signal and generates a number of different speech recognition data, For each of the multiple different speech recognition data generated by the speech recognition unit, the speech recognition data is divided into units of the same number of characters as the raw subtitle data extracted by the subtitle extraction unit, generating multiple different divided speech recognition data. The raw subtitle data is used as the ground truth data, and the similarity between the ground truth data and each of the multiple different divided speech recognition data is calculated. The divided speech recognition data with the highest similarity is determined to be the target for matching. A matching unit that calculates a text matching rate with the raw subtitle data for each of the multiple different matching targets corresponding to the multiple different speech recognition data, determines the matching target with the highest text matching rate, corrects the subtitle time information of the raw subtitle data using the speech time information indicating the time when the audio of the matching target is output, and generates new raw subtitle data. A program characterized by having the following features.