Screen control device and program
The screen control device optimizes closed caption display based on open caption presence and speech recognition accuracy, addressing low recognition rates and interference issues to improve viewer comprehension and satisfaction.
Patent Information
- Application Number
- JP2021190935
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-11-25
AI Technical Summary
Conventional speech recognition technologies struggle with low recognition rates for non-broadcast announcer speech, leading to difficulties in understanding and interference between closed and open captions, especially in environments with varying audio quality, and result in reduced viewer satisfaction.
A screen control device that integrates open caption detection, speech recognition, and subtitle generation, dynamically controlling closed caption display based on recognition accuracy and open caption presence to avoid interference and adjust display positions or patterns accordingly.
Enhances viewer understanding by minimizing caption interference and maintaining subtitle visibility even with low recognition accuracy, ensuring clear program content delivery.
Smart Images

Figure 0007784871000005 
Figure 0007784871000006 
Figure 0007784871000007
Abstract
Description
[Technical Field]
[0001] The present invention relates to a screen control device and a program. [Background technology]
[0002] In content provision businesses, such as those providing services via broadcasting or internet distribution, there is a desire to automatically generate subtitles using speech recognition results and add them to video. The background to this is that, for example, relatively small broadcasting stations (including local stations of broadcasting companies connected via networks across Japan) find it difficult to secure the equipment and personnel required for manual subtitle generation. Furthermore, even if subtitles are generated using a respeak method, it is difficult to secure respeakers. On the other hand, there is also a need for subtitled broadcasts, even if the accuracy is lower, for people with hearing impairments, for example. These circumstances give rise to the need to recognize the speech contained in video and add subtitles using the speech recognition results directly, without the need for an operator or other intervention.
[0003] Generally, if the recognition rate (the rate at which correct answers are output as recognition results) of the speech recognition process is above a certain value (for example, 90% or above), the subtitle text is of high value even if it contains some recognition errors. According to Non-Patent Document 1, when such a recognition rate is achieved, even if the recognition result contains errors, the viewer of the displayed subtitle text can mentally fill in the errors with the correct words.
[0004] Conversely, if the speech recognition rate is too low (for example, around 70% or lower), it becomes difficult for viewers of the subtitle text to mentally complete the sentences and understand them as correct words. For example, speech recognition rates tend to be high for parts spoken by announcers, but low for parts spoken by other people (such as street interviews). In other words, if you read the subtitle text during parts with low speech recognition rates, such as street interviews, it can be difficult to understand the program content.
[0005] In other words, it is necessary to control the display of subtitles according to the speech recognition rate.
[0006] Patent Document 1 describes a technology that estimates in real time whether the speech recognition rate is declining while performing speech recognition processing. This technology makes it possible to control so that the recognition result is not output (displayed) when the speech recognition rate is declining.
[0007] Patent Document 2 describes a technology for detecting noise in the external environment and controlling whether closed captions are displayed based on the detection result. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Publication No. 2020-187313 [Patent Document 2] Japanese Patent Application Laid-Open No. 2005-064599 [Non-patent literature]
[0009] [Non-Patent Document 1] Tatsuya Kawahara, "Advances in Speech Recognition of Spoken Language: From Proceedings of Parliamentary Meetings to Subtitling of Lectures and Presentations," Media Education Research, Vol. 9, No. 1, pp. S1-S8, 2012. Summary of the Invention [Problem to be solved by the invention]
[0010] However, conventional technologies have the following problems. For example, speech recognition rates tend to be low for speech by people other than broadcast station announcers (e.g., speech by people being interviewed in street interviews). Furthermore, such speech is often difficult for people to understand due to differences in the audio recording environment and pronunciation. In such cases, open captions may be added as a content enhancement. Open captions are subtitles whose display cannot be switched on or off. In other words, open captions are information included in video signals. Even in news programs and other programs on relatively small broadcast stations, open captions are added for approximately 5% of the time. Thus, open captions play an important role in conveying the content of speech to viewers in a more understandable manner.
[0011] In the technology described in Patent Document 1, the content of closed captions is controlled based solely on the speech recognition rate. This means that there is a possibility that the content of closed captions and the content of open captions may interfere with each other. In other words, when closed captions and open captions are displayed simultaneously, similar but different subtitles are displayed on the screen at the same time, which may hinder the viewer's understanding of the program.
[0012] Furthermore, if only the technology described in Patent Document 1 is used without using open captions, there is a problem that if the recognition rate drops, the subtitle information will no longer be displayed, which will reduce viewer satisfaction.
[0013] Furthermore, the technology described in Patent Document 2 has a problem in that it is not possible to control the display of closed captions based on whether or not open captions are displayed.
[0014] The present invention was made based on the above-mentioned problem recognition, and aims to provide a screen control device and program that can automatically control the display of subtitles (closed captions) based on speech recognition results, depending on whether or not open captions are available. [Means for solving the problem]
[0015] [1] In order to solve the above problem, a screen control device according to one aspect of the present invention includes an open caption detection unit that detects an open caption area included in an input image frame; a speech recognition unit that performs a recognition process on the input speech and outputs text that is the recognition result; a subtitle generation unit that generates closed caption subtitles based on the text output by the speech recognition unit; a subtitle display control unit that controls at least one of whether or not to display the closed caption subtitles or the display position of the closed caption subtitles, depending on whether or not the open caption detection unit has detected the open caption area; and a subtitle display unit that displays the closed caption subtitles in accordance with the control of the subtitle display control unit.
[0016] [2] Furthermore, one aspect of the present invention is that the above-mentioned screen control device further comprises a recognition accuracy determination unit that outputs information regarding the accuracy of the recognition process by grasping the status of the recognition process in the voice recognition unit, and the subtitle display control unit further controls at least one of whether to display the closed caption subtitles or whether to display a specific pattern instead of the closed caption subtitles based on the information regarding the accuracy received from the recognition accuracy determination unit.
[0017] [3] Furthermore, one aspect of the present invention is that in the above-mentioned screen control device, when the open caption detection unit detects the open caption area, the subtitle display control unit controls the closed caption subtitles to be displayed in a position that does not cause interference with the open caption area.
[0018] [4] In another aspect of the present invention, in the above-mentioned screen control device, the subtitle display control unit controls the closed caption subtitles to be displayed at a position overlapping the display of the image frame.
[0019] [5] In another aspect of the present invention, in the above-mentioned screen control device, the subtitle display control unit controls the closed caption subtitles to be displayed at a position that does not overlap with the display of the image frame.
[0020] [6] In addition, in one aspect of the present invention, in the above-mentioned screen control device, the subtitle display control unit (1) controls, when the open caption detection unit does not detect the open caption area, (1A) based on information regarding the accuracy received from the recognition accuracy determination unit, to display the specific pattern instead of the closed caption subtitles if the accuracy is worse than a predetermined first threshold, and (1B) based on information regarding the accuracy received from the recognition accuracy determination unit, to display the closed caption subtitles if the accuracy is equal to or better than the first threshold, and (2) controls the open caption detection unit to display the specific pattern instead of the closed caption subtitles if the accuracy is worse than a predetermined first threshold, based on information regarding the accuracy received from the recognition accuracy determination unit. When the unit detects the open caption area, (2A) based on the information regarding the accuracy received from the recognition accuracy determination unit, if the accuracy is worse than a predetermined second threshold, the unit controls so that the closed caption subtitles are not displayed, and (2B) based on the information regarding the accuracy received from the recognition accuracy determination unit, if the accuracy is equal to or better than the second threshold, the unit controls so that the closed caption subtitles are displayed in a position that does not cause interference with the open caption area, where the second threshold corresponds to an accuracy equal to or better than the first threshold.
[0021] [7] Another aspect of the present invention is a program for causing a computer to function as a screen control device, including: an open caption detection unit that detects an open caption area contained in an input image frame; a speech recognition unit that performs a recognition process on the input speech and outputs text that is the recognition result; a subtitle generation unit that generates closed caption subtitles based on the text output by the speech recognition unit; a subtitle display control unit that controls at least one of whether or not to display the closed caption subtitles or the display position of the closed caption subtitles, depending on whether or not the open caption detection unit has detected the open caption area; and a subtitle display unit that displays the closed caption subtitles in accordance with the control of the subtitle display control unit. [Effects of the Invention]
[0022] According to the present invention, the screen control device can avoid interference between open caption subtitles and closed caption subtitles in a video section consisting of image frames containing open captions. [Brief explanation of the drawings]
[0023] [Figure 1] 1 is a block diagram showing a schematic functional configuration of a screen control device according to an embodiment of the present invention; [Figure 2] 2 is a block diagram showing a more detailed configuration of a part of the functions of a voice recognition unit of the screen control device according to the embodiment. FIG. [Figure 3] FIG. 2 is a block diagram showing the configuration of a content supply system including a screen control device according to the embodiment. [Figure 4] 10 is a flowchart (1 / 2) showing the procedure of processing by the screen control device according to the embodiment. [Figure 5] 10 is a flowchart (2 / 2) showing the procedure of processing by the screen control device according to the embodiment. [Figure 6]10 is a schematic diagram showing an example of the position of a target area in which an open caption detection unit according to the embodiment detects open captions. FIG. [Figure 7] 10 is a schematic diagram showing an example of the relationship between the position of an open caption detected by an open caption detection unit and the position at which a closed caption is displayed in the embodiment. FIG. [Figure 8] 10 is a schematic diagram showing another example of the relationship between the position of an open caption detected by the open caption detection unit and the position at which a closed caption is displayed in the same embodiment. FIG. [Figure 9] 2 is a block diagram showing an example of the internal configuration of each device of a screen control device, a video supply device, and an audio supply device in the embodiment. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0024] Next, an embodiment of the present invention will be described with reference to the drawings. A screen control device 1 of this embodiment controls the display of closed caption subtitles on a screen based on acquired frame images and audio. Closed caption subtitles are generated by performing speech recognition processing.
[0025] Specifically, the screen control device 1 determines whether or not open caption subtitles are included in the frame image, and if open caption subtitles are included, detects their position on the screen. Furthermore, while performing a recognition process on the acquired voice, the screen control device 1 determines whether or not the recognition accuracy in the recognition process has decreased. The screen control device 1 determines whether or not the accuracy of voice recognition has decreased in real time.
[0026] In a situation where open caption subtitles are displayed in an image and it is expected that the recognition accuracy has decreased, the screen control device 1 may perform control so as not to display closed caption subtitles. Furthermore, in a situation where open caption subtitles are displayed in an image and it is not determined that the recognition accuracy has decreased, the screen control device 1 performs control so that the display position of closed caption subtitles is different from the display position of open caption subtitles. In other words, the screen control device 1 controls the display positions so that the open caption subtitles and closed caption subtitles do not interfere with each other. Furthermore, instead of displaying closed caption subtitles based on the results of speech recognition processing, the screen control device 1 may output a specific pattern indicating that speech recognition is difficult (low recognition accuracy). The specific pattern indicating that speech recognition is difficult is, for example, a specific character string such as "...". The screen control device 1 may process, for example, image frames and audio extracted from television programs or video streamed over the Internet.
[0027] FIG. 1 is a block diagram showing a schematic functional configuration of a screen control device according to this embodiment. As shown in the figure, the screen control device 1 includes an image frame acquisition unit 11, an open caption detection unit 12, an audio acquisition unit 21, an audio recognition unit 22, a subtitle generation unit 23, a recognition accuracy determination unit 31, a subtitle display control unit 32, and a subtitle display unit 33. Each of these functional units can be realized, for example, by a computer and a program. Each functional unit also includes a storage unit as needed. The storage unit is, for example, a program variable or memory allocated by program execution. Non-volatile storage units such as a magnetic hard disk drive or a solid-state drive (SSD) may also be used as needed. At least some of the functions of each functional unit may be realized as a dedicated electronic circuit rather than a program.
[0028] The image frame acquisition unit 11 acquires image frames from the outside. Image frames are elements that make up video (for example, video content such as a broadcast program). In other words, video is a series of image frames at a predetermined frame rate. Each image frame has time code information that indicates the presentation time. This time code itself may be input from the outside, or may be generated by the screen control device 1 itself. The time code of an image frame is associated with the time code of an audio frame, which will be described below.
[0029] The image frame acquisition unit 11 passes the acquired image frames to the open caption detection unit 12 in a format that allows processing on a frame-by-frame basis. The image frames are acquired at a frequency of, for example, 30 frames or 60 frames per second. In each case, the image frame period is approximately 33 milliseconds (msec) or approximately 17 msec. The frame image frequency is not limited to the example given here, and other frequencies may be used. In this embodiment, the frame length and period are 30 frames per second (approximately 33 msec). The frame length and period are not limited to the example given here, and are arbitrary.
[0030] The open caption detection unit 12 detects open caption subtitles in the image frames acquired by the image frame acquisition unit 11. The open caption detection unit 12 notifies the subtitle display control unit 32 whether or not open caption subtitles are included in the image frames. Furthermore, if open caption subtitles are included in the image frames, the open caption detection unit 12 notifies the subtitle display control unit 32 of information on the position of the open caption subtitles (such as coordinate information indicating the position of the open caption subtitle area). In other words, the open caption detection unit 12 detects the open caption area included in the input image frames.
[0031] Specifically, the open caption detection unit 12 detects open caption subtitles using the following method. Specifically, the open caption detection unit 12 sets a small area called a scanning window within a target image frame and determines whether the scanning window area is part of an open caption while gradually moving the position of the scanning window. One method for determining whether a scanning window area is part of an open caption subtitle is to distinguish the area using a support vector machine (SVM) or the like based on edge-related features. Edge-related features are effective as quantities for distinguishing between text and background images (which may include people, etc.). The open caption detection unit 12 then extracts areas where scanning windows determined to be part of open captions overlap by a predetermined number or more to determine the extracted area. Furthermore, the open caption detection unit 12 removes erroneous detection areas. The determination of whether a detection is erroneous is based on edge-related features and the size of the detected area. If the detected area is too small (smaller than a predetermined threshold), the area is deemed to be erroneously detected. Based on these processes, the open caption detection unit 12 determines whether or not open caption subtitles exist. If it is determined that open caption subtitles exist, the open caption detection unit 12 outputs information about the position of the open caption subtitles. If the area of the open caption subtitles is rectangular, the information about the position of the open caption subtitles may be, for example, the coordinate values of the vertices of the rectangle. Alternatively, it may be the coordinate values of one vertex of the rectangle (for example, the upper left vertex).
[0032] The open caption detection unit 12 may detect an open caption area using other methods, such as using a machine learning model trained based on a set of pixel values of the area and information on the presence or absence of open captions.
[0033] The audio acquisition unit 21 acquires an audio signal. The audio signal acquired by the audio acquisition unit 21 has content associated with the image frame acquired by the image frame acquisition unit 11. The audio acquisition unit 21 may acquire, for example, an analog signal representing an audio waveform. Alternatively, the audio acquisition unit 21 may acquire digital data representing audio. Alternatively, the audio acquisition unit 21 may acquire a time code signal associated with the audio signal. The time code signal enables synchronization between the audio acquired by the audio acquisition unit 21 and the image frame acquired by the image frame acquisition unit 11.
[0034] The voice acquisition unit 21 passes the voice signal acquired from the outside to the voice recognition unit 22 in a format that allows processing in voice frame units. In this embodiment, one frame has a length of 25 milliseconds and starts every 10 milliseconds. In other words, there is a time period in which multiple voice frames overlap. Note that the length and period of the frames are not limited to those exemplified here and may be different.
[0035] The speech recognition unit 22 performs a recognition process on the speech acquired by the speech acquisition unit 21. The speech recognition unit 22 performs a recognition process on the input speech and outputs text as the recognition result. The speech recognition unit 22 generates hypotheses about the speech recognition result based on the acquired speech and an acoustic model and a language model stored in advance. The speech recognition unit 22 also searches for hypotheses based on scores assigned to hypotheses about the recognition result. The speech recognition unit 22 then finds the most likely hypothesis based on the scores and outputs it as the recognition result. The speech recognition unit 22 passes the text as the recognition result to the subtitle generation unit 23. Note that the speech recognition process itself can be realized using existing technology.
[0036] The speech recognition unit 22 has an internal function for searching speech recognition hypotheses for speech recognition processing. This function is called the speech recognition hypothesis search function. The speech recognition hypothesis search function searches for speech recognition hypotheses based on the scores of multiple speech recognition hypotheses related to the speech recognition result for each time interval of the speech signal. The speech recognition hypothesis search function then determines the speech recognition result from the searched speech recognition hypotheses and outputs the speech recognition result. In other words, the speech recognition hypothesis search function generates multiple speech recognition hypotheses based on the speech acquired by the speech acquisition unit 21, calculates the score of each speech recognition hypothesis, and searches for the speech recognition hypotheses based on the score. More specifically, the speech recognition hypothesis search function calculates acoustic features for each frame (time interval) of the input speech. The speech recognition hypothesis search function may also have a voice activity detection function. In other words, the speech recognition hypothesis search function uses the voice activity detection function to extract segments for each utterance from the speech signal that is continuously input in real time. The speech recognition hypothesis search function then performs speech recognition processing for each utterance according to the following procedure. The speech recognition hypothesis search function generates speech recognition hypotheses for a section of an utterance using pre-stored acoustic models and language models. Here, an acoustic model has information on the probabilistic relationship between acoustic features and phonemes. Also, a language model has information on the probability indicating whether a phoneme sequence is a language. A speech recognition hypothesis is a candidate word sequence that can be considered as corresponding to the input speech. For example, a speech recognition hypothesis can be expressed as a network whose states transition probabilistically over time. The speech recognition hypothesis search function calculates a score for each speech recognition hypothesis based on its correspondence with the input speech.
[0037] The speech recognition hypothesis search function searches the network based on the score and finds the most likely recognition result. In the above processing steps, the speech recognition hypothesis search function performs pruning as appropriate. Generally, in large vocabulary speech recognition processing such as that used in subtitling broadcast programs, it is difficult to match speech with all recognition hypotheses for a word string in real time. Therefore, each time a portion of speech is input, the score of the most likely hypothesis is compared with the score of each hypothesis, and hypotheses with low probability are pruned and discarded. The speech recognition hypothesis search function of this embodiment completes the search in real time by pruning hypotheses. The processing of this speech recognition hypothesis search function itself can be realized using existing technology. This pruning method is called beam search. As an example of beam search, when the search beam width is W, the recognition hypothesis search function calculates the score S of the most likely hypothesis as follows: |SS i | <W Score S i Pruning is performed so that only alternative hypothesis i with
[0038] As a result of hypothesis generation and pruning (discarding), the number of speech recognition hypotheses at a given time (a given speech frame) is variable. The number of speech recognition hypotheses at a given time may reach, for example, several thousand. In parts of the input speech that are highly intelligible and can achieve high speech recognition accuracy, the scores of the other alternative hypotheses are sufficiently small compared to the score S of the most likely hypothesis. Therefore, in these parts, the number of alternative hypotheses that survive in the beam is small. On the other hand, in parts where intelligibility is low and recognition accuracy is expected to decrease, there is not enough difference between the scores of the most likely hypothesis and the alternative hypotheses, so the number of alternative hypotheses that remain as candidates in the beam increases.
[0039] In a situation where there are many alternative hypotheses as described above, for example, the search processing time will increase. Therefore, by measuring the processing time required for the search, it is possible to estimate the deterioration of the accuracy of the speech recognition. In other words, the longer the search processing time, the more the accuracy of the speech recognition processing can be considered to decrease. Alternatively, by measuring the number of hypotheses to be searched, it is possible to estimate the deterioration of the accuracy of the speech recognition. In other words, the more the number of hypotheses to be searched, the more the accuracy of the speech recognition can be considered to decrease.
[0040] The subtitle generation unit 23 generates closed caption subtitle data using text resulting from the speech recognition output by the speech recognition unit 22. The subtitle generation unit 23 passes the generated closed caption subtitles to the subtitle display control unit 32. This enables the subtitle display unit 33 to display closed caption subtitles based on the speech recognition results. In other words, the subtitle generation unit 23 generates subtitle data for video content based on the audio contained in the video content.
[0041] The recognition accuracy determination unit 31 determines the recognition accuracy of the speech recognition unit 22. The recognition accuracy determination unit 31 grasps the status of the recognition processing in the speech recognition unit 22 and outputs information regarding the accuracy of the recognition processing. More specifically, the recognition accuracy determination unit 31 determines the degree to which the recognition accuracy of the speech recognition unit 22 has deteriorated. The recognition accuracy determination unit 31 passes information regarding the recognition accuracy of the speech recognition unit 22 to the subtitle display control unit 32. Specifically, the recognition accuracy determination unit 31 grasps the status of the processing in the speech recognition unit 22 and estimates whether the accuracy of the speech recognition processing has deteriorated based on the status. When making this determination, the recognition accuracy determination unit 31 makes a judgment based on how likely the most likely recognition hypothesis is compared to other hypotheses. More specifically, the recognition accuracy determination unit 31 makes the above judgment based on information such as whether there are a large number of speech recognition hypotheses, whether there is a large delay in the time from inputting a speech frame to processing, whether the processing time per speech frame is long, or whether the processor load (e.g., CPU load) for the recognition processing is large. That is, the recognition accuracy determination unit 31 acquires information regarding the number of speech recognition hypotheses from the speech recognition unit 22, and estimates that the greater the number of speech recognition hypotheses, the more the speech recognition accuracy deteriorates. Alternatively, the recognition accuracy determination unit 31 may acquire the amount of delay from the input of a speech frame to its processing, and estimate that the greater the delay, the more the speech recognition accuracy deteriorates. Alternatively, the recognition accuracy determination unit 31 may acquire the processing time per speech frame, and estimate that the longer the processing time per frame, the more the speech recognition accuracy deteriorates. Furthermore, the recognition accuracy determination unit 31 may acquire information regarding the processor load for recognition processing by the speech recognition unit 22, and estimate that the higher the processor load, the more the speech recognition accuracy deteriorates.
[0042] That is, the voice recognition unit 22 outputs information indicating whether or not the recognition accuracy has decreased, and information indicating the degree to which the recognition accuracy has decreased. The recognition accuracy determination unit 31 determines the degree to which the recognition accuracy has decreased based on the information from the voice recognition unit 22. The recognition accuracy determination unit 31 passes information indicating whether or not the accuracy of the voice recognition process has decreased to the subtitle display control unit 32. This information may be a binary determination value indicating whether or not the accuracy of the voice recognition process has decreased. Alternatively, this information may be numerical information indicating the degree to which the accuracy of the recognition process has decreased. Alternatively, the recognition accuracy determination unit 31 may pass numerical information indicating the recognition accuracy itself at that time to the subtitle display control unit 32.
[0043] The subtitle display control unit 32 controls the display of subtitles (closed captions) based on multiple factors. Specifically, the subtitle display control unit 32 controls at least one of whether to display closed caption subtitles or the display position of closed caption subtitles, depending on whether the open caption detection unit 12 has detected an open caption area. The subtitle display control unit 32 may further control at least one of whether to display the closed caption subtitles or whether to display a specific pattern instead of the closed caption subtitles, based on information regarding the accuracy received from the recognition accuracy determination unit. When the open caption detection unit has detected the open caption area, the subtitle display control unit 32 may control the display of the closed caption subtitles to a position that does not interfere with the open caption area.
[0044] The details of the processing of the subtitle display control unit 32 are as follows. The subtitle display control unit 32 controls the display of subtitles based on whether the open caption detection unit 12 has detected open captions in the image frame. The subtitle display control unit 32 also controls the display of subtitles based on the position (coordinates, etc.) in the image frame of the open caption detected by the open caption detection unit 12. The subtitle display control unit 32 also controls the display of subtitles based on information on the accuracy of voice recognition by the voice recognition unit 22, which is passed from the recognition accuracy determination unit 31.
[0045] Based on the above information, the subtitle display control unit 32 controls whether to display or hide closed caption subtitles. The subtitle display control unit 32 also controls whether to display text (the result of speech recognition) passed from the subtitle generation unit 23 or information of a specific pattern as subtitles. Specific examples of information of a specific pattern will be described later. The subtitle display control unit 32 also controls the position at which the closed caption subtitles are displayed. In other words, the subtitle display control unit 32 determines the display position of the closed caption so that the open captions included in the image frame and the closed caption subtitles based on the speech recognition results do not interfere with each other.
[0046] In other words, if the determination result indicates that there is no decrease in the accuracy of the speech recognition process (the accuracy is equal to or greater than a predetermined threshold), the caption display control unit 32 displays the closed caption generated by the caption generation unit 23 as is. If the determination result indicates that there is a decrease in the accuracy of the speech recognition process (the accuracy is less than a predetermined threshold), the caption display control unit 32 displays a specific pattern (a predetermined pattern) indicating that there is a decrease in recognition accuracy. This specific pattern may be a specific character string, a specific image, or the like. As an example, the specific pattern may be the character string "...".
[0047] As described above, the subtitle display control unit 32 may control the display position of closed caption subtitles. Therefore, the subtitle display control unit 32 has a subtitle display position adjustment function. The subtitle display position adjustment function determines the display position of closed caption subtitles based on the display position of the detected open captions and the number of characters in the closed caption subtitles. For example, the subtitle display position adjustment function can calculate the display start position of closed caption subtitles based on the display position of the open captions passed from the open caption detection unit 12, the font size of the closed caption subtitles to be displayed, and the number of characters in the closed caption subtitles. If the open caption subtitle area is near the bottom center of the screen, the subtitle display position adjustment function determines the position so that the closed caption subtitles are displayed above the open caption subtitles.
[0048] The subtitle display unit 33 displays closed caption subtitles under the control of the subtitle display control unit 32. Note that the closed caption subtitles may not be displayed as a result of control by the subtitle display control unit 32. Also, the specific pattern described above may be displayed instead of the closed caption subtitles.
[0049] The screen control device 1 may include a video signal acquisition unit (not shown). The video signal includes an image frame signal and an audio signal. The video signal may be, for example, a signal conforming to a television broadcasting standard, or may be data encoded using a standard encoding method for distribution via communications (such as the Internet). The video signal acquisition unit extracts (separates) frame images and audio from the acquired video signal and passes them to the image frame acquisition unit 11 and audio acquisition unit 21, respectively. The video signal acquisition unit may also extract a time code signal from the video signal along with the frame images and audio, and pass the time code signal to the image frame acquisition unit 11 and audio acquisition unit 21. The video signal acquired by the video signal acquisition unit may be an analog signal or a digital signal. Whether the video signal is an analog signal or a digital signal, the video signal acquisition unit acquires image, audio, and time code information from the video signal. The video signal may be compressed and encoded using a method such as MPEG. "MPEG" is an abbreviation for "moving picture experts group." The MPEG data includes a stream of image frames, a stream of audio, and a stream of time codes.
[0050] The processing of the screen control device 1 having the above-described functional configuration will be explained in further detail below.
[0051] The speech recognition unit 22 analyzes the input speech in units of speech frames. At time t, the speech recognition unit 22 recognizes a set of hypotheses H t The speech recognition unit 22 passes to the recognition accuracy determination unit 31 information indicating whether or not the recognition accuracy has decreased, depending on the status of the search process performed by the speech recognition hypothesis search function.
[0052] In a situation where there is no deterioration in recognition accuracy, the subtitle display control unit 32 controls to display the closed caption subtitles generated by the subtitle generation unit 23 as is. In a situation where there is expected to be a deterioration in recognition accuracy, the subtitle display control unit 32 can control to output a specific pattern such as "....". Even in such a case, if the recognition accuracy improves when the next utterance begins, for example, the subtitle display control unit 32 controls to display the closed caption subtitles generated by the subtitle generation unit 23. In a portion where there is expected to be no deterioration in speech recognition accuracy, the speech recognition unit 22 outputs a word that can be determined at time t, if any. The speech recognition unit 22 then expands the hypothesis set to the next word to generate a hypothesis set H'. t+1 Then, the speech recognition unit 22 calculates the set of hypotheses H' based on the difference in score from the highest-scoring maximum likelihood hypothesis. t+1 By pruning, the hypothesis set H t+1 Then, the speech recognition unit 22 moves on to processing the next speech frame. The speech recognition unit 22 repeats this processing sequentially while advancing the time t.
[0053] 2 is a block diagram showing a configuration of some of the functions of the speech recognition unit 22. The speech recognition unit 22 has a situation grasping unit 220 as a function for grasping the situation of the speech recognition processing and passing it to the recognition accuracy determination unit 31. As shown in the figure, the situation grasping unit 220 is configured to include a recognition hypothesis amount grasping unit 221, a time difference grasping unit 222, a processing time grasping unit 223, and a CPU load grasping unit 224. Note that "CPU" stands for "central processing unit."
[0054] Each of the above four functional units constituting the situation grasping unit 220 grasps the situation of the recognition hypothesis search process by the speech recognition unit 22. Note that in this embodiment, the situation grasping unit 220 includes four units: the recognition hypothesis quantity grasping unit 221, the time difference grasping unit 222, the processing time grasping unit 223, and the CPU load grasping unit 224. However, the situation grasping unit 220 may be configured to include only some of the four functional units listed here. Furthermore, the situation grasping unit 220 may have functions other than the four functional units listed here. In other words, the situation grasping unit 220 may grasp situations other than the recognition hypothesis quantity, the time difference, the processing time, and the CPU load as the situation of the process of searching for recognition hypotheses.
[0055] The recognition hypothesis quantity grasping unit 221 grasps information on the quantity of recognition hypotheses that are search targets in the process of searching for recognition hypotheses by the speech recognition unit 22. Specifically, the recognition hypothesis quantity grasping unit 221 grasps information on the quantity of recognition hypotheses that are search targets in the process of searching for recognition hypotheses by the speech recognition unit 22. Specifically, the recognition hypothesis quantity grasping unit 221 grasps information on the quantity of recognition hypotheses that are search targets in the process of searching for recognition hypotheses by the speech recognition unit 22. t The number of hypotheses (number of active nodes) in t Alternatively, the recognition hypothesis quantity grasping unit 221 may grasp a value expressed by the following equation (1) as the recognition hypothesis quantity.
[0056]
number
[0057] In equation (1), T is a positive integer that is determined appropriately. That is, in this case, the recognition hypothesis quantity grasping unit 221 grasps, as the recognition hypothesis quantity, the sum of the number of hypotheses in each frame from the speech frame at time (tT), which is T before time t, to the speech frames up to time t.
[0058] Furthermore, the recognition hypothesis quantity grasping unit 221 may grasp, as the recognition hypothesis quantity, another numerical value correlated with the number of recognition hypotheses. As an example, the recognition hypothesis quantity grasping unit 221 may grasp, as the recognition hypothesis quantity, the memory capacity for at least temporarily storing hypothesis information in the processing of the recognition hypothesis search function.
[0059] The situation grasping unit 220 can pass information on the recognition hypothesis quantity grasped by the recognition hypothesis quantity grasping unit 221 to the recognition accuracy determining unit 31. The recognition accuracy determining unit 31 estimates whether or not there is a decrease in the accuracy of the speech recognition processing based on this recognition hypothesis quantity. The recognition accuracy determining unit 31 estimates that the recognition accuracy has decreased when the recognition hypothesis quantity is larger than a predetermined threshold. Furthermore, the recognition accuracy determining unit 31 estimates that the recognition accuracy has not decreased when the recognition hypothesis quantity is equal to or smaller than the predetermined threshold.
[0060] The time difference ascertaining unit 222 ascertains the difference between the original time of an audio frame and the time at which recognition processing for that audio frame is executed. The original time of an audio frame is the time at which that audio frame is input to the screen control device 1. In other words, the time difference ascertained by the time difference ascertaining unit 222 corresponds to the time delay from the input time of the audio frame to the processing time (e.g., the start time of processing). The value of the time difference ascertained by the time difference ascertaining unit 222 can be used as an approximate value according to the amount of recognition hypotheses at that time.
[0061] Generally, in hypothesis search in speech recognition processing, increasing the beam width W is expected to improve recognition accuracy. In real-time speech recognition, the beam width W is set to be large enough to prevent delays in the recognition process. However, in a time period in which the number of active nodes continues to increase for a certain period of time, it becomes necessary to calculate the scores of a large number of hypothesis candidates for each speech frame. As a result, the calculation process for recognition cannot be completed within the time of one speech frame, and the processing time of the frame is delayed. Therefore, the time difference grasping unit 222 grasps the difference between the original time of the frame at time t (for example, the time when the speech of the frame is input) and the time when the frame is processed. The time difference grasping unit 222 can obtain information about the current time from a clock function (not shown) of the screen control device 1 itself.
[0062] The situation grasping unit 220 can pass information about the time difference grasped by the time difference grasping unit 222 to the recognition accuracy determining unit 31. The recognition accuracy determining unit 31 estimates whether or not there has been a decrease in the accuracy of the recognition processing by the recognition hypothesis search function based on this time difference (amount of delay). For example, the recognition accuracy determining unit 31 estimates that the recognition accuracy has decreased when the time difference is greater than a predetermined threshold. Furthermore, the recognition accuracy determining unit 31 estimates that the recognition accuracy has not decreased when the time difference is equal to or less than the predetermined threshold.
[0063] The processing time grasping unit 223 grasps the length of time required for recognition processing for each voice frame. This length of processing time can be used to grasp the recognition accuracy as a numerical value corresponding to the amount of recognition hypotheses. As mentioned above, in a section where the number of active nodes has increased, it is necessary to calculate the scores of a large number of hypothesis candidates for each frame. This increases the processing time required for calculations for recognition. Therefore, the processing time grasping unit 223 grasps the time P required for processing the frame at time t. t is acquired from the recognition hypothesis search function. Alternatively, the processing time ascertaining unit 223 may ascertain the value expressed by the following formula (2) as the processing time.
[0064]
number
[0065] In equation (2), T is a positive integer that is determined appropriately. That is, in this case, the processing time ascertaining unit 223 ascertains the total processing time of the frames from the frame at time (tT), which is T before time t, to time t.
[0066] The situation grasping unit 220 can pass information on the processing time grasped by the processing time grasping unit 223 to the recognition accuracy determining unit 31. Based on this processing time, the recognition accuracy determining unit 31 estimates whether or not there has been a decrease in accuracy in the recognition processing by the recognition hypothesis search function. For example, the recognition accuracy determining unit 31 estimates that the recognition accuracy has decreased if the processing time is greater than a predetermined threshold. Furthermore, the recognition accuracy determining unit 31 estimates that the recognition accuracy has not decreased if the processing time is equal to or less than the predetermined threshold.
[0067] The CPU load ascertaining unit 224 ascertains information about the CPU load associated with the processing of the recognition hypothesis search function. The magnitude of the CPU load can be used as a value corresponding to the number of recognition hypotheses. As a prerequisite, in the speech recognition processing of the screen control device 1, at least the processing of the recognition hypothesis search function is implemented as a program executed by the CPU. Specifically, for example, the CPU load ascertaining unit 224 acquires a numerical value of the CPU load provided by a process management function within the operating system (OS). Note that the CPU load ascertaining unit 224 may acquire information about the overall load of the CPU or the overall load of all user processes running on the CPU as an approximation of the CPU load associated with the processing of the recognition hypothesis search function. When the number of hypotheses to be evaluated (the number of active nodes) is large, such as when the processing for recognizing each speech frame does not fit within the time frame of that speech frame, the CPU load associated with the processing of the recognition hypothesis search function becomes very high (possibly reaching or approaching 100%). In other words, it is possible to estimate whether or not the accuracy of the recognition processing is deteriorating based on the CPU load.
[0068] The relationship between the number of active nodes and the recognition accuracy is as follows: Let the set of hypotheses at time t be H t In addition, this hypothesis set H t The number of hypotheses (number of active nodes) that t ) in the interval (time interval) where high recognition accuracy is expected to be obtained, the number of active nodes N(H t) is generally small. In other words, in the section where high recognition accuracy is expected, the number of hypotheses that may conflict with the most likely hypothesis is generally small. Also, in such a situation where recognition accuracy is high, the number of hypotheses N(H t ) is roughly constant around the lower limit. On the other hand, in the section where the recognition accuracy is degraded, the number of active nodes N(H t ) increases. In other words, the number of hypotheses at time t, N(H t ), it can be estimated whether the recognition accuracy is being maintained at a high level or is declining.
[0069] When the recognition accuracy is high, the number of hypotheses N(H t The reason why the number of hypotheses N(H) is relatively small is that the score of the most likely hypothesis is significantly higher (high peak) than the scores of other hypotheses, so many hypotheses are discarded by pruning during the search. Conversely, when the recognition accuracy is low, the number of hypotheses N(H t ) is relatively high because the score of the most likely hypothesis is less prominent than the scores of other hypotheses (less peaky), so more hypotheses survive without being pruned.
[0070] 3 is a block diagram showing a schematic functional configuration of a content supply system according to this embodiment. As shown in the figure, the content supply system 8 includes a screen control device 1, a video supply device 6, and an audio supply device .
[0071] As already explained, the screen control device 1 controls the display of closed caption subtitles on the screen based on the acquired audio. In the content supply system 8, the screen control device 1 acquires image frames supplied by the video supply device 6 and detects open captions in the image frames. The screen control device 1 acquires audio data supplied by the audio supply device 7 and performs a recognition process for the audio. The screen control device 1 also controls the display of closed caption subtitles generated using text resulting from the audio recognition. That is, the screen control device 1 controls the display of closed caption subtitles depending on whether or not open captions are detected in the image frames. The screen control device 1 also controls the display of closed caption subtitles based on information (such as coordinate values) about the position of the open captions detected in the image frames. The screen control device 1 also controls the display of closed caption subtitles depending on the estimated accuracy of the recognition process for the audio supplied by the audio supply device 7 (depending on whether or not the accuracy has decreased).
[0072] The video supply device 6 supplies a sequence of image frames to the screen control device 1.
[0073] The audio supply device 7 supplies audio data (a series of audio frames) to the screen control device 1. The audio supplied by the audio supply device 7 may be a time domain signal or a frequency domain signal.
[0074] The sequence of image frames supplied by the video supply device 6 and the audio supplied by the audio supply device 7 represent the same video content. These image frames and audio can be synchronized using common time code information or the like. The screen control device 1 outputs a sequence of image frames. These image frames may include closed caption subtitles. The position of the closed caption subtitles may be adjusted based on the position of the open caption subtitles. The video (sequence of image frames) output by the screen control device 1 and the audio output by the audio supply device 7 can be viewed as a single video content.
[0075] 4 and 5 are flowcharts showing the processing procedure by the screen control device 1. FIGS. 4 and 5 together form a flowchart showing one procedure, and the procedures shown in these two figures are connected using connectors. The processing procedure of the screen control device 1 will be explained below in accordance with this flowchart. Note that the time t is appropriately initialized before starting this processing. Also, in this embodiment, the audio frame update cycle will be explained as 10 ms (milliseconds), but in reality the length of the frame cycle may be different.
[0076] In step S1, the image frame acquisition unit 11 determines whether or not an image frame update has occurred at time t. If the frame rate is, for example, 30 fps (frames per second), the image frame is updated 30 times per second. Therefore, the image frame is updated once every time the audio frame, which shifts every 10 milliseconds, is updated approximately 3.33 times. If an image frame update has occurred (step S1: YES), the process proceeds to step S2. If an image frame update has not occurred (step S1: NO), the process proceeds to step S5.
[0077] When the process proceeds to step S2, in this step the open caption detection unit 12 determines whether or not open captions exist in the current image frame (the image frame whose update was confirmed in step S1). The method for determining whether or not open captions exist is as will be explained separately. If open caption subtitles exist in the image frame (step S2: YES), the process proceeds to step S3. If open caption subtitles do not exist (step S2: NO), the process proceeds to step S5.
[0078] When the process proceeds to step S3, in this step, open caption detection unit 12 identifies the position of the detected open caption. If the shape of the area in which the open caption is displayed is rectangular, the position is represented, for example, by the coordinates of the upper left point and the lower right point of the area on the screen. If the shape of the area is not rectangular, open caption detection unit 12 identifies information representing the range of the area as appropriate, depending on the shape of the area.
[0079] Next, in step S4, the open caption detection unit 12 stores information about the image frame having the open caption in a storage area. Specifically, the open caption detection unit 12 stores information indicating the presence or absence of the open caption subtitle and information indicating the position of the open caption subtitle in association with the time t of the image frame. After processing in this step, the process proceeds to step S5.
[0080] In step S5, the speech recognition unit 22 identifies a set of recognition hypotheses based on the data of the speech frame at time t and calculates a score for each recognition hypothesis belonging to the set. The speech recognition unit 22 determines a most likely hypothesis based on the score of each recognition hypothesis. The speech recognition unit 22 outputs text of the speech recognition result based on the most likely hypothesis, which is the search result.
[0081] Next, in step S6, the recognition accuracy determination unit 31 grasps the status of the processing by the speech recognition unit 22. Specifically, the recognition accuracy determination unit 31 grasps at least one of the amount of recognition hypotheses, the delay time from input of a speech frame to processing, the processing time of the speech frame, and the CPU load. Note that the recognition accuracy determination unit 31 may grasp the status of the processing of the recognition hypothesis search by the speech recognition unit 22 based on information other than those listed here. Details of the processing by the recognition accuracy determination unit 31 are as described separately.
[0082] Next, in step S7, the subtitle display control unit 32 determines whether or not open captions are present in the screen frame. The determination in step S7 may be made based on the result of the determination in step S2. That is, the subtitle display control unit 32 may determine the presence or absence of open captions based on the information saved in step S4. If open captions are present in the image frame (step S7: YES), the process proceeds to step S8. If it is determined that open captions are not present in the image frame (step S7: NO), the process jumps to step S11 (FIG. 5).
[0083] Next, when the process proceeds to step S8, in that step, the recognition accuracy determination unit 31 determines whether or not the recognition accuracy has decreased below a predetermined threshold (threshold A) (whether or not the recognition accuracy is less than threshold A). If the determination result in step S8 is that the recognition accuracy is less than threshold A (step S8: YES), the process proceeds to step S9. If the recognition accuracy is equal to or greater than threshold A (step S8: NO), the process proceeds to step S10.
[0084] Specifically, the recognition accuracy determination unit 31 determines whether the recognition accuracy is high or low (good or bad) according to the following (1) to (4). Here, it is assumed that the above-mentioned recognition accuracy, amount of recognition hypotheses, amount of time delay, processing time, and CPU load can all be expressed as numerical values.
[0085] (1) The case where the recognition accuracy determination unit 31 determines the magnitude of the recognition accuracy using the recognition hypothesis quantity is as follows: (1-a) The recognition accuracy being less than a threshold (threshold for the recognition accuracy) corresponds to the recognition hypothesis quantity being greater than the threshold (threshold for the recognition hypothesis quantity). (1-b) The recognition accuracy being equal to or greater than the threshold (threshold for the recognition accuracy) corresponds to the recognition hypothesis quantity being equal to or less than the threshold (threshold for the recognition hypothesis quantity).
[0086] (2) When the recognition accuracy determination unit 31 determines the magnitude of the recognition accuracy using the amount of time delay from input of the audio frame to processing, the following applies: (2-a) When the recognition accuracy is less than a threshold (threshold for recognition accuracy), the amount of time delay is greater than the threshold (threshold for time delay). (2-b) When the recognition accuracy is equal to or greater than the threshold (threshold for recognition accuracy), the amount of time delay is equal to or less than the threshold (threshold for time delay).
[0087] (3) When the recognition accuracy determination unit 31 uses the processing time of the audio frame to determine the magnitude of the recognition accuracy, it is as follows: (3-a) When the recognition accuracy is less than a threshold (threshold for recognition accuracy), it means that the processing time is greater than a threshold (threshold for processing time). (3-b) When the recognition accuracy is equal to or greater than the threshold (threshold for recognition accuracy), it means that the processing time is equal to or less than the threshold (threshold for processing time).
[0088] (4) When the recognition accuracy determination unit 31 uses the CPU load to determine the magnitude of the recognition accuracy, it is as follows: (4-a) When the recognition accuracy is less than a threshold (threshold for recognition accuracy), it means that the CPU load is higher than the threshold (threshold for CPU load). (4-b) When the recognition accuracy is equal to or greater than the threshold (threshold for recognition accuracy), it means that the CPU load is equal to or less than the threshold (threshold for CPU load).
[0089] Note that in each of (1) to (4) above, we have explained the case where the recognition accuracy determination unit 31 evaluates the recognition accuracy using a single criterion, but the recognition accuracy determination unit 31 may also measure the recognition accuracy by combining conditions for multiple criterions.
[0090] If the process proceeds to step S9, in that step the subtitle display control unit 32 performs control to hide the closed caption subtitles. As a result of this control, the closed caption subtitles generated by the subtitle generation unit 23 are no longer displayed on the screen. After the process of step S9, the process proceeds to step S13 (FIG. 5).
[0091] If the process proceeds to step S10, the subtitle display control unit 32 calculates the display position of the closed caption subtitles and performs control to display the closed caption subtitles at that position. Specifically, the subtitle display control unit 32 sets the display start position of the closed caption subtitles two lines above the position of the open caption subtitles detected by the open caption detection unit 12. Alternatively, the subtitle display control unit 32 may set the left edge of the screen safety zone as the display start position of the closed caption subtitles. Alternatively, the subtitle display control unit 32 may set another position as the display start position of the closed caption subtitles. In either case, the subtitle display control unit 32 calculates the display position of the closed caption subtitles so that the open caption subtitles and the closed caption subtitles do not interfere with each other. This makes it possible to avoid a situation in which at least one of the open caption subtitles or the closed caption subtitles is difficult for the viewer to view. After the process of step S10, the process proceeds to step S13 (FIG. 5).
[0092] 5, when the process proceeds to step S11, the recognition accuracy determination unit 31 determines whether the recognition accuracy has decreased below a predetermined threshold (threshold B) (whether the recognition accuracy is less than threshold B). If the determination result in step S11 is that the recognition accuracy is less than threshold B (step S11: YES), the process proceeds to step S12. If the recognition accuracy is equal to or greater than threshold B (step S11: NO), the process jumps to the processing of step S13.
[0093] If the process proceeds to step S12, the subtitle display control unit 32 controls the display of a specific pattern in this step. That is, in S12, control is performed so that the specific pattern is output to the outside, regardless of whether or not a recognition result is output from the speech recognition unit 22 at that time, in other words, regardless of whether or not a subtitle is generated by the subtitle generation unit 23. The specific pattern is an arbitrary pattern that is determined appropriately. An example of a specific pattern is a character string such as "...." (a series of periods). This specific pattern indicates that no recognition result is output (that output is suppressed). Note that patterns other than those exemplified here (such as character strings or images) may also be used as the specific pattern. After the process of step S12, the process proceeds to step S13.
[0094] The relationship between threshold A and threshold B may be threshold B<threshold A. In other words, threshold B is a threshold for determining whether the recognition accuracy has further decreased than threshold A. However, threshold B may be equal to threshold A.
[0095] If threshold B<threshold A, the determinations in steps S7, S8, and S11 are divided into cases as shown in Table 1 below, and the respective processes are carried out.
[0096] [Table 1]
[0097] That is, if an open caption is not detected and the recognition accuracy is less than threshold B, the caption display control unit 32 performs control such that a specific pattern is displayed as the process of step S12.
[0098] Furthermore, if no open caption is detected and the recognition accuracy is equal to or greater than threshold B, the caption display control unit 32 performs control such that the generated caption (closed caption) is displayed on the screen as is.
[0099] Furthermore, if open captions are detected and the recognition accuracy is less than threshold B, or if the recognition accuracy is equal to or greater than threshold B and less than threshold A, the caption display control unit 32 performs control in step S9 to hide the closed captions (not display the captions generated by the caption generation unit 23). When hiding the closed captions, a specific pattern indicating that "the captions are turned off" may be optionally displayed.
[0100] Furthermore, if open captions are detected and the recognition accuracy is equal to or greater than threshold A, the subtitle display control unit 32 performs control in step S10 to calculate a display position that will not interfere with the display of the open captions, and then display the closed captions.
[0101] If threshold B=threshold A, the determinations in steps S7, S8, and S11 are divided into cases as shown in Table 2 below, and the respective processes are performed.
[0102] [Table 2]
[0103] In other words, if open captions are not detected and the recognition accuracy is less than threshold B (threshold B is equal to threshold A), the subtitle display control unit 32 performs control to display a specific pattern as the processing of step S12.
[0104] In addition, if open captions are not detected and the recognition accuracy is equal to or greater than threshold B (threshold B is equal to threshold A), the subtitle display control unit 32 controls the generated subtitles (closed captions) to be displayed on the screen as is.
[0105] Furthermore, if open captions are detected and the recognition accuracy is less than threshold B (threshold B is equal to threshold A), the caption display control unit 32 performs control in step S9 to hide the closed captions (not display the captions generated by the caption generation unit 23). When hiding the closed captions, a specific pattern indicating that "the captions are turned off" may be optionally displayed.
[0106] Furthermore, if open captions are detected and the recognition accuracy is equal to or greater than threshold B (threshold B is equal to threshold A), the subtitle display control unit 32 performs control in step S10 to calculate a display position that will not interfere with the display of the open captions, and then display the closed captions.
[0107] In step S13, if there is a confirmed speech recognition result at that time, the caption display unit 33 displays captions (closed captions) based on the recognition result. The closed captions are generated by the caption generation unit 23.
[0108] The subtitle display unit 33 displays the subtitles under the control of the subtitle display control unit 32. That is, when the subtitle display control unit 32 exercises control to hide the closed caption subtitles (processing of step S9), the subtitle display unit 33 does not display the closed caption subtitles. When the subtitle display control unit 32 exercises control to display a specific pattern (processing of step S12), the subtitle display unit 33 displays the specific pattern of subtitles (for example, a pattern such as "..."). When the subtitle display control unit 32 calculates a subtitle position and exercises control to display the closed caption subtitles at that position (processing of step S10), the subtitle display unit 33 displays the closed caption subtitles generated by the subtitle generation unit 23 at that position.
[0109] However, if there is no confirmed recognition result at that point, the caption generation unit 23 does not generate captions, and the caption display unit 33 does not do anything in particular (does not display closed captions).
[0110] After step S13 is completed, the process proceeds to step S14, where the speech recognition unit 22 prunes the set of recognition hypotheses. The speech recognition unit 22 discards hypotheses with relatively low likelihood based on the likelihood of each hypothesis. Specifically, the speech recognition unit 22 performs pruning based on the difference between the score of the most likely hypothesis and the score of each hypothesis. That is, in this pruning process, hypotheses with scores (relatively low scores) that differ more significantly from the score of the most likely hypothesis (relatively high score) are more likely to be discarded. If the scores among hypotheses are relatively uniform, a relatively larger number of hypotheses will survive the pruning process in this step without being pruned. If the score of the most likely hypothesis is relatively peaky, a relatively smaller number of hypotheses will survive the pruning process in this step without being pruned.
[0111] Next, in step S15, the screen control device 1 advances the time t to the next time. Specifically, if the time t is an integer value, the time t is updated so that t:=t+1.
[0112] Next, in step S16, the screen control device 1 determines whether or not a termination condition is met. Examples of the termination condition include input of a stop instruction from outside, or a state in which no utterance that can be a target for voice recognition continues for more than a predetermined time. If the termination condition is met (step S16: YES), the screen control device 1 ends the entire processing of this flowchart. If the termination condition is not met (step S16: NO), the screen control device 1 returns to the processing of step S1 (FIG. 4) to process the next voice frame.
[0113] The method of controlling the subtitle display explained with reference to the flowchart can be summarized as follows: (1) when the open caption detection unit 12 does not detect an open caption area, (1A) based on the information regarding accuracy received from the recognition accuracy determination unit 31, the subtitle display control unit 32 controls so that a specific pattern is displayed instead of the closed caption subtitles if the accuracy is worse than a predetermined first threshold, and (1B) based on the information regarding accuracy received from the recognition accuracy determination unit 31, the closed caption subtitles are displayed if the accuracy is equal to or better than the first threshold. In addition, (2) when the open caption detection unit 12 detects an open caption area, the subtitle display control unit 32 controls (2A) based on information regarding accuracy received from the recognition accuracy determination unit 31, to prevent the closed caption subtitles from being displayed if the accuracy is worse than a predetermined second threshold, and (2B) based on information regarding accuracy received from the recognition accuracy determination unit 31, to display the closed caption subtitles in a position that does not cause interference with the open caption area if the accuracy is equal to or better than the second threshold.
[0114] The second threshold value is a threshold value corresponding to accuracy equal to or better than the first threshold value. The first threshold value corresponds to threshold B in the previous description. The second threshold value corresponds to threshold A in the previous description.
[0115] FIG. 6 is a schematic diagram showing an example of the position of a target area for open caption detection unit 12 to detect open captions. In the figure, 101 is a screen. Area 102 located at the bottom of screen 101 is an area for open caption detection unit 12 to detect open captions. That is, open caption detection unit 12 identifies an area within area 102 that has image features that resemble open captions as an area in which open captions are displayed. Area 103 located within area 102 is an example of an open caption area detected by open caption detection unit 12. As in this example, the position where open captions are displayed is usually near the bottom center of the screen. When open caption detection unit 12 detects area 103 as an open caption area, open caption detection unit 12 passes position information of area 103 (e.g., information representing the coordinates of the four vertices of this rectangle) to caption display control unit 32.
[0116] FIG. 7 is a schematic diagram showing an example of the relationship between the position of detected open captions and the position at which closed captions are displayed. Within screen 101, area 201 is the area of open captions detected by open caption detection unit 12. Open caption detection unit 12 passes information about the position of area 201 to subtitle display control unit 32. Based on the position information of area 201, subtitle display control unit 32 controls closed captions to be displayed within area 202. Area 202 does not overlap with area 201. In other words, when subtitle display control unit 32 controls closed captions to be displayed within area 202, the closed captions (area 202) do not interfere with the detected open captions (area 201). Note that subtitle display control unit 32 determines the size and position of area 202 from, for example, information about the position of area 201, the number of characters in the closed caption subtitles to be displayed, and the font size to be used.
[0117] FIG. 8 is a schematic diagram showing another example of the relationship between the position of detected open captions and the position at which closed captions are displayed. In this example, the subtitle display control unit 32 controls the display of closed caption subtitles outside the area displaying the video content. In this diagram, an area 301 in the screen 101 is an area for displaying the video content (e.g., the video of a broadcast program). An area 302 in the area 301 is an area where the open caption subtitles are displayed. In other words, the area 302 is an area for the open captions detected by the open caption detection unit 12. Furthermore, an area 304 is an area for displaying information other than the video content. In this example, an area 303 is an area for displaying the closed caption subtitles. In other words, the subtitle display control unit 32 controls the display of the closed caption subtitles in the area 303. As in this example, the subtitle display control unit 32 may control the display position so that the closed caption subtitles are displayed outside the area displaying the video (area 301).
[0118] As explained in the example of multiple screens, the subtitle display control unit 32 may control the display of closed caption subtitles so that they overlap with the display of image frames (for example, in the case of Fig. 7), or may control the display of closed caption subtitles so that they do not overlap with the display of image frames (for example, in the case of Fig. 8). That is, in both the case of the example shown in Fig. 7 (where the display position of the closed caption subtitles overlaps with the video content) and the case of the example shown in Fig. 8 (where the display position of the closed caption subtitles does not overlap with the video content), the subtitle display control unit 32 controls the display position of the closed caption subtitles.
[0119] FIG. 9 is a block diagram showing an example of the internal configuration of each of the screen control device 1, video supply device 6, and audio supply device 7 in the above embodiment. Each device can be realized using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, and a bus 906. The computer itself can be realized using existing technology. The central processing unit 901 executes instructions contained in a program read from the RAM 902 or the like. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices. The input / output devices 904 and 905 are input / output devices. Input / output devices 904 and 905 exchange data with the central processing unit 901 via an input / output port 903. A bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from and to RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906.
[0120] At least some of the functions of the screen control device 1, video supply device 6, and audio supply device 7 in the above-described embodiments can be implemented by a computer. In this case, a program for implementing these functions may be recorded on a computer-readable recording medium and then loaded and executed by a computer system. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, CD-ROMs, DVD-ROMs, and USB flash drives, as well as storage devices such as hard disks built into computer systems. In other words, a "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the term "computer-readable recording medium" may also include media that temporarily and dynamically store programs, such as communication lines used when transmitting programs via networks such as the Internet or telephone lines, or media that store programs for a certain period of time, such as volatile memory within a computer system that serves as a server or client. The program may be designed to implement some of the above-described functions, or may be capable of implementing the above-described functions in combination with a program already stored in the computer system.
[0121] Although the embodiment has been described above, the present invention may be realized by a modification of the above embodiment. Furthermore, the specific configuration of the device is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. An example of a modification will be described next.
[0122] [Variations] As a modified example, the screen control device 1 may control the display of closed caption subtitles regardless of the speech recognition accuracy. In this case, the screen control device 1 may have a functional configuration that does not include the recognition accuracy determination unit 31. Alternatively, the subtitle display control unit 32 may not receive information related to the recognition accuracy from the recognition accuracy determination unit 31. In either of these cases, the subtitle display control unit 32 controls the display of closed caption subtitles regardless of the recognition accuracy. In this modified example, the subtitle display control unit 32 controls at least one of whether or not to display closed caption subtitles and the display position, if any, of the closed caption subtitles, depending on whether or not the open caption detection unit 12 has detected an open caption area. Furthermore, in this modified example, when the open caption detection unit 12 has detected an open caption area, the subtitle display control unit 32 may control the display of closed caption subtitles to a position that does not interfere with the open caption area. [Industrial Applicability]
[0123] The present invention can be used, for example, in businesses that create and distribute content (including, but not limited to, broadcasting businesses), but the scope of use of the present invention is not limited to the examples given here. [Explanation of symbols]
[0124] 1 Screen Control Device 6. Video supply equipment 7 Audio supply device 8 Contents Supply System 11 Image frame acquisition unit 12 Open Caption Detector 21 Voice acquisition unit 22 Voice Recognition Unit 23 Subtitle generation section 31 Recognition accuracy judgment section 32 Subtitle display control unit 33 Subtitle display area 220 Situation Assessment Department 221 Recognition hypothesis quantity grasping unit 222 Time difference understanding unit 223 Processing Time Grasping Unit 224 CPU load understanding part 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus
Claims
1. an open caption detection unit that detects an open caption area included in an input image frame; a speech recognition unit that performs a recognition process on an input speech and outputs a text that is a recognition result; a subtitle generation unit that generates closed caption subtitles based on the text output by the speech recognition unit; a caption display control unit that controls at least one of whether or not to display the closed caption subtitles and the display position of the closed caption subtitles when displayed, depending on whether or not the open caption detection unit has detected the open caption area; a caption display unit that displays the closed caption captions under the control of the caption display control unit; A screen control device comprising:
2. a recognition accuracy determination unit that outputs information regarding accuracy of the recognition processing by grasping the status of the recognition processing in the voice recognition unit; Furthermore, the subtitle display control unit further controls at least one of whether to display the closed caption subtitles or whether to display a specific pattern instead of the closed caption subtitles, based on the information related to the accuracy received from the recognition accuracy determination unit. The screen control device according to claim 1 .
3. When the open caption detection unit detects the open caption area, the subtitle display control unit controls the closed caption subtitle to be displayed at a position that does not cause interference with the open caption area. The screen control device according to claim 1 or 2.
4. the subtitle display control unit controls to display the closed caption subtitles at a position overlapping the display of the image frame; The screen control device according to claim 3 .
5. the subtitle display control unit controls to display the closed caption subtitles at a position that does not overlap with the display of the image frame. The screen control device according to claim 3 .
6. The subtitle display control unit (1) When the open caption detection unit does not detect the open caption area, (1A) based on information regarding the accuracy received from the recognition accuracy determination unit, when the accuracy is lower than a predetermined first threshold, control is performed so that the specific pattern is displayed instead of the closed caption subtitle; (1B) based on the information on the accuracy received from the recognition accuracy determination unit, when the accuracy is equal to or better than the first threshold, control is performed so that the closed caption subtitle is displayed; (2) When the open caption detection unit detects the open caption area, (2A) based on the information on the accuracy received from the recognition accuracy determination unit, if the accuracy is lower than a predetermined second threshold, control is performed so that the closed caption subtitle is not displayed; (2B) based on the information on the accuracy received from the recognition accuracy determination unit, when the accuracy is equal to or better than the second threshold, control is performed to display the closed caption subtitle at a position that does not cause interference with the open caption area. That is what it means. the second threshold corresponds to an accuracy equal to or better than the first threshold; The screen control device according to claim 2 .
7. an open caption detection unit that detects an open caption area included in an input image frame; a speech recognition unit that performs a recognition process on an input speech and outputs a text that is a recognition result; a subtitle generation unit that generates closed caption subtitles based on the text output by the speech recognition unit; a caption display control unit that controls at least one of whether or not to display the closed caption subtitles and the display position of the closed caption subtitles when displayed, depending on whether or not the open caption detection unit has detected the open caption area; a caption display unit that displays the closed caption captions under the control of the caption display control unit; A program for causing a computer to function as a screen control device comprising:
Citation Information
Patent Citations
Method for controlling subtitles display for open caption
JP2002344805A
Apparatus, method and program for caption control
JP2005064599A
Digital broadcast receiving device and its control method
JP2007300270A
Voice recognition device, recognition result output control device, and program
JP2020187313A
Digital broadcast receiving apparatus and control method therefor
US20070252913A1