Information processing device, information processing program, and information processing method

The information processing device addresses subtitle errors and workload by separating image and audio data, generating support information for proofreading, and associating it with text data, enhancing subtitle accuracy and efficiency.

JP7764209B2Active Publication Date: 2025-11-05KK TOSHIBA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2021186690
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-11-05
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Existing technologies for superimposing performer remarks as subtitles in live broadcasts result in errors and a heavy workload for subtitle producers due to real-time transcription requirements.

Method used

An information processing device that separates image and audio data, generates support information for proofreading using image data, and associates it with text information to assist subtitle creators, reducing errors and workload.

Benefits of technology

The system reduces subtitle errors and workload by providing accurate support information for proofreading, allowing quicker and more precise subtitle creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764209000001
    Figure 0007764209000001
  • Figure 0007764209000002
    Figure 0007764209000002
  • Figure 0007764209000003
    Figure 0007764209000003
Patent Text Reader

Abstract

To reduce errors in subtitles and a work load of a subtitle creator in a case where remarks of a performer or the like who appears in a program of a broadcast or the like are superimposed on a screen as subtitles.SOLUTION: An information processing device according to an embodiment comprises: a first acquisition unit; a second acquisition unit; a support information generation unit; and an output control unit. The first acquisition unit acquires image data and voice data associated with each other by time information. The second acquisition unit acquires character information generated by transcription processing using the voice data. The support information generation unit generates support information supporting calibration of the character information using image data. The output control unit outputs the character information and the support information in association with each other.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing program, and an information processing method. [Background technology]

[0002] For example, when a news program is broadcast live, there is a technology that superimposes the remarks of the performers on the screen as closed captions. However, this technology requires that the remarks of the performers during the program be transcribed in real time, which can lead to errors in the subtitles and places a heavy workload on the subtitle producer. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-68783 Summary of the Invention [Problem to be solved by the invention]

[0004] One of the problems that the present invention aims to solve is to reduce errors in subtitles and the workload of subtitle producers when the remarks of performers appearing on a program such as a broadcast are superimposed on the screen as subtitles. [Means for solving the problem]

[0005] An information processing device according to an embodiment includes a first acquisition unit, a second acquisition unit, a support information generation unit, and an output control unit. The first acquisition unit acquires image data and audio data that are associated with each other by time information. The second acquisition unit acquires text information generated by a transcription process using the audio data. The support information generation unit uses the image data to generate support information that supports proofreading of the text information. The output control unit outputs the text information and the support information in association with each other. [Brief explanation of the drawings]

[0006] [Figure 1] 1 is a diagram illustrating an example of a configuration of an information processing system according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device according to the first embodiment. [Figure 3] 2 is a diagram illustrating an example of functional blocks of a processor of the information processing device according to the first embodiment. FIG. [Figure 4] 10 is a flowchart showing an example of the flow of information processing as subtitle information generation processing executed by the first information processing system. [Figure 5] 10 is a flowchart showing an example of the flow of support information generation processing executed by the information processing device according to the first embodiment. [Figure 6] 10A and 10B are diagrams for explaining color reduction processing in the support information generation processing. [Figure 7] 10A and 10B are diagrams illustrating an example of a flattening process executed by an information processing apparatus. [Figure 8] 10A and 10B are diagrams illustrating an example of a peripheral portion removal process executed by an information processing device. [Figure 9] FIG. 10 is a diagram illustrating an example of extraction processing of a first character candidate region executed by the information processing device. [Figure 10] 10A and 10B are diagrams illustrating an example of a first background region removal process executed by the information processing device. [Figure 11] 10A and 10B are diagrams illustrating an example of a second background region removal process executed by the information processing device. [Figure 12] 10A and 10B are diagrams illustrating an example of a character candidate region determination process executed by an information processing device. [Figure 13] 10A and 10B are diagrams illustrating an example of a process for generating a reversed image executed by an information processing device. [Figure 14] 10A and 10B are diagrams showing an example of text information and support information displayed in the subtitle information generation process. [Figure 15]10 is a flowchart showing an example of the flow of support information generation processing executed by an information processing device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] The information processing device, information processing program, information processing method, and information processing system according to the embodiments will be described below. When broadcasting or distributing a news program or the like, the information processing device superimposes remarks of performers appearing in the program as subtitles onto a corresponding screen with as little time lag as possible.

[0008] In the following embodiments, for the sake of specificity, an information processing device or the like used for live broadcasting of a program wirelessly or via a cable will be described as an example. However, the information processing device or the like according to the embodiments can be used not only for broadcasting a program but also for distributing a program via a communication network, for example.

[0009] [First embodiment] Fig. 1 is a diagram showing an example of the configuration of an information processing system SY according to the first embodiment. As shown in Fig. 1, the information processing system SY includes a switcher 1, an information processing device 3, an inserter 5, a terminal device 6, and a first external device 7. The information processing device 3 is connected to the terminal device 6 and the first external device 7 via a network N so as to be able to communicate with them.

[0010] The switcher 1, information processing device 3, and inserter 5 are installed, for example, in a master room that transmits broadcasts. Here, the master room is one of the facilities installed in a broadcasting station and is also called a master control room. In the master room, image and audio data captured inside and outside the broadcasting station is edited and sent to the transmitting station as program material according to a broadcast progress schedule. Note that the switcher 1, information processing device 3, and inserter 5 can also be installed in a location other than the master room, for example, in a room equipped with equipment for creating subtitles in real time, or on the cloud.

[0011] The switcher 1 receives image and audio data from program production equipment in a sub-control room (also called a "news sub-room" or "sub-control room") and relay equipment at the broadcast site, and performs switching processing based on instructions from an operator. Here, image and audio data refers to data including image data and audio data that are associated with each other by time information. Through the switching processing of the switcher 1, image and audio data for broadcasting is generated from multiple image and audio data. The switcher 1 outputs the generated image and audio data for broadcasting to the information processing device 3.

[0012] The information processing device 3 executes subtitle information generation processing using the image and audio data acquired from the switcher 1. That is, the information processing device 3 separates the image and audio data acquired from the switcher 1 into image data and audio data. The information processing device 3 acquires character information using the separated audio data. The information processing device 3 generates support information for subtitle proofreading using the separated image data. The information processing device 3 outputs (transmits) the acquired character information and the generated support information to the terminal device 6.

[0013] The information processing device 3 also encodes the character information received from the terminal device 6 and generates subtitle information including the encoded character information and additional information. Here, the additional information includes the position (coordinates) on the screen where the character information is displayed, the character size, color, scrolling speed, time information for associating the character information with the image data, etc. The information processing device 3 outputs (transmits) the audio-visual data and the generated subtitle information to the inserter 5.

[0014] In this example, text information is obtained using separated voice data through a transcription process using voice recognition application software built into the first external device 7 on the cloud.

[0015] The subtitle information generation process and the support information generation process executed by the information processing device 3 will be described in detail later.

[0016] The inserter 5 uses the subtitle information received from the information processing device 3 to perform an insert process that inserts the subtitle information into the image and audio data received from the switcher 1. The inserter 5 outputs the image and audio data that has been subjected to the insert process to a transmission device 8 provided at the transmission station. Note that the insert process by the inserter 5 is not synchronized with the time information of the image and audio data. This assumes the presence of a delay time related to the calibration described below. Therefore, from the viewer's perspective, the subtitle information can be viewed a few seconds after the image and audio.

[0017] The terminal device 6 is a computer that is communicatively connected to the information processing device 3 and is used by a subtitler. The terminal device 6 has a dedicated application for use in proofreading installed, and outputs (displays) a proofreading screen that includes text information and support information acquired from the information processing device 3. The subtitler can proofread the text information using the proofreading screen that includes the text information and support information. The terminal device 6 transmits the proofread text information to the information processing device 3.

[0018] The first external device 7 is communicably connected to the information processing device 3 via, for example, a network N. The first external device 7 has built-in voice recognition application software. The first external device 7 performs transcription processing using the voice data received from the information processing device 3 to generate text information.

[0019] (Information processing device 3) Next, the configuration and functions of the information processing device 3 will be described in detail with reference to FIGS.

[0020] 2 is a diagram showing an example of the hardware configuration of an information processing device 3 according to the first embodiment. The information processing device 3 has a hardware configuration similar to that of a typical computer. That is, as shown in FIG. 2, the information processing device 3 includes a processor 301, a main memory device 303, a device interface 305, an auxiliary memory device 307, and a network interface 309.

[0021] The processor 301 is a processing circuit that performs overall control of the information processing device 3 and the second external device 9 connected to the information processing device 3. Specifically, the processor 301 is a CPU (Central Processing Unit), GPU (Graphics Processing Unit), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), or the like, which is an electronic circuit including a control device and an arithmetic device of a computer. The processor 301 may also be realized by an optical circuit using optical logic elements.

[0022] The main memory device 303 is a storage device that stores instructions to be executed by the processor 301, various data, etc., and information stored in the main memory device 303 is read by the processor 301. More specifically, the main memory device 303 or the auxiliary memory device 307 stores image and audio data acquired from the switcher 1, and programs for realizing the subtitle information generation process and the support information generation process, which will be described later.

[0023] The device interface 305 directly or indirectly connects the second external device 9 and the processor 301 via a bus. The device interface 305 may have a connection terminal such as a USB. An external storage medium or a storage device (memory) may also be connected to the device interface 305 via the connection terminal.

[0024] The auxiliary storage device 307 is a storage device other than the main storage device 303 .

[0025] The main storage device 303 and the auxiliary storage device 307 refer to any electronic components capable of storing electronic information, and may be semiconductor memories. Typically, the main storage device 303 and the auxiliary storage device 307 are configured with semiconductor memory elements such as RAM (Random Access Memory), flash memory, hard disks, optical disks, etc. The main storage device 303 and the auxiliary storage device 307 may also be configured with portable media such as USB (Universal Serial Bus), memory, and DVD (Digital Versatile Disk).

[0026] The network interface 309 is an interface for connecting wirelessly or by wire to the network N. The network interface 309 can transmit and receive information to and from the first external device 7 via the network N.

[0027] The second external device 9 is, for example, an input device (such as a microphone, keyboard, mouse, or touch panel), an output device (such as a display device as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel), or a storage device (memory).

[0028] 3 is a diagram showing an example of functional blocks of the processor 301 of the information processing device 3 according to the first embodiment. As shown in FIG. 3, the processor 301 includes a first acquisition function 30a, a second acquisition function 30b, an image / audio separation function 30c, a support information generation function 30d, a transmission / reception control function 30e, and a subtitle information generation function 30f.

[0029] The first acquisition function 30a, second acquisition function 30b, image / audio separation function 30c, support information generation function 30d, transmission / reception control function 30e, and subtitle information generation function 30f of processor 301 are stored in main memory device 303 or the like in the form of programs executable by a computer, for example. That is, processor 301 realizes the function corresponding to each program by reading the program from main memory device 303 or the like and executing it. In other words, when each program has been read, processor 301 has first acquisition function 30a, second acquisition function 30b, image / audio separation function 30c, support information generation function 30d, transmission / reception control function 30e, and subtitle information generation function 30f shown in processor 301 in FIG. 3.

[0030] The first acquisition function 30a acquires image and sound data for broadcasting from the switcher 1.

[0031] The second acquisition function 30b acquires character information (character data) generated by the transcription process in the first external device 7. The second acquisition function 30b is an example of a second acquisition unit.

[0032] The image / audio separation function 30c separates image data and audio data from the image / audio data acquired by the first acquisition function 30a. The separated image data and audio data are associated with each other by time information. The image / audio separation function 30c is an example of a first acquisition unit.

[0033] The support information generation function 30d generates support information that supports proofreading of text information using image data separated from the image and audio data. The support information generation process automatically generates support information that supports the proofreading process performed by a subtitle producer using image data separated from the image and audio data. The support information generation function 30d is an example of a support information generation unit.

[0034] The transmission / reception control function 30e associates the image and audio data, text information, and support information using time information and transmits them to the terminal device 6. The transmission / reception control function 30e receives text information transmitted from the terminal device 6. The transmission / reception control function 30e also outputs the image and audio data and the subtitle information generated by the subtitle information generation function 30f to the inserter 5. The transmission / reception control function 30e is an example of a transmission / reception control unit.

[0035] The subtitle information generating function 30f generates subtitle information based on character information acquired from the terminal device 6. The subtitle information generating function 30f is an example of a subtitle information generating unit.

[0036] (Subtitle information generation process) Next, the subtitle information generation process executed by the information processing system SY will be described with reference to FIGS.

[0037] Fig. 4 is a flowchart showing an example of the flow of information processing as subtitle information generation processing executed by the information processing system SY. The subtitle information generation processing shown in Fig. 4 is executed sequentially in response to acquisition of image and audio data. As shown in Fig. 4, the first acquisition function 30a of the information processing device 3 acquires image and audio data for broadcasting from the switcher 1 (step S1).

[0038] The image / audio separation function 30c of the information processing device 3 separates image data and audio data from the image / audio data acquired by the first acquisition function 30a (step S2).

[0039] The support information generating function 30d of the information processing device 3 executes support information generating processing using the image data separated from the image and sound data (step S3a).

[0040] (Support information generation process) The support information generation process executed in step S3a will be described in detail with reference to FIGS.

[0041] Fig. 5 is a flowchart showing an example of the flow of support information generation processing executed by the information processing device according to the first embodiment. Note that the support information generation processing shown in Fig. 5 can be executed on, for example, image data at a fixed time interval among time-series image data obtained by separating it from image and sound data.

[0042] The support information generating function 30d executes color reduction processing (step S11).

[0043] Fig. 6 is a diagram for explaining color reduction processing in the support information generation processing. The support information generation function 30d performs color reduction processing by reducing the number of RGB gradations (red, green, and blue gradations) of the image data shown on the left side of Fig. 6 to the number of RGB gradations shown on the left side of Fig. 6 (for example, 128 red gradations × 128 green gradations × 128 blue gradations). This color reduction processing absorbs subtle color differences between pixels and removes noise.

[0044] Next, the support information generating function 30d performs a smoothing process to remove localized steep gradient noise (step S12).

[0045] FIG. 7 is a diagram for explaining the flattening process in the support information generation process. In the flattening process, as shown on the left side of FIG. 7, the support information generation function 30d sets a mask M1 having, for example, a 5×5 grid for the image data I1. Note that the size of each grid is not limited to 5×5 pixels, but may be N×M pixels (N>1 and M>1). For example, the support information generation function 30d starts from the upper left corner of the image data I1, and scans the image data I1 along a trajectory L1 while shifting the position of this mask M1 from left to right and top to bottom.

[0046] At each position in the scan of mask M1, support information generation function 30d replaces the color of the central square in mask M1 with the most common color among the colors of the multiple squares other than the central square. For example, support information generation function 30d assumes that the colors of the squares of mask M1 at a certain position on image data I1 are distributed as shown in the center of Figure 7, and the color of the central square is "purple." In this case, support information generation function 30d converts the color "purple" of the central square to "red," the most common color among the squares other than the central square, as shown on the right side of Figure 7. This flattening process makes it possible to remove steep gradient noise from squares where the color changes locally.

[0047] Next, the support information generating function 30d executes a peripheral portion removal process (step S13).

[0048] 8 is a diagram for explaining the peripheral portion removal process in the support information generation process. As shown in Fig. 8, the support information generation function 30d removes a certain range (peripheral portion) from the top, bottom, left, and right of the image data I2 obtained by the flattening process in step S12, and generates image data I3 from the image data I2. The reason for performing such peripheral portion removal process is that it is assumed that no text information is included in the peripheral portion of the image data I2.

[0049] Next, the support information generating function 30d extracts a first character candidate region from the image data I3 (step S14).

[0050] FIG. 9 is a diagram illustrating the extraction process of a first character candidate region in the support information generation process. As shown on the left side of FIG. 9, the support information generation function 30d sets a mask M2, for example, having a 9×9 grid, for image data I3. Note that the size of each grid may be one pixel or may include multiple pixels. For example, the support information generation function 30d starts from the upper left corner of the image data I3 and scans the image data I3 along a trajectory L2 while shifting the position of the mask M2 from left to right and top to bottom. If the mask M2 is composed of two colors and the color difference (contrast) between the two colors is equal to or greater than a certain value, the support information generation function 30d generates image data I4 from the image data I3, in which the pixel region included in the mask M2 at that position is extracted as the first character candidate region P1, as shown on the right side of FIG. 9. This process is performed because the character candidate region to be extracted is composed of two colors, the character color and the background color, and the color difference between these colors can be considered large.

[0051] Next, the support information generating function 30d executes a first background region removal process using the first character candidate region P1 (step S15).

[0052] Figure 10 is a diagram illustrating the first background region removal process in the support information generation process. As shown on the left side of Figure 10, the support information generation function 30d considers the extracted first character candidate region P1 to be not a character if the same color is lined up over a long distance vertically or horizontally, and removes it from the image data I4. As a result, the support information generation function 30d generates image data I5 including the second character candidate region P2 from the image data I4, as shown on the right side of Figure 10.

[0053] Next, the support information generating function 30d executes a second background region removal process using the second character candidate region P2 (step S16).

[0054] FIG. 11 is a diagram illustrating the second background region removal process in the support information generation process. As shown on the left side of FIG. 11, the support information generation function 30d calculates the area of ​​each extracted second character candidate region P2. The support information generation function 30d considers second character candidate regions P2 whose area is greater than (or smaller than) a certain threshold value to not be characters and removes them from the image data I5. As a result, the support information generation function 30d generates image data I6 including a third character candidate region P3 from the image data I5, as shown on the right side of FIG. 11.

[0055] Next, the support information generating function 30d executes a character candidate area determination process to extract a character candidate area image (step S17).

[0056] FIG. 12 is a diagram illustrating the character candidate area determination process in the support information generation process. As shown on the left side of FIG. 12, the support information generation function 30d determines an area of ​​the third character candidate area P3 that is longer in height than in width by a certain amount (for example, an area that is twice as long vertically) and removes it from the image data I6. This is because the target characters extracted by the OCR (Optical Character Recognition) used in this embodiment are assumed to be horizontally written characters. In an actual device, the size criteria for removal can be adjusted by adjusting the settings of the OCR used. Through this character candidate area determination process, the support information generation function 30d generates image data I7 including a character candidate area image P4 from the image data I6, as shown on the right side of FIG. 12.

[0057] Next, the support information generating function 30d generates a reverse image using the character candidate area image P4 (step S18).

[0058] Fig. 13 is a diagram for explaining the inverted image generation process in the support information generation process. The support information generation function 30d generates the inverted image (negative) shown on the right side of Fig. 13 from the character candidate area image P4 (positive) shown on the left side of Fig. 13. The reason for performing such processing is that, as a rule of thumb in character recognition using OCR, there are cases where the recognition accuracy is higher when the character color and background color are inverted.

[0059] Next, the support information generation function 30d executes character determination processing by OCR using the character candidate area image P4 (positive) and the inverted image (negative) as input data (step S19). As a result, support information is generated based on the image data separated from the image and sound data.

[0060] Returning to FIG. 4, the second acquisition function 30b of the information processing device 3 acquires character information (character data) generated by the transcription process in the first external device 7 (step S3b). That is, the second acquisition function 30b transmits the audio data separated from the image and audio data to the first external device 7. The first external device 7 executes transcription process using the received audio data, generates character information, and transmits it to the information processing device 3. The second acquisition function 30b acquires the character information generated in the first external device 7.

[0061] The transmission / reception control function 30e of the information processing device 3 associates the image and sound data, the text information, and the support information with each other using time information, and transmits them to the terminal device 6 (step S4).

[0062] The terminal device 6 used by the subtitle creator outputs (displays) a proofreading screen including the text information and support information acquired from the information processing device 3.

[0063] Fig. 14 is a diagram showing an example of text information and support information displayed in the subtitle information generation process, and shows an example of a screen DP of the terminal device 6. As shown in Fig. 14, the screen DP includes a display area R1, a display area R2, a display area R3, and a display area R4.

[0064] Display area R1 displays video based on the audio-visual data. Display area R2 in step S3b displays text information obtained by the transcription process. Display area R3 displays support information obtained by the support information generation process in step S3a. Display area R4 displays time information indicating how long the display of subtitle information is currently delayed from a certain reference time.

[0065] For example, as shown in Figure 14, let us consider a case where the text information "...Triangle in terrestrial digital broadcasting experiment..." is displayed in display area R2, and the support information "Participation" is displayed in display area R3. By immediately visually recognizing the support information "Participation" obtained from the image data corresponding to the video in display area R1, the subtitler can determine that the "triangle" in the text information should be corrected to "Participation."

[0066] The terminal device 6 performs proofreading processing by character conversion or the like in response to an input from the subtitle production staff via the proofreading screen shown in Fig. 14 (step S6). The terminal device 6 transmits the proofread character data to the information processing device 3.

[0067] The subtitle information generation function 30f generates subtitle information based on the character information received from the terminal device 6 via the transmission / reception control function 30e (step S7). The subtitle information generation function 30f encodes the generated subtitle information (step S8). The transmission / reception control function 30e outputs the encoded subtitle information to the inserter 5.

[0068] The inserter 5 executes an insert process to insert the subtitle information received from the information processing device 3 into the image and audio data received from the switcher 1 (step S9). The inserter 5 outputs the image and audio data that has been subjected to the insert process to a transmission device 8 provided in a transmission station.

[0069] As described above, the information processing device 3 according to this embodiment includes an image / audio separation function 30c as a first acquisition unit, a second acquisition function 30b as a second acquisition unit, a support information generation function 30d as a support information generation unit, and a transmission / reception control function 30e as an output control unit. The image / audio separation function 30c acquires image data and audio data associated with each other by time information from the image and audio data acquired from the switcher 1. The second acquisition function 30b acquires text data generated by a transcription process using the audio data from, for example, the first external device 7. The support information generation function 30d generates support information using the image data to assist in proofreading the text data. The transmission / reception control function 30e associates the text data with the support information and outputs it. The output text information and support information are displayed on a terminal device 6 used by a subtitle creator for proofreading.

[0070] With this configuration, the subtitler can proofread the subtitles using support information generated from image data related to the text information. This reduces the number of subtitle errors and the workload of the subtitler, compared to the conventional method of proofreading using only audio information. Furthermore, the support information allows the subtitler to proofread more quickly than before.

[0071] Furthermore, the support information generation function 30d of the information processing device 3 according to this embodiment extracts text information contained in image data as an image, and generates support information by optical character recognition using the extracted image.

[0072] With this configuration, an image including character information contained in the image data can be read by OCR, and the character information contained in the image data can be effectively used as support information.

[0073] In addition, the support information generation function 30d of the information processing device 3 according to this embodiment extracts an image containing character information by using image data to perform color reduction processing, flattening processing, character candidate area extraction processing, background area removal processing, and character candidate area determination processing.

[0074] According to this configuration, character information contained in image data can be extracted as an image with high accuracy.

[0075] (Variation 1) In the above embodiment, the acquisition of text information using voice data by the information processing device 3 is exemplified by a case where the text information is acquired using separated voice data through a transcription process using voice recognition application software built into the first external device 7 on the cloud. However, for example, voice recognition application software can be built into the information processing device 3 as the second acquisition function 30b, and text information can be acquired through a transcription process using this.

[0076] [Second embodiment] Next, an information processing device 3 according to a second embodiment will be described. The information processing device according to the second embodiment generates support information by performing object recognition processing using image data separated from image and sound data.

[0077] 3, the support information generation function 30d according to the second embodiment inputs image data, executes object recognition processing, and outputs the results. Here, the object recognition processing in this embodiment includes person detection processing (for example, processing using AI face recognition technology), landscape detection processing, etc.

[0078] In the second embodiment, the configuration of the information processing system SY and the hardware configuration of the information processing device 3 are the same as those in the first embodiment, and therefore a description thereof will be omitted.

[0079] The subtitle information generation process executed by the information processing system SY according to the second embodiment is the same as the support information generation process shown in step S3a of Fig. 4. Therefore, the support information generation process according to the second embodiment will be described below.

[0080] Fig. 15 is a flowchart showing an example of the flow of support information generation processing executed by the information processing device according to the second embodiment. Note that the support information generation processing shown in Fig. 15 can be executed on, for example, image data at a fixed time interval among time-series image data obtained by separating it from image and audio data.

[0081] As shown in FIG. 15, the support information generating function 30d receives as input image data separated from the image and sound data (step S21).

[0082] Next, the support information generating function 30d executes a person detection process using the input image and sound data (step S22).

[0083] Next, the support information generating function 30d executes a scenery detection process using the input image and sound data (step S23).

[0084] Next, the support information generation function 30d outputs the results of the person detection process and the landscape detection process as support information (step S24).

[0085] The order of the person detection process in step S22 and the landscape detection process in step S23 may be reversed.Also, the configuration may be such that either the person detection process or the landscape detection process is executed.

[0086] As described above, the support information generation function 30d of the information processing device 3 according to this embodiment executes object recognition processing using image data to recognize objects included in the image data, and generates the support information based on the recognized objects. For example, the support information generation function 30d executes at least one of face recognition processing and landscape recognition processing as the object recognition processing.

[0087] This configuration allows for the generation of support information for proofreading subtitles using information about people and scenery contained in image data, thereby reducing the risk of subtitle errors and the workload of subtitle producers compared to conventional proofreading that relies solely on audio information.

[0088] (Variation 2) The support information generation process according to the second embodiment, which uses information about people and scenery contained in image data, can also be combined with the support information generation process according to the first embodiment, which uses text information contained in image data.

[0089] Although several embodiments of the present invention have been described above, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the inventions and their equivalents as defined in the claims. [Explanation of symbols]

[0090] 1...Switcher 3...Information processing device 5...Insert 6...Terminal device 7...First external device 8…Delivery device 30a...First acquisition function 30b...Second acquisition function 30c...Image and audio separation function 30d…Support information generation function 30e...Transmission and reception control function 301...Processor 303…Main storage device 305...Device Interface 307…Auxiliary storage device 309...Network Interface

Claims

1. a first acquisition unit that acquires image data and audio data that are associated with each other by time information; a second acquisition unit that acquires character data generated by a transcription process using the audio data; a support information generating unit that generates support information for supporting a user in proofreading the character data using the image data; an output control unit that outputs the character data and the support information in association with each other; An information processing device comprising:

2. The support information generation unit extracting character information contained in the image data as an image; generating the support information by optical character recognition using the extracted image; The information processing device according to claim 1 .

3. The support information generation unit extracting the image by performing color reduction processing, flattening processing, character candidate region extraction processing, background region removal processing, and character candidate region determination processing using the image data; The information processing device according to claim 2 .

4. The support information generation unit Execute an object recognition process using the image data to recognize an object included in the image data; generating the assistance information based on the recognized object; The information processing device according to claim 1 .

5. the support information generation unit executes at least one of face recognition processing and scenery recognition processing as the object recognition processing; The information processing device according to claim 4 .

6. On the computer, a first acquisition function for acquiring image data and audio data associated with each other by time information; a second acquisition function for acquiring text information generated by a transcription process using the audio data; a support information generation function that generates support information that supports a user in proofreading the character information using the image data; an output control function for outputting the character information and the support information in association with each other; An information processing program to achieve this.

7. a first acquisition step of acquiring image data and audio data associated with each other by time information; a second acquisition step of acquiring character information generated by a transcription process using the audio data; a support information generating step of generating support information that supports a user in proofreading the character information using the image data; an output control step of outputting the character information and the support information in association with each other; An information processing method comprising:

Citation Information

Patent Citations

  • Device and method for voice recognition using video information

    JP2004333738A

  • Image processing apparatus

    JP2007266793A

  • Voice recognition device, voice recognition method and program

    JP2013068783A

  • Voice confirmation system, voice confirmation method, and program

    JP2020166262A