Artificial intelligence-based subtitle management device, method, and program
The AI-based subtitle management device addresses errors and synchronization issues in conventional subtitle generation by using time and motion information for synchronization and user-driven modifications, resulting in improved subtitle accuracy and user experience.
Patent Information
- Application Number
- PCT/KR2024/004715
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-04-09
- Publication Date
- 2025-05-30
AI Technical Summary
Conventional subtitle generation technologies often produce errors and synchronization issues with video data, leading to inefficient and inaccurate subtitle management.
An artificial intelligence-based subtitle management device and method that includes a subtitle generation unit for creating synchronized subtitle data using time information, a content provision unit for resynchronizing subtitles based on motion information, and a subtitle modification unit for modifying subtitles based on user requests, while learning from modification data to improve subtitle generation accuracy.
The solution enables efficient and accurate subtitle creation, modification, and management, improving synchronization with video content and enhancing user understanding by addressing errors and synchronization issues in existing technologies.
Smart Images

Figure KR2024004715_30052025_PF_FP_ABST
Abstract
Description
Artificial intelligence-based subtitle management device, method, and program
[0001] Embodiments of the present disclosure relate to an artificial intelligence-based subtitle management device, method, and program, and more particularly, to an apparatus, method, and program that automatically generate subtitle data for content data and manage modification of the subtitle data.
[0002] Recently, content containing video data on a variety of topics has been made available to users online. Depending on the purpose of the video content, subtitles for the audio included in the video content are increasingly being provided.
[0003] In line with this trend, active development is underway in technologies to automate subtitle generation. A prime example is Speech-to-Text (STT) technology, which allows a computer to interpret human speech and convert it into text.
[0004] In the process of automatically converting voice data into text data using such technology and providing it in the form of subtitles in sync with video content, there is a need for technology to improve the efficiency and accuracy of quality management and modification of automatically generated subtitles.
[0005] However, these conventional subtitle generation technologies have problems in that automatically generated subtitles frequently contain errors or are improperly synchronized with video data.
[0006] Embodiments of the present disclosure are intended to address various issues, including the aforementioned ones, and provide an AI-based subtitle management device, method, and program. However, these tasks are exemplary and do not limit the scope of the present disclosure.
[0007] According to one aspect of the present disclosure, an artificial intelligence-based subtitle management device is provided, including a subtitle generation unit that obtains content data including video data and audio data from a first user terminal and generates first subtitle data synchronized based on time information of the audio data through a subtitle generation model, a content provision unit that resynchronizes the first subtitle data based on motion information of the video data and provides the content data and the first subtitle data to a second user terminal by matching them, and a subtitle modification unit that obtains a subtitle modification request including modification data from the second user terminal and generates second subtitle data by modifying the first subtitle data based on the modification data.
[0008] According to the present embodiment, the image data includes a plurality of sequentially consecutive frames, and the content provider classifies the plurality of frames into a plurality of groups according to the motion information, but when the Nth (where N is a positive integer) frame and the N+1th frame included in the plurality of frames include different motion information, the Nth frame and the N+1th frame can be classified into different groups.
[0009] According to the present embodiment, the plurality of groups include a first group and a second group that are sequentially consecutive, and the content provider can synchronize the starting point of the part matched to the first group to match the starting point of the first frame of the second group when the part matched to the first group in the first subtitle data corresponds to the motion information of the second group.
[0010] According to the present embodiment, the subtitle modification unit may determine the suitability of the subtitle modification request based on at least one of a matching rate between the modification data and the first subtitle data and subtitle modification requester information, and if the suitability is greater than a threshold, modify the first subtitle data to generate the second subtitle data.
[0011] According to the present embodiment, the content provider may, based on the subtitle modification request, synchronize the starting point of the second subtitle data to match the starting point of the first subtitle data when the starting point of the first subtitle data and the starting point of the second subtitle data are different for the modified portion.
[0012] According to the present embodiment, when the second subtitle data is generated, the subtitle modification unit can extract a modification keyword from the modification data and update the second subtitle data by identically modifying a portion of the second subtitle data that includes the modification keyword.
[0013] According to the present embodiment, the subtitle modification unit can search for similar content data related to the modification keyword among other content data acquired by the subtitle generation unit and used to generate subtitle data, and can identically modify a portion of the subtitle data generated for the similar content data that includes the modification keyword.
[0014] According to the present embodiment, the subtitle modification unit provides the modification data to the subtitle generation unit, the subtitle generation unit trains the subtitle generation model by using the modification data as learning data, and the learned subtitle generation model can generate subtitle data by reflecting the modification data when generating subtitle data for new content data.
[0015] According to one aspect of the present disclosure, an artificial intelligence-based subtitle management method is provided, including the steps of obtaining content data including video data and audio data from a first user terminal, generating first subtitle data synchronized based on time information of the audio data through a subtitle generation model, resynchronizing the first subtitle data based on motion information of the video data, and providing the content data and the first subtitle data to a second user terminal by matching them, and obtaining a subtitle modification request including modification data from the second user terminal, and generating second subtitle data by modifying the first subtitle data based on the modification data.
[0016] According to one aspect of the present disclosure, a computer-readable recording medium is provided, which stores a program for executing the artificial intelligence-based open market management method by being coupled to a computer.
[0017] Other aspects, features and advantages other than those described above will become apparent from the following detailed description, claims and drawings for carrying out the invention.
[0018] Additionally, these general and specific aspects may be implemented using any system, method, computer program, or combination of any system, method, or computer program.
[0019] According to the exemplary embodiments of the present disclosure described above, an AI-based subtitle management device, method, and program can be implemented that efficiently performs everything from subtitle creation to editing and management for content data without the involvement of the content creator. Of course, the scope of the present disclosure is not limited by these effects.
[0020] FIG. 1 is a conceptual diagram schematically illustrating a content provision system according to an exemplary embodiment of the present disclosure.
[0021] FIG. 2 is a conceptual diagram schematically illustrating the operation of a subtitle management device of a content provision system according to an exemplary embodiment of the present disclosure.
[0022] FIG. 3 is an exemplary diagram schematically illustrating a subtitle data synchronization function according to an exemplary embodiment of the present disclosure.
[0023] FIG. 4 is an exemplary diagram schematically illustrating a subtitle data synchronization function according to an exemplary embodiment of the present disclosure.
[0024] FIG. 5 is a conceptual diagram schematically illustrating a subtitle data modification function according to an exemplary embodiment of the present disclosure.
[0025] FIG. 6 is a conceptual diagram schematically illustrating a subtitle data modification function according to an exemplary embodiment of the present disclosure.
[0026] FIG. 7 is a flowchart schematically illustrating a subtitle management method according to an exemplary embodiment of the present disclosure.
[0027] FIGS. 8 to 11 are exemplary diagrams schematically illustrating screens provided by a subtitle management device according to an exemplary embodiment of the present disclosure.
[0028] The present disclosure is capable of various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present disclosure, as well as methods for achieving them, will become clearer with reference to the embodiments described in detail below, along with the drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various forms.
[0029] In the examples below, the terms first, second, etc. are not used in a limiting sense, but are used for the purpose of distinguishing one component from another.
[0030] In the examples below, singular expressions include plural expressions unless the context clearly indicates otherwise.
[0031] In the following examples, terms such as “include” or “have” mean that a feature or component described in the specification is present, and do not preclude the possibility that one or more other features or components may be added.
[0032] In the following examples, when a part such as a layer, region, or component is said to be on or above another part, it includes not only the case where it is directly above the other part, but also the case where another region, component, or the like is interposed in between.
[0033] For convenience of explanation, the sizes of components in the drawings may be exaggerated or reduced. For example, the sizes and thicknesses of each component shown in the drawings are arbitrarily indicated for convenience of explanation, and thus the present disclosure is not necessarily limited to the figures shown.
[0034] In some embodiments, where implementations are otherwise feasible, specific sequences of operations may be performed in a different order than described. For example, two steps described in succession may be performed substantially simultaneously, or in a reverse order from the described order.
[0035] In this specification, “A and / or B” refers to the case where it is A, or B, or both A and B. And, “at least one of A and B” refers to the case where it is A, or B, or both A and B.
[0036] In the following examples, when it is said that layers, regions, components, etc. are connected, it includes cases where the layers, regions, components, etc. are directly connected, and / or cases where other layers, regions, components, etc. are interposed between the layers, regions, and components and are indirectly connected. For example, when it is said in this specification that layers, regions, components, etc. are electrically connected, it refers to cases where the layers, regions, components, etc. are directly electrically connected, and / or cases where other layers, regions, components, etc. are interposed between them and are indirectly electrically connected.
[0037] The advantages and features of the present disclosure, and methods for achieving them, will become clearer with reference to the embodiments described below in detail with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure is complete and to fully inform those skilled in the art of the scope of the present disclosure, and the present disclosure is defined solely by the scope of the claims.
[0038] The terminology used in this disclosure is for the purpose of describing embodiments only and is not intended to limit the present disclosure. In this disclosure, the singular may also include the plural unless specifically stated otherwise. The terms "comprises" and / or "comprising" as used herein do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the disclosure, and "and / or" may include each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present disclosure.
[0039] The word "exemplary" is used herein to mean "serving as an example or illustration." Any embodiment described herein as "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments.
[0040] Embodiments of the present disclosure may be described in terms of a function or a block that performs a function. A block, which may be referred to as a "unit" or a "module" in the present disclosure, may be physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memories, passive electronic components, active electronic components, optical components, hardwired circuits, etc., and may optionally be driven by firmware and software. Furthermore, the term "unit" as used in the disclosure refers to software, hardware elements such as FPGAs or ASICs, and the "unit" may perform certain roles. However, the "unit" is not limited to software or hardware. The "unit" may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, a "part" may include elements such as software elements, object-oriented software elements, class elements, and task elements, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the elements and "parts" may be combined into a smaller number of elements and "parts" or further separated into additional elements and "parts."
[0041] Embodiments of the present disclosure can be implemented using at least one software program running on at least one hardware device and capable of performing network management functions to control elements.
[0042] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" may be used to readily describe the relationship between one component and other components as depicted in the drawings. Spatially relative terms may be understood to encompass different orientations of components during use or operation in addition to the orientations depicted in the drawings. For example, if a component depicted in the drawings were flipped over, a component described as "below" or "beneath" another component may end up "above" the other component. Thus, the exemplary term "below" may encompass both the above and below orientations. Components may also be oriented in other directions, and thus spatially relative terms may be interpreted accordingly.
[0043] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure may be used with the meaning commonly understood by those skilled in the art to which this disclosure pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals and redundant descriptions thereof will be omitted.
[0045] FIG. 1 is a conceptual diagram schematically illustrating a content provision system according to an exemplary embodiment of the present disclosure.
[0046] As illustrated in FIG. 1, a content providing system (1) according to an exemplary embodiment of the present disclosure may include a subtitle management device (10), a first user terminal (20), and a second user terminal (30).
[0047] A subtitle management device (10) is a device that processes and provides content data including video data and / or audio data. There is no limitation on the form of the subtitle management device (10), and it may include any of a variety of devices capable of performing computational processing and providing results to a user. For example, the subtitle management device (10) according to the present disclosure may take the form of one or a combination of two or more of a computer, a server device, and a portable terminal.
[0048] The subtitle management device (10) can communicate with the first user terminal (20) and / or the second user terminal (30) and transmit and receive data.
[0049] The subtitle management device (10) can obtain content data from the first user terminal (20). The subtitle management device (10) can process the content data obtained from the first user terminal (20). For example, the subtitle management device (10) can generate subtitle data corresponding to the content data obtained from the first user terminal (20). The subtitle management device (10) can match the content data obtained from the first user terminal (30) and the subtitle data generated by the subtitle management device (10) and provide the matched content data to the second user terminal (30).
[0050] Here, content data may include video data and audio data. Video data is data including a video signal that conveys visual information. Audio data is data including an audio signal that conveys voice-based auditory information. Content data may further include data including an audio signal that conveys non-voice-based auditory information (e.g., background noise, sound effects, etc.).
[0051] Meanwhile, the format of the content data may be, but is not limited to, any one of MP4, MOV, WMV, AVI, AVCHD, FLV, F4V, SWF, MKV, WEBM, and HTML5. In addition, the format of the subtitle data may be, but is not limited to, any one of SRT, SBV, SUB, MPSUB, LRC, CAP, SMI, SAMI, RT, VTT, TTML, and DFXP.
[0052] The subtitle management device (10) can perform modification and management tasks on generated subtitle data. In one embodiment, the subtitle management device (10) can obtain a modification request for subtitle data from a second user terminal (30) and modify the subtitle data based on modification data included in the obtained modification request. In other words, the subtitle management device (10) can not only arbitrarily modify subtitle data, but also perform subtitle modification tasks through communication with a user of the content provision system (1).
[0053] The first user terminal (20) is a terminal of the first user that provides content data to the subtitle management device (10). In other words, the first user terminal (20) is a terminal of a user that produces a video for supplying to other users.
[0054] The second user terminal (30) is a terminal of a second user that receives content data from the subtitle management device (10). For example, the second user terminal (30) may be a terminal of a user that views content managed by the subtitle management device (10).
[0055] The second user terminal (30) may provide a request for modification of subtitle data for the provided content data to the subtitle management device (10). The subtitle management device (10) may modify the subtitle data based on modification data included in the subtitle modification request obtained from the second user terminal (30). In addition, the subtitle management device (10) may determine the suitability of the subtitle modification request obtained from the second user terminal (30) and perform subtitle modification work only if it is determined to be suitable. A detailed description thereof will be provided later.
[0056] Meanwhile, the first user terminal (20) and the second user terminal (30) are devices capable of wireless communication, and their forms are not limited. For example, the first user terminal (20) and the second user terminal (30) according to the present disclosure may be portable terminals such as computers, smartphones, etc.
[0057] FIG. 2 is a conceptual diagram schematically illustrating the operation of a subtitle management device of a content provision system according to an exemplary embodiment of the present disclosure.
[0058] As illustrated in FIG. 2, the subtitle management device (10) may include a subtitle generation unit (100), a content provision unit (200), and a subtitle modification unit (300).
[0059] The subtitle generation unit (100) acquires content data including video data and audio data from a second user terminal (20), and automatically generates first subtitle data synchronized based on the time information of the audio data through a subtitle generation model for the acquired content data. To this end, the subtitle generation unit (100) may include a subtitle generation model. Here, the subtitle generation model may include an STT (Speech-to-Text) model, and there is no limitation on the type of STT API (Application Programming Interface).
[0060] Specifically, the subtitle generation unit (100) can selectively extract audio data from the video data and audio data included in the content data acquired from the second user terminal (20). The subtitle generation unit (100) can selectively extract voice data from the voice data and non-voice data included in the extracted audio data.
[0061] The subtitle generation unit (100) can generate subtitle data by converting the extracted voice data into text data. At this time, the subtitle data generated by the subtitle generation unit (100) may be subtitle data synchronized based on the time information of the voice data. That is, the subtitle generation unit (100) recognizes voice data according to the flow of time, and generates text data that has undergone natural language processing by matching it to the time information at which the audio signal of the voice data occurred, thereby generating subtitle data synchronized based on the time information of the voice data.
[0062] Meanwhile, the subtitle generation unit (100) can use the correction data of the subtitle correction unit (300) described below as learning data to reinforce learning the subtitle generation model. That is, when the subtitle correction unit (300) performs a subtitle correction task, it provides the correction data to the subtitle generation unit (100), and the subtitle generation unit (100) trains the subtitle generation model based on the acquired correction data. When generating subtitle data for new content data, the subtitle generation model that has learned the correction data can reflect the correction data to generate subtitle data. In this way, the subtitle generation unit (100) can perform a more accurate subtitle generation task by continuously accumulating and learning the correction data.
[0063] The content provision unit (200) matches content data obtained from a first user terminal (20) with subtitle data generated by a subtitle generation unit (100) for the content data and provides the same to a second user terminal (30).
[0064] Meanwhile, subtitle data synchronized based on the time information of the audio data generated by the subtitle generation unit (100) may have a part that is inconsistent with the motion information of the video data (e.g., the screen provided) during the process of providing or viewing the content.
[0065] In one embodiment, the content data may be lecture content data. The lecture content data may include video data containing lecture materials and audio data containing explanations of the lecture content. In this case, the video data containing the lecture materials may include operation information related to the lecture materials (e.g., displaying a specific page in the lecture materials, playing a video inserted within a specific page, executing a special effect inserted within a specific page, turning to the next page, etc.). Furthermore, the audio data containing the lecture content explanation may include an instructor's voice audio signal generated based on time information.
[0066] The subtitle generation unit (100) can generate synchronized subtitle data based on the time information of audio data including the lecture content description. Since the subtitle data generated by the subtitle generation unit (100) is synchronized with the time information of the audio data, in cases where the video data including the lecture material and the audio data including the lecture content description do not match, the subtitle data also becomes inconsistent with the video data. For example, if the instructor of the lecture content data explains the content of the next page in advance before turning the page of the lecture material, the motion information of the video data is the lecture material for the current page, and the audio data is the lecture content description for the next page, so the video data and the audio data become inconsistent. Similarly, the subtitle data also becomes inconsistent with the video data.
[0067] According to embodiments of the present disclosure, by resynchronizing at least a portion of subtitle data based on motion information rather than time information, subtitle data can be provided in a more accurate manner matching the content data. Accordingly, viewers receive more accurate subtitle data corresponding to video data, thereby enhancing their understanding of the content data.
[0068] In one embodiment, the content provider (200) may re-synchronize the first subtitle data generated by the subtitle generator (100) based on the time information of the audio data of the content data, based on the motion information of the video data of the content data, and provide the content data and the first subtitle data to the second user terminal (20) by matching them. A detailed description of such re-synchronization based on motion information will be described later with reference to FIGS. 3 and 4.
[0069] The subtitle correction unit (300) performs the role of modifying the first subtitle data generated by the subtitle generation unit (100) to generate second subtitle data.
[0070] In one embodiment, the subtitle correction unit (300) may obtain a subtitle correction request including correction data from a second user terminal (30), and correct the first subtitle data based on the correction data to generate second subtitle data.
[0071] The above correction data may include a correction request portion and a correction draft among the first subtitle data. The subtitle correction unit (300) may manage the subtitle data in corpus units divided by a preset standard to distinguish the correction request portion. For example, the subtitle correction unit (300) may manage the subtitle data in corpus units divided by at least one of a sentence unit, a phrase unit, a word unit, a character unit, a morpheme unit, and a sentence component unit (e.g., a subject, a predicate, a complement, an object, an adverb, an adjective, an independent word). Accordingly, the correction data of the subtitle correction request may be identified for which portion of the subtitle data the correction request is made, and correction work may be performed only on the corresponding portion.
[0072] The subtitle correction unit (300) can determine the suitability of a subtitle correction request to ensure the reliability of subtitle correction, and perform subtitle correction work only when it is determined to be suitable, and can process the correction request as unsuitable when it is determined to be unsuitable.
[0073] In one embodiment, the subtitle correction unit (300) may determine the suitability of the subtitle correction request based on at least one of the matching rate between the correction data and the first subtitle data and the subtitle correction requester information, and if the suitability is greater than a threshold, may correct the first subtitle data to generate second subtitle data.
[0074] The subtitle correction unit (300) can extract a corpus corresponding to a correction request portion of the correction data within the first subtitle data, and analyze the matching rate between the extracted corpus and the correction proposal of the correction data. At this time, the subtitle correction unit (300) can adjust the range of the corpus unit to be wide or narrow depending on the range of the correction request portion. The subtitle correction unit (300) can calculate a higher suitability as the matching rate between the original corpus extracted from the subtitle data and the correction proposal for the corresponding corpus is higher. The subtitle correction unit (300) can determine that the subtitle correction request is suitable when the matching rate between the original corpus and the correction proposal for the corresponding corpus is higher than a threshold value (or when the calculated suitability is higher than a threshold value).
[0075] In addition, the subtitle correction unit (300) can analyze the subtitle correction requester information that provided the subtitle correction request. The subtitle correction unit (300) can analyze the subtitle correction requester information based on the type of the subtitle correction requester (content provider or content viewer), subtitle correction history, similar content viewing history, or supply history, etc. to determine the suitability of the subtitle correction request. For example, the subtitle correction unit (300) can calculate a higher suitability if the type of the subtitle correction requester is a content provider than if the type of the subtitle correction requester is a content viewer, and can calculate a higher suitability as the subtitle correction history, similar content viewing history, or supply history increases. If the suitability calculated by analyzing the subtitle correction requester information is greater than a threshold, the subtitle correction unit (300) can modify the first subtitle data to generate second subtitle data.
[0076] Meanwhile, the subtitle correction unit (300) can also calculate the suitability of a correction request by utilizing the matching rate between the correction data and the first subtitle data and the subtitle correction requester information. The subtitle correction unit (300) can set different weights for the matching rate between the correction data and the first subtitle data and the subtitle correction requester information.
[0077] As a specific example, the subtitle correction unit (300) may assign a higher weight to the content of the subtitle correction request than to the subject of the subtitle correction request. In other words, the subtitle correction unit (300) may set a higher weight for the matching rate between the correction data and the first subtitle data than for the weight for the analysis of the subtitle correction requester information.
[0078] As another specific example, the subtitle correction unit (300) may assign a lower weight to the content of the subtitle correction request than to the subject of the subtitle correction request. That is, the subtitle correction unit (300) may set a lower weight for the matching rate between the correction data and the first subtitle data than for the weight for the analysis of the subtitle correction requester information.
[0079] In one embodiment, the subtitle correction unit (300) may provide compensation to the subtitle correction requester who provided the subtitle correction request when the subtitle correction work is performed based on the subtitle correction request.
[0080] FIG. 3 is an exemplary diagram schematically illustrating a subtitle data synchronization function according to an exemplary embodiment of the present disclosure.
[0081] As illustrated in FIG. 3, content provided to a second user terminal (30, see FIG. 2) may include video data (11), audio data (12), and subtitle data (13).
[0082] The video data (11) may include a plurality of sequentially consecutive frames. Here, the number of frames per unit time may vary depending on the frame rate. For example, the frame rate may be 24 fps, 30 fps, 60 fps, etc., but is not limited thereto.
[0083] The content provider (300, see FIG. 2) can classify multiple frames included in the video data (11) into multiple groups based on motion information. The content provider (200) can analyze the motion information included in each of the multiple frames, classify frames with the same motion information into the same group, and classify frames with different motion information into different groups.
[0084] In one embodiment, the content provider (200) may classify the Nth frame (where N is a positive integer) and the N+1th frame, which are included in a plurality of frames included in the video data (11), into different groups when they include different motion information.
[0085] Likewise, the content provider (200) can classify the Nth frame (where N is a positive integer) and the N+1th frame, which are included in the plurality of frames included in the video data (11), into the same group if they include the same motion information.
[0086] As illustrated in FIG. 3, the plurality of groups classified by the content provider (200) may include a first group (11a) and a second group (11b) that are sequentially consecutive. That is, the frames included in the first group (11a) are frames that include the same motion information, and the frames included in the second group (11b) are frames that include the same motion information, but include motion information that is different from the frames included in the first group (11a).
[0087] The content provider (200) can sequentially perform a comparison of motion information between the 1st and 2nd frames of the video data (11), and a comparison of motion information between the last frame and the frame immediately before the last frame. The content provider (200) sequentially performs the motion information comparison between the two consecutive frames as described above, and can classify groups when the Nth frame and the N+1th frame having different motion information are found. In this case, the 1st frame to the Nth frame can be classified into a first group (11a), and the N+1th frame and the M+1th frame can be classified into a second group (11b). Similarly, when the Mth frame (wherein M is a positive integer greater than N+1) and the M+1th frame having different motion information are found, the N+1th frame to the Mth frame can be classified into a second group (11b), and the M+1th frame and the M+1th frame can be classified into a third group.
[0088] Meanwhile, in one embodiment, the content provider (200) may determine whether motion information between frames is the same based on at least one of the code, category, topic, content introduction, provider information, viewer information, video progress, visual information (e.g., lecture material images and text) analysis, and subtitle data content assigned to the content.
[0089] The first subtitle data (13) may include first subtitle data (13a) synchronized based on the time information of the audio data (12) and first subtitle data (13b) synchronized based on the motion information of the video data (11). The content provider (200) may resynchronize the first subtitle data (13a) synchronized based on the time information of the audio data (12) based on the motion information of the video data (11).
[0090] In one embodiment, the content provider (200) may synchronize the starting point (t1) of the part matched to the first group (11a) of the video data (11) in the first subtitle data (13a) synchronized based on the time information of the audio data (12), if the part corresponds to the motion information of the second group (11b) of the video data (11).
[0091] For example, referring to FIG. 3, the "CCCCCCCCCCC" portion of the first subtitle data (13) is matched to the first group (11a), but since the "CCCCCCCCCCC" portion corresponds to the motion information of the second group (11b), it can be confirmed that the starting point of the "CCCCCCCCCCC" portion is synchronized to match the starting point (t2) of the first frame of the second group (11b).
[0092] FIG. 4 is an exemplary diagram schematically illustrating a subtitle data synchronization function according to an exemplary embodiment of the present disclosure.
[0093] As illustrated in FIG. 4, the subtitle correction unit (300, see FIG. 2) can modify the first subtitle data (13) generated by the subtitle generation unit (100, see FIG. 2) to generate second subtitle data (14).
[0094] The second subtitle data (14) may include second subtitle data (14a) before synchronization due to subtitle modification and second subtitle data (14b) after synchronization due to subtitle modification.
[0095] In one embodiment, the content provider (200, see FIG. 2) may, based on a subtitle modification request, synchronize the start point (t4) of the second subtitle data (14) to match the start point (t3) of the first subtitle data (13) with the start point (t3) of the first subtitle data (14) when the start point (t3) of the first subtitle data (14) is different from the start point (t4) of the second subtitle data (14) with respect to the modified portion. That is, the content provider (200) may synchronize the start point (t4) of the portion modified through the subtitle modification operation in the second subtitle data (14a) prior to synchronization due to subtitle modification with the start point (t3) of the first subtitle data (13) prior to modification of the portion, thereby generating the second subtitle data (14b) after synchronization due to subtitle modification.
[0096] For example, referring to FIG. 4, it can be confirmed that the starting point has changed as the "CCCCCCCCCCCCCCC" part of the first subtitle data (13) has been modified to "XXXXXX", and the starting point of the "XXXXXX" part has been synchronized to match the starting point (t3) of the "CCCCCCCCCCCCCCCC" part of the existing first subtitle data (13).
[0097] FIG. 5 is a conceptual diagram schematically illustrating a subtitle data modification function according to an exemplary embodiment of the present disclosure.
[0098] As illustrated in FIG. 5, the subtitle correction unit (300, see FIG. 2) can generate second subtitle data (14) by correcting at least a portion of the first subtitle data (13). The subtitle correction unit (300) can update the second subtitle data (14) by performing additional subtitle correction work based on the correction data. Accordingly, the second subtitle data (14) can include second subtitle data (14a) before updating and second subtitle data (14c) after updating.
[0099] In one embodiment, when the subtitle correction unit (300) generates second subtitle data (14), it can extract a correction keyword from the correction data and update the second subtitle data (14) by identically modifying the portion of the second subtitle data (14) that includes the correction keyword. That is, the subtitle correction unit (300) can extract a correction keyword that is the core of the content of the modified first subtitle data (13), and additionally search for a portion of the second subtitle data (14a) before updating that includes the extracted correction keyword and perform the same correction operation, thereby generating the second subtitle data (14c) after updating.
[0100] For example, referring to FIG. 5, the first subtitle data (13) includes the "BBB" part in two places, the front and the back. A subtitle modification request for the "BBB" part in the front of the first subtitle data (13) to be changed to "XXX" can be obtained from the second user terminal (30, refer to FIG. 2). Accordingly, the subtitle modification unit (300) can modify the "BBB" part in the front to "XXX" to generate the second subtitle data (14a) before updating. Subsequently, the subtitle modification unit (300) can set "BBB" as a modification keyword for the subtitle modification task, and additionally search for the part containing "BBB" in the second subtitle data (14a) before updating to extract the "BBB" part in the back. The subtitle modification unit (300) can modify the extracted "BBB" part in the back to "XXX", which is the same as the "BBB" part in the front, to generate the second subtitle data (14c) after updating.
[0101] This additional subtitle editing work allows the subtitle requester to perform the editing work on the entire subtitle data, enabling more efficient subtitle editing and improving the overall subtitle quality.
[0102] Meanwhile, if the above-described additional subtitle modification work is performed, the subtitle modification unit (300) can increase the compensation to the subtitle modification requester in proportion to the number of parts for which the additional subtitle modification work is performed.
[0103] FIG. 6 is a conceptual diagram schematically illustrating a subtitle data modification function according to an exemplary embodiment of the present disclosure.
[0104] As illustrated in FIG. 6, a content provision system (1, see FIG. 2) can obtain multiple content data from a first user terminal (20, see FIG. 2) and provide multiple contents to a second user terminal (30, see FIG. 2).
[0105] The subtitle correction unit (300, see FIG. 2) can perform additional subtitle correction work on other content data that is likely to have similar subtitle errors based on correction data of a subtitle correction work performed on a single content data.
[0106] In one embodiment, the subtitle correction unit (300) may perform a subtitle correction operation on a certain content data and extract a correction keyword from the correction data of the subtitle correction operation. The subtitle correction unit (300) may search for similar content data related to the correction keyword among other content data acquired by the subtitle generation unit (100, see FIG. 2) and from which subtitle data was generated, and may similarly correct a portion of the subtitle data generated for the similar content data that includes the correction keyword.
[0107] For example, referring to FIG. 6, the subtitle correction unit (300) can search for one or more other content data having a high correlation with the correction keyword. The subtitle correction unit (300) can analyze the subtitle data generated for one or more of the searched content data to determine whether a subtitle error related to the correction keyword exists. If a subtitle error related to the correction keyword is found, the subtitle correction unit (300) can perform the same subtitle correction task based on the correction data.
[0108] Through this kind of serial additional subtitle modification work, there is an effect of improving the overall subtitle quality of multiple content data registered in the content provision system (1).
[0109] Meanwhile, the subtitle correction unit (300) can increase the compensation to the subtitle correction requester in proportion to the number of additional subtitle correction tasks performed when additional subtitle correction tasks are performed.
[0110] FIG. 7 is a flowchart schematically illustrating a subtitle management method according to an exemplary embodiment of the present disclosure.
[0111] As illustrated in FIG. 7, a subtitle management method according to an exemplary embodiment of the present disclosure may include a step of obtaining content data (S100), a step of generating first subtitle data (S200), a step of resynchronizing the first subtitle data (S300), a step of providing content data and first subtitle data (S400), a step of obtaining a subtitle modification request (S500), a step of determining suitability of the modification request (S600), a step of processing the modification request as unsuitable (S710), a step of generating second subtitle data (S720), a step of performing an additional modification task (S800), and a step of learning the modification data (S900).
[0112] Hereinafter, the same drawing symbols in the drawings represent the same components, and explanations of content that overlaps with the above content are omitted.
[0113] The method is a step of obtaining content data (S100) in which a subtitle generation unit (100, see FIG. 2) obtains content data including video data and audio data from a first user terminal (20, see FIG. 2).
[0114] The step of generating first subtitle data (S200) is a step in which the subtitle generation unit (100) generates first subtitle data synchronized based on the time information of the voice data through a subtitle generation model.
[0115] The step (S300) of resynchronizing the first subtitle data is a step in which the content provider (200, see FIG. 2) resynchronizes the first subtitle data based on the motion information of the video data.
[0116] The step (S400) of providing content data and first subtitle data is a step in which the content provider (200) matches the content data and the first subtitle data and provides them to a second user terminal (30, see FIG. 2).
[0117] The step of obtaining a subtitle modification request (S500) is a step in which the subtitle modification unit (300, see FIG. 2) obtains a subtitle modification request including modification data from a second user terminal (30).
[0118] The step (S600) of determining the suitability of the modification request is a step in which the subtitle modification unit (300) determines the suitability of the subtitle modification request based on at least one of the matching rate between the modification data and the first subtitle data and the subtitle modification requester information.
[0119] The step of processing a modification request as unsuitable (S710) is a step of processing a subtitle modification request as unsuitable if the subtitle modification unit (300) determines that the subtitle modification request is unsuitable (if the suitability of the subtitle modification request is below the threshold).
[0120] The step of generating second subtitle data (S720) is a step of generating second subtitle data by modifying the first subtitle data based on the modification data when the subtitle modification unit (300) determines that the subtitle modification request is appropriate (when the suitability of the subtitle modification request is greater than or equal to a threshold value).
[0121] The step (S800) of performing additional correction work is a step in which the subtitle correction unit (300) performs additional subtitle correction work based on correction data.
[0122] In one embodiment, the step of performing additional modification work (S800) may include a step of updating the second subtitle data (14) by performing additional subtitle modification work based on the modification data by the subtitle modification unit (300).
[0123] Specifically, the step (S800) of performing additional modification work may include a step of extracting a modification keyword from the modification data when the subtitle modification unit (300) generates the second subtitle data (14), and updating the second subtitle data (14) by identically modifying a portion of the second subtitle data (14) that includes the modification keyword.
[0124] In one embodiment, the step of performing an additional correction operation (S800) may include a step of performing an additional subtitle correction operation on other content data that is likely to have similar subtitle errors based on correction data of a subtitle correction operation performed by the subtitle correction unit (300) on one content data.
[0125] Specifically, the step (S800) of performing an additional modification task may include a step in which the subtitle modification unit (300) performs a subtitle modification task on a certain content data, extracts a modification keyword from the modification data of the subtitle modification task, searches for similar content data related to the modification keyword among other content data acquired by the subtitle generation unit (100) and from which subtitle data was generated, and identically modifies a portion of the subtitle data generated for the similar content data that includes the modification keyword.
[0126] The step (S900) of learning correction data is a step in which the subtitle correction unit (300) provides the correction data to the subtitle generation unit (100), and the subtitle generation unit (100) uses the correction data as learning data to learn the subtitle generation model. When generating subtitle data for new content data, the subtitle generation model that has learned the correction data can generate subtitle data by reflecting the correction data.
[0127] FIGS. 8 to 11 are exemplary diagrams schematically illustrating screens provided by a subtitle management device according to an exemplary embodiment of the present disclosure.
[0128] As illustrated in FIGS. 8 to 10, a subtitle management device (10, see FIG. 2) according to an exemplary embodiment of the present disclosure may include at least one of a subtitle function activation button, a script activation button, and a subtitle (or script) modification request button on a screen providing content.
[0129] The subtitle function activation button can set whether to display subtitle data corresponding to the video data included in the provided content. For example, the subtitle management device (10) can provide a menu that allows users to turn subtitles on / off or set the display position of subtitles when the subtitle function activation button is clicked. This subtitle function activation button may be displayed at the bottom right of the screen, but is not limited thereto, and may be positioned at any location on the content provision screen.
[0130] In one embodiment, the subtitle management device (10) may provide subtitle data by overlapping at least a portion of the video data. In another embodiment, the subtitle management device (10) may place the subtitle data on one side of the video data (e.g., the lower side, upper side, left side, right side, etc.) so that the subtitle data does not overlap with the video data. The display position of such subtitle data can be controlled through a subtitle function activation button.
[0131] The script activation button can set whether to display script data corresponding to video data included in the provided content. Here, the script data is data including all subtitle data provided according to the flow of the video data playback time. Referring to FIG. 11, the script data can classify subtitle data according to preset criteria, and display the classified subtitle data by matching the time information of audio data. In addition, the subtitle data can be displayed sequentially on the script data according to the time information. In one embodiment, when a part of the subtitle data included in the script data is selected, video data corresponding to the time information matched to the corresponding subtitle data can be displayed. Through this, content viewers can conveniently search for video data corresponding to specific subtitle data.
[0132] In one embodiment, a user viewing content can activate script data and immediately request modification of at least a portion of the subtitle data included in the script data. For example, as illustrated in FIG. 11, the content viewer can select at least a portion of the subtitle data included in the subtitle script data and input a modification proposal for the selected subtitle data. The subtitle modification unit (300, see FIG. 2) can obtain the modification data entered by the user, determine the suitability of the aforementioned subtitle modification request, and perform the subtitle modification task.
[0133] The subtitle management device (10) may provide a menu that allows turning the script on / off or setting the script display position when a script activation button is clicked. This script function activation button may be displayed on the right side of the screen, but is not limited thereto, and may be positioned at any location on the content provision screen.
[0134] While this disclosure has been described with reference to the embodiments illustrated in the drawings, these are merely exemplary, and those skilled in the art will appreciate that various modifications and equivalent alternative embodiments are possible. Therefore, the true scope of technical protection of this disclosure should be determined by the technical spirit of the appended claims.
Claims
1. A subtitle generation unit that obtains content data including video data and audio data from a first user terminal and generates first subtitle data synchronized based on time information of the audio data through a subtitle generation model; A content providing unit that resynchronizes the first subtitle data based on the motion information of the video data and provides the content data and the first subtitle data to a second user terminal by matching them; and A subtitle modification unit is included, which obtains a subtitle modification request including modification data from the second user terminal, and modifies the first subtitle data based on the modification data to generate second subtitle data. The above image data includes a plurality of sequentially consecutive frames, The above content provider is, Classify the above multiple frames into multiple groups according to the above motion information, If the Nth frame (where N is a positive integer) and the N+1th frame included in the above plurality of frames contain different motion information, the Nth frame and the N+1th frame are classified into different groups, The above subtitle correction section, An artificial intelligence-based subtitle management device that determines the suitability of the subtitle modification request based on at least one of the matching rate between the modification data and the first subtitle data and the subtitle modification requester information, and modifies the first subtitle data to generate the second subtitle data if the suitability is greater than a threshold.
2. In paragraph 1, The above multiple groups include a first group and a second group which are sequentially consecutive, The above content provider is, An artificial intelligence-based subtitle management device that synchronizes the starting point of the part matched to the first group to match the starting point of the first frame of the second group when the part matched to the first group in the first subtitle data corresponds to the motion information of the second group.
3. In paragraph 1, The above content provider is, An artificial intelligence-based subtitle management device that synchronizes the starting point of the second subtitle data to match the starting point of the first subtitle data when the starting point of the first subtitle data and the starting point of the second subtitle data are different for the modified portion based on the subtitle modification request.
4. In paragraph 3, The above subtitle correction section, An artificial intelligence-based subtitle management device that, when generating the second subtitle data, extracts a modification keyword from the modification data and updates the second subtitle data by identically modifying a portion of the second subtitle data containing the modification keyword.
5. In paragraph 4, The above subtitle correction section, An artificial intelligence-based subtitle management device, which searches for similar content data related to the above-mentioned modification keyword among other content data acquired by the above-mentioned subtitle generation unit to generate subtitle data, and identically modifies a portion of the subtitle data generated for the similar content data that includes the above-mentioned modification keyword.
6. In paragraph 5, The above subtitle modification unit provides the above modification data to the above subtitle generation unit, The above subtitle generation unit trains the subtitle generation model by using the above modification data as learning data, The above-mentioned learned subtitle generation model is an artificial intelligence-based subtitle management device that generates subtitle data by reflecting the above-mentioned modification data when generating subtitle data for new content data.
7. In paragraph 1, The above subtitle correction section, An artificial intelligence-based subtitle management device that manages the subtitle data in corpus units divided by preset criteria, extracts a corpus corresponding to a modification request portion of the modification data within the first subtitle data, and calculates the suitability as the match rate between the original of the extracted corpus and the modification to the corpus is higher.
8. In paragraph 1, The above subtitle correction section, An artificial intelligence-based subtitle management device that calculates the suitability higher when the type of the subtitle modification requester is a content provider than when the type of the content viewer is the content provider, and calculates the suitability higher the more subtitle modification history, similar content viewing history, or supply history of the subtitle modification requester.
9. Performed by a computer, A step in which a subtitle generation unit obtains content data including video data and audio data from a first user terminal, and generates first subtitle data synchronized based on time information of the audio data through a subtitle generation model; A step in which the content provider resynchronizes the first subtitle data based on the motion information of the video data and provides the content data and the first subtitle data to the second user terminal by matching them; and The subtitle modification unit comprises a step of obtaining a subtitle modification request including modification data from the second user terminal, and modifying the first subtitle data based on the modification data to generate second subtitle data. The above image data includes a plurality of sequentially consecutive frames, The above content provider is, Classify the above multiple frames into multiple groups according to the above motion information, If the Nth frame (where N is a positive integer) and the N+1th frame included in the above plurality of frames contain different motion information, the Nth frame and the N+1th frame are classified into different groups, The above subtitle correction section, An artificial intelligence-based subtitle management method, wherein the suitability of the subtitle modification request is determined based on at least one of the matching rate between the modification data and the first subtitle data and the subtitle modification requester information, and if the suitability is greater than a threshold, the first subtitle data is modified to generate the second subtitle data.
10. A computer-readable recording medium coupled with a computer and storing a program for executing the method of claim 9.
Citation Information
Patent Citations
Caption correction apparatus
JP2007256714A
Method and apparatus for controlling playing video
KR1020150057591A
Glued fiberboard floor
KR1020240168136A
Mobile phone-based payment solution when the buyer and payer are different
KR1020240177753A
A Method for manufacturing vegan immune health fiber composition containing natural vegetable protein and method for manufacturing careware products containing the same
KR102604522B1