Video subtitle extraction method and device

The method enhances subtitle extraction by grouping and time-stamping video frames to address merging and synchronization issues, improving accuracy and efficiency in subtitle generation.

CN120318809AActive Publication Date: 2025-07-15BEIJING XIAODU INTERACTIVE ENTERTAINMENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510786885.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-15
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

In the prior art, there are many problems in the merging of video subtitles and time synchronization, which affects the completeness, accuracy and visual perception of subtitles, resulting in a decrease in translation accuracy.

Method used

By disassembling the video into an image frame, subtitles are identified using OCR technology, and based on the similarity of adjacent frames, the number of recurrences of subtitles and the text length of the subtitle group, the subtitle merging and time synchronization strategies are optimized to determine the start and ending moments of the target subtitles.

Benefits of technology

It improves the accuracy of subtitles recognition, time synchronization and processing efficiency, and achieves efficient and accurate subtitle generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318809A_ABST
    Figure CN120318809A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a video subtitle extraction method and device. The specific embodiment of the method comprises the following steps: splitting a target video into a plurality of image frames, and numbering each image frame according to a time sequence; determining an initial subtitle of each image frame; dividing the initial subtitles into a plurality of subtitle groups based on the similarity between the initial subtitles of the adjacent image frames; and for a subtitle group comprising a plurality of initial subtitles, determining a target subtitle in the subtitle group based on the repeated occurrence frequency and the text length of each initial subtitle in the subtitle group, and setting a starting moment and an ending moment of the target subtitle based on the serial number of the image frame corresponding to the subtitle group. According to the embodiment, the subtitles in the video can be continuously and accurately extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and particularly to a method and apparatus for video subtitle extraction. Background Art

[0002] Videos obtained on the Internet currently generally include videos with embedded subtitles and do not have independent subtitle files. Therefore, when subtitle files need to be extracted or subtitle translation is required, technology or manual labor is needed to obtain subtitles. After obtaining the subtitles, traditional translation generally relies on manual labor. Therefore, it generally takes a long time for a video to appear on the Internet and be translated into subtitles in other languages.

[0003] In the prior art, the method of using OCR (Optical Character Recognition) to extract video subtitles has been relatively mature, and usually, text recognition is performed on the pictures parsed frame by frame. However, there are still many problems in subtitle merging and time synchronization, which directly affect the integrity, accuracy, and visual perception of subtitles, and subtitles in turn affect the final accuracy of translation. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method and apparatus for video subtitle extraction.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for video subtitle extraction, including: disassembling a target video into multiple image frames, and numbering each image frame in chronological order; respectively determining initial subtitles of each image frame; based on the similarity between the initial subtitles of adjacent image frames, dividing each initial subtitle into multiple subtitle groups; for a subtitle group containing multiple initial subtitles, determining a target subtitle in the subtitle group based on the number of repeated occurrences and text length of each initial subtitle in the subtitle group, and setting the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

[0006] In some embodiments, the method further includes: calculating the similarity between target subtitles in adjacent subtitle groups; in response to detecting that the similarity between target subtitles is greater than a predetermined value, using the target subtitle with the longest text length as the target subtitle for each subtitle group in adjacent subtitle groups.

[0007] In some embodiments, respectively determining initial subtitles of each image frame includes: performing text recognition on each image frame through optical character recognition technology to obtain text information, text position, and confidence; filtering out invalid text information based on at least one of the text language, text length, confidence, and text position to obtain the initial subtitle.

[0008] In some embodiments, disassembling a target video into multiple image frames includes: decoding the middle part of the target video frame by frame to obtain at least one middle video frame; inputting the at least one middle video frame into a subtitle area detection model to output the subtitle areas of the respective middle video frames; determining the subtitle area of the target video according to the subtitle areas of the respective middle video frames; cropping the non-subtitle area of the target video according to the subtitle area of the target video to obtain a cropped video; and decoding the cropped video frame by frame to obtain multiple image frames.

[0009] In some embodiments, based on the similarity between the initial subtitles of adjacent image frames and dividing the initial subtitles into multiple subtitle groups includes: setting the initial subtitle of the first image frame as the comparison subtitle, setting the initial subtitle of the second image frame as the current subtitle, creating a subtitle group, and adding the comparison subtitle to the subtitle group; performing the following grouping steps: in response to the difference between the number of the image frame where the current subtitle is located and the number of the image frame where the comparison subtitle is located being within a predetermined range, calculating the similarity between the current subtitle and the comparison subtitle; in response to the similarity being greater than a predetermined threshold, adding the current subtitle to the subtitle group where the comparison subtitle is located, resetting the trial count to zero, and updating the current subtitle based on the initial subtitle of the next image frame of the image frame where the current subtitle is located, and continuing to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count not reaching the predetermined number, incrementing the trial count, and updating the current subtitle based on the initial subtitle of the next image frame of the image frame where the current subtitle is located, and continuing to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count reaching the predetermined number, creating a new subtitle group, using the current subtitle when the trial count is 1 as the new comparison subtitle, adding it to the new subtitle group, using the initial subtitle of the next image frame of the image frame where the new comparison subtitle is located as the new current subtitle, and continuing to perform the grouping steps.

[0010] In some embodiments, updating the current subtitle based on the initial subtitle of the next image frame of the image frame where the current subtitle is located and continuing to perform the grouping steps includes: in response to the length of the current subtitle being greater than the length of the comparison subtitle, updating the comparison subtitle based on the current subtitle, and using the initial subtitle of the next image frame of the image frame where the current subtitle is located as the new current subtitle, and continuing to perform the grouping steps.

[0011] In some embodiments, the method further includes: in response to the difference between the number of the image frame where the current subtitle is located and the number of the image frame where the comparison subtitle is located not being within a predetermined range, detecting whether the number of lines of the current subtitle is greater than 1 and equal to the number of lines of the comparison subtitle; in response to detecting that the number of lines of the current subtitle is greater than 1 and equal to the number of lines of the comparison subtitle, adding the current subtitle to the subtitle group where the comparison subtitle is located.

[0012] In some embodiments, determining the target subtitle in the subtitle group based on the number of repeated occurrences and the text length of each initial subtitle in the subtitle group includes: selecting a predetermined number of initial subtitles in descending order of the number of repeated occurrences of the initial subtitles in the subtitle group to obtain a candidate subtitle set; and selecting the candidate subtitle with the longest text length from the candidate subtitle set as the target subtitle in the subtitle group.

[0013] In some embodiments, the method further includes: obtaining the translation of the target subtitle in each subtitle group; and adding the translation to the target video according to the start time, end time, and text position of the target subtitle in each subtitle group.

[0014] In a second aspect, an embodiment of the present disclosure provides a video subtitle extraction device, including: a splitting unit configured to disassemble a target video into multiple image frames and number each image frame in chronological order; an extraction unit configured to respectively determine the initial subtitles of each image frame; a grouping unit configured to divide the initial subtitles into multiple subtitle groups based on the similarity between the initial subtitles of adjacent image frames; and a determining unit configured to, for a subtitle group including multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and the text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

[0015] In some embodiments, the determining unit is further configured to: calculate the similarity between the target subtitles in adjacent subtitle groups; and in response to detecting that the similarity between the target subtitles is greater than a predetermined value, use the target subtitle with the longest text length as the target subtitle for each subtitle group in the adjacent subtitle groups.

[0016] In some embodiments, the extraction unit is further configured to: perform optical character recognition on each image frame to obtain text information, text position, and confidence; and filter out invalid text information based on at least one of the language of the text information, the text length of the text information, the confidence, and the text position to obtain the initial subtitles.

[0017] In some embodiments, the splitting unit is further configured to: decode the middle part of the target video frame by frame to obtain at least one middle video frame; input the at least one middle video frame into a subtitle area detection model to output the subtitle areas of each middle video frame; determine the subtitle area of the target video according to the subtitle areas of each middle video frame; crop the non-subtitle area of the target video according to the subtitle area of the target video to obtain a cropped video; and decode the cropped video frame by frame to obtain multiple image frames.

[0018] In some embodiments, the grouping unit is further configured to: set the initial caption of the first image frame as the comparison caption, set the initial caption of the second image frame as the current caption, create a caption group, and add the comparison caption to the caption group; perform the following grouping steps: in response to the difference between the number of the image frame where the current caption is located and the number of the image frame where the comparison caption is located being within a predetermined range, calculate the similarity between the current caption and the comparison caption; in response to the similarity being greater than a predetermined threshold, add the current caption to the caption group where the comparison caption is located, reset the trial count to zero, and update the current caption based on the initial caption of the next image frame after the image frame where the current caption is located, and continue to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count not reaching the predetermined number, accumulate the trial count, and update the current caption based on the initial caption of the next image frame after the image frame where the current caption is located, and continue to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count reaching the predetermined number, create a new caption group, use the current caption when the trial count is 1 as the new comparison caption, and add it to the new caption group, use the initial caption of the next image frame after the image frame where the new comparison caption is located as the new current caption, and continue to perform the grouping steps.

[0019] In some embodiments, the grouping unit is further configured to: in response to the length of the current caption being greater than the length of the comparison caption, update the comparison caption based on the current caption, and use the initial caption of the next image frame after the image frame where the current caption is located as the new current caption, and continue to perform the grouping steps.

[0020] In some embodiments, the grouping unit is further configured to: in response to the difference between the number of the image frame where the current caption is located and the number of the image frame where the comparison caption is located not being within a predetermined range, detect whether the number of lines of the current caption is greater than 1 and equal to the number of lines of the comparison caption; in response to detecting that the number of lines of the current caption is greater than 1 and equal to the number of lines of the comparison caption, add the current caption to the caption group where the comparison caption is located.

[0021] In some embodiments, the determining unit is further configured to: select a predetermined number of initial captions in descending order of the number of repeated occurrences of the initial captions in the caption group to obtain a candidate caption set; select the candidate caption with the longest text length from the candidate caption set as the target caption in the caption group.

[0022] In some embodiments, the device further includes an editing unit, which is configured to: obtain the translation of the target caption in each caption group; add the translation to the target video according to the start time, end time, and text position of the target caption in each caption group.

[0023] In a third aspect, embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more computer programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement the method according to any one of the first aspect.

[0024] In a fourth aspect, embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the method according to any one of the first aspect.

[0025] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, which, when executed by a processor, implements the method according to any one of the first aspect.

[0026] The video subtitle extraction method and device provided by the embodiments of the present disclosure optimize subtitle merging and time synchronization strategies based on the existing OCR recognition results, provide an efficient, accurate, and intelligent subtitle generation method, can significantly improve the recognition accuracy, time synchronization, rationality of the merging strategy, and processing efficiency of subtitles, and have high application value and market prospects.

[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied; Figure 2 is a flowchart of an embodiment of the video subtitle extraction method according to the present disclosure; Figure 3 is a schematic diagram of an application scenario of the video subtitle extraction method according to the present disclosure; Figure 4 is a flowchart of another embodiment of the video subtitle extraction method according to the present disclosure; Figure 5 is a schematic structural diagram of an embodiment of the video subtitle extraction device according to the present disclosure; Figure 6 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only the parts related to the relevant invention are shown in the drawings.

[0030] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The present disclosure will be described in detail below with reference to the drawings and embodiments.

[0031] Figure 1 An exemplary system architecture 100 is shown, which is an embodiment of a video subtitle extraction method or a video subtitle extraction device to which the present disclosure can be applied.

[0032] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0033] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as video players, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0034] The terminal devices 101, 102, 103 may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices with a display screen and supporting video playback, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (for example, to provide distributed services), or it may be implemented as a single software or software module. No specific limitation is made here.

[0035] Server 105 can be a server that provides various services. For example, it can be a background video playback server that supports videos displayed on terminal devices 101, 102, and 103. The background video playback server can extract the embedded subtitles from the video, translate the subtitles according to the language required by the user, and finally add the translated subtitles to the video. The extracted subtitles can also be used for other purposes, such as content compliance detection, etc.

[0036] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or as a single software or software module. No specific limitation is made here. The server can also be a server of a distributed system or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0037] It should be noted that the video subtitle extraction method provided by the embodiments of the present disclosure is generally executed by server 105. Correspondingly, the video subtitle extraction device is generally set in server 105. It can also be executed by the terminal device. Correspondingly, the video subtitle extraction device is set in the terminal device.

[0038] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in

[0039] Continue to refer to Figure 2 , which shows a flow 200 of an embodiment of the video subtitle extraction method according to the present disclosure. The video subtitle extraction method includes the following steps: Step 201, disassemble the target video into multiple image frames and number each image frame in chronological order.

[0040] In this embodiment, the execution subject of the video subtitle extraction method (such as Figure 1 the server shown) can obtain the target video whose subtitles are to be extracted from the terminal device through a wired connection method or a wireless connection method. The execution subject can disassemble the target video into multiple image frames through a decoder such as ffmpeg. Then number each image frame in chronological order. The numbering can be consecutive natural numbers from small to large.

[0041] Step 202, respectively determine the initial subtitles of each image frame.

[0042] In this embodiment, text can be recognized from each image frame through OCR, and if text is recognized, initial subtitles can be obtained. If no text is recognized in the image frame, the image frame can be filtered out. Each initial subtitle will be associated with the number of the image frame where it is located, that is, the number of the image frame can be used as the index of the initial subtitle.

[0043] Step 203: Based on the similarity between the initial subtitles of adjacent image frames, divide each initial subtitle into multiple subtitle groups.

[0044] In this embodiment, the similarity between the initial subtitles of adjacent image frames can be calculated by calculating the edit distance, string matching degree, etc. The similarity between the initial subtitles of adjacent image frames can be calculated pairwise, and then the adjacent image frames with a similarity greater than a predetermined threshold can be divided into a group. Optionally, the initial subtitles of multiple image frames can be clustered and grouped through a clustering algorithm.

[0045] Step 204: For a subtitle group containing multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

[0046] In this embodiment, the same subtitle group shares a target subtitle. Therefore, select the best subtitle from the subtitle group as the target subtitle. The initial subtitle with the most repeated occurrences and the longest text length in the subtitle group can be used as the target subtitle. Optionally, weights can be set for the number of repeated occurrences and text length respectively, and the score of the initial subtitle can be calculated according to the weighted sum of the number of repeated occurrences and text length, and the initial subtitle with the highest score is used as the target subtitle.

[0047] Take the number of the first image frame in each subtitle group as the start time of the target subtitle of the subtitle group, and take the number of the last image frame in the subtitle group as the end time of the target subtitle of the subtitle group.

[0048] In the prior art, since the display of subtitles in a video is dynamic, the same subtitle content may be detected in different frames, resulting in repeated storage. If the text recognized by OCR is simply concatenated frame by frame, it will cause subtitle duplication, misalignment or loss. For example, a complete subtitle may be split into multiple frames, but due to discontinuous timeline information, direct concatenation will result in semantic errors or content loss. The method provided in the above embodiments of the present disclosure can make the segmentation and merging of subtitles more accurate.

[0049] In the prior art, the subtitle text recognized by OCR usually does not have accurate time information. The current mainstream approach is to calculate the time range when the subtitle appears based on the frame rate. However, since the subtitle switch is not instantaneous but has a certain fade-in, fade-out or overlap, it is difficult to accurately locate the start and end times of the subtitle. For example, a subtitle may continuously appear within 100 frames, but the same text may not be completely recognized in every frame, resulting in deviation in the calculation of the subtitle start and end times. The method provided in the above embodiments of the present disclosure can make the display time of the subtitle accurately match.

[0050] In the prior art, in some cases, a complete subtitle sentence may be misrecognized by OCR as two independent sentences, and in other cases, two different subtitles may be wrongly merged due to close time proximity, affecting the correct presentation of the subtitle. For example, movie dialogues may span multiple shots, but the subtitle generation cannot correctly distinguish these scenes, resulting in the subtitles of different speakers being mixed together, affecting readability. The method provided in the above embodiments of the present disclosure can standardize the subtitle splitting strategy.

[0051] In summary, the method for extracting subtitles in this application can overcome the limitations in deduplication, time matching, and logical splitting in the prior art, so as to obtain coherent and accurate subtitle text.

[0052] In some alternative implementation manners of this embodiment, the method further includes: calculating the similarity between target subtitles in adjacent subtitle groups; in response to detecting that the similarity between the target subtitles is greater than a predetermined value, using the target subtitle with the longest text length as the target subtitle for each subtitle group in the adjacent subtitle groups.

[0053] Subtitles can be further merged by merging adjacent subtitle groups. Here, the subtitle groups where the target subtitles appear successively in the video playback order are adjacent subtitle groups. Similarly, the similarity can be calculated according to methods such as the edit distance and string matching degree. If the similarity is high, the target subtitle with the longest text length is used as the target subtitle for each subtitle group in the adjacent subtitle groups.

[0054] In some alternative implementation manners of this embodiment, determining the initial subtitles for each image frame respectively includes: performing text recognition on each image frame through optical character recognition technology to obtain text information, text position, and confidence; filtering out invalid text information according to at least one of the text language, text length, confidence, and text position to obtain the initial subtitles.

[0055] Optical character recognition technology is based on a neural network for image recognition. In addition to being able to detect text boxes and obtain text positions, it can also obtain the confidence of the detection results.

[0056] Invalid subtitles can be filtered out according to the language of the text information. For example, if all the videos to be recognized are English videos now, and non-English characters appear in the subtitles, they can be determined as invalid subtitles.

[0057] Invalid subtitles can be filtered out according to the text length of the text information. For example, if the text length <= 1 (in unicode mode, one Chinese character is 3 characters), it is determined as an invalid subtitle.

[0058] Text information with a confidence level lower than the predetermined confidence threshold can be determined as invalid subtitles. For example, if the confidence level >= 0.89, it is determined as a valid subtitle. If the confidence level >= 0.87 and (lowercase letters > 5 or spaces >= 1 or punctuation marks >= 1), it is also determined as a valid subtitle. Otherwise, the text information is determined as an invalid subtitle.

[0059] Invalid subtitles can be filtered out according to the text position. For example, if the subtitle is close to symmetric with respect to the abscissa of the center point of the picture (absolute value < 20), it is determined as a valid subtitle. Otherwise, it is a filtered subtitle (which can effectively filter the rolling cast list at the end of the video, non-valid subtitles at the beginning of the video, etc.).

[0060] In some alternative implementation manners of this embodiment, the target video is disassembled into multiple image frames, including: decoding the middle part of the target video frame by frame to obtain at least one middle video frame; inputting at least one middle video frame into a subtitle area detection model to output the subtitle areas of each middle video frame; determining the subtitle area of the target video according to the subtitle areas of each middle video frame; cropping the non-subtitle area of the target video according to the subtitle area of the target video to obtain a cropped video; decoding the cropped video frame by frame to obtain multiple image frames.

[0061] There is no need to detect subtitles from the complete video frames. Instead, detect the position of subtitles from the middle section of the video, then crop the video, and only decode the cropped video, which can not only reduce the decoding time but also reduce the time for OCR to recognize subtitles.

[0062] Even if there are characters at the beginning and end of the video, they may not necessarily be subtitles and may be the cast list. Therefore, determining the position of subtitles from the middle section of the video can avoid misjudgment.

[0063] The subtitle area detection model can be a neural network for pixel point classification, that is, a binary classifier. The probability that each pixel point in each image frame output by the binary classifier belongs to the subtitle area can be obtained. Based on the probability that each pixel point in each image frame belongs to the subtitle area, obtain the text position information of the target video; determine the subtitle area of the target video based on the text position information.

[0064] In some alternative implementation manners of this embodiment, based on the similarity between the initial captions of adjacent image frames, and dividing each initial caption into multiple caption groups, including: setting the initial caption of the first image frame as the comparison caption, setting the initial caption of the second image frame as the current caption, creating a caption group, and adding the comparison caption to the caption group; performing the following grouping steps: in response to the difference between the number of the image frame where the current caption is located and the number of the image frame where the comparison caption is located being within a predetermined range, calculating the similarity between the current caption and the comparison caption; in response to the similarity being greater than a predetermined threshold, adding the current caption to the caption group where the comparison caption is located, resetting the trial count to zero, and updating the current caption based on the initial caption of the next image frame of the image frame where the current caption is located, and continuing to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count not reaching the predetermined number of times, accumulating the trial count, and updating the current caption based on the initial caption of the next image frame of the image frame where the current caption is located, and continuing to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count reaching the predetermined number of times, creating a new caption group, using the current caption when the trial count is 1 as the new comparison caption, and adding it to the new caption group, using the initial caption of the next image frame of the image frame where the new comparison caption is located as the new current caption, and continuing to perform the grouping steps.

[0065] Traverse the valid initial captions and group them according to the similarity. Merge the similar initial captions into the target caption. To avoid jitter, try a few more times after detecting dissimilar initial captions. If the trial is successful, even if the initial captions of the previous few frames are dissimilar, they will be combined into a caption group and use the same target caption.

[0066] To prevent the caption group from being too large, control the difference in the numbers of the initial captions to be compared by the fps (frames per second).

[0067] If the difference in numbers is within fps / 2, directly merge the similar captions. If the difference in numbers exceeds fps / 2, then judge whether to merge according to the number of lines of the captions.

[0068] In some alternative implementation manners of this embodiment, updating the current caption based on the initial caption of the next image frame of the image frame where the current caption is located, and continuing to perform the grouping steps includes: in response to the length of the current caption being greater than the length of the comparison caption, updating the comparison caption based on the current caption, and using the initial caption of the next image frame of the image frame where the current caption is located as the new current caption, and continuing to perform the grouping steps.

[0069] Update the comparison caption according to the length of the caption, avoiding comparing only with adjacent captions. Thereby improving the accuracy of caption extraction.

[0070] In some optional implementations of the present embodiment, the method further includes: in response to the difference between the number of the image frame where the current subtitle is located and the number of the image frame where the comparison subtitle is located being not within a predetermined range, detecting whether the number of lines of the current subtitle is greater than 1 and is equal to the number of lines of the comparison subtitle; in response to detecting that the number of lines of the current subtitle is greater than 1 and is equal to the number of lines of the comparison subtitle, adding the current subtitle to the subtitle group where the comparison subtitle is located.

[0071] If similar subtitle contents are displayed in separate lines, they are divided into different subtitle groups.

[0072] In some optional implementations of the present embodiment, the target subtitles in the subtitle group are determined based on the number of repetitions and the text length of each initial subtitle in the subtitle group, including: selecting a predetermined number of initial subtitles in the subtitle group in descending order of the number of repetitions of the initial subtitles to obtain a set of candidate subtitles; and selecting a candidate subtitle with the longest text length from the set of candidate subtitles as the target subtitle in the subtitle group.

[0073] The accuracy of the subtitles can be improved by sorting them in descending order of the number of repetitions and then selecting the subtitles with the longest text length.

[0074] Continue to see Figure 3 , Figure 3 FIG. 1 is a schematic diagram of an application scenario of the video subtitle extraction method according to this embodiment. Figure 3 In the application scenario, the process of extracting subtitles is as follows: Step 1. Obtain the video, decode the 1-3 minutes of the middle part of the video into frame-by-frame images, use OCR technology to recognize the images, determine the most likely video area of the images corresponding to these few minutes based on OCR, and appropriately enlarge the area, such as the bottom 1 / 5; Step 2. Use ffmpeg to decode the video into frame-by-frame pictures with numbers (for example, cut off the bottom 1 / 5); Step 3. Use OCR to identify the captured image, determine whether the subtitles are available, and divide the subtitles into valid subtitles and filtered subtitles. Take English subtitles as an example: Language judgment: For example, if the videos to be recognized are all in English, and non-English subtitles appear, they can be considered as filtered subtitles. Length determination: if the text length is less than or equal to 1 (in unicode mode, a Chinese character is 3 characters), the subtitles are filtered. Confidence judgment: if the confidence >= 0.89, it is considered a valid subtitle; if the confidence >= 0.87 and (lowercase letters > 5 or spaces >= 1 or punctuation >= 1); Coordinate judgment: If the abscissa of the subtitle relative to the center point of the picture is approximately symmetric (absolute value < 20), it is determined as a valid subtitle; otherwise, it is a filtered subtitle (which can effectively filter the cast list scrolling at the end of the film, non-valid subtitles at the beginning of the film, etc.).

[0075] Step 4. Traverse the valid subtitles:

[0076] If it is the first traversal, then beg = 0, create a temporary subtitle group group, and the comparison subtitle firstword of the temporary subtitle group;

[0077] If it is within fps / 2 frames, determine whether the edit distance between the current subtitle currentword and firstword is within N times (N = 3, text length <= 3, N = 1) or 75% of the same substring. If they are similar, add currentword to group. If the length of currentword is greater than the length of firstword, update the value of firstword to currentword; otherwise, probe forward M times (M = 5) and perform the above similarity judgment process; If it exceeds fps / 2 frames, but after splitting with the delimiter, the number of array elements is greater than 1 (i.e., the number of subtitle lines is greater than 1), and the number of lines of the two subtitles is exactly the same, then they are considered subtitles in the same group; If the edit distance is more than N times (and after performing the operations in b and c), then it is considered that they should be different subtitles. At this time, use the lucky word strategy (first, sort by the number of occurrences from largest to smallest, and take the top 10; second, sort by length from largest to smallest. When the lengths are the same, then sort by frequency from largest to smallest and take the first one) to obtain the most suitable subtitle txt in group, add the start frame number beg and end frame number end corresponding to the group of subtitles in group, the most suitable subtitle txt to the srtfile subtitle group, clear the subtitle group group, assign the current frame number to beg and end, and assign the current subtitle currentword to firstword.

[0078] Step 5. Loop the above steps to finally obtain the first version of the subtitle srtfile; Step 6. To prevent subtitles such as x, y, xy that should be merged from appearing, perform a final traversal of srtfile. Traverse srtfile to obtain beg, end, txt, and perform an edit distance comparison (N = 3, same strategy as above) or 75% of the same substring between the current subtitle and the next subtitle. If they are similar, update the current subtitle text to the longer text, and extend the current subtitle time range to the end time of the adjacent subtitle; otherwise, keep it unchanged. Finally, obtain the final subtitle and output it.

[0079] Further referenceFigure 4 , which shows the process 400 of another embodiment of the video subtitle extraction method. The process 400 of the video subtitle extraction method includes the following steps: Step 401, disassemble the target video into multiple image frames and number each image frame in chronological order.

[0080] Step 402, respectively determine the initial subtitles of each image frame.

[0081] Step 403, based on the similarity between the initial subtitles of adjacent image frames, divide each initial subtitle into multiple subtitle groups.

[0082] Step 404, for a subtitle group containing multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and the text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

[0083] Steps 401-404 are basically the same as steps 201-204, so they will not be elaborated here.

[0084] Step 405, obtain the translations of the target subtitles in each subtitle group.

[0085] In this embodiment, the translations of the target subtitles in each subtitle group can be obtained through manual translation or AI translation.

[0086] Step 406, add the translation to the target video according to the start time, end time and text position of the target subtitle in each subtitle group.

[0087] In this embodiment, through the position coordinates of the text obtained in step 402, the exact position of the coordinates in the target video can be calculated. After translation, the corresponding subtitles and time (start time, end time) can be obtained. Through the subtitles, time and coordinates, the translated subtitles can be accurately displayed below the embedded subtitles in the front-end player, achieving a perfect display of the subtitles. Optionally, if bilingual subtitles are not required, the original subtitles in the video can be erased and replaced with the target subtitles.

[0088] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a video subtitle extraction device. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0089] As Figure 5As shown in the figure, the video subtitle extraction device 500 of this embodiment includes: a splitting unit 501, an extraction unit 502, a grouping unit 503, and a determination unit 504. Among them, the splitting unit 501 is configured to disassemble the target video into multiple image frames and number each image frame in chronological order; the extraction unit 502 is configured to respectively determine the initial subtitles of each image frame; the grouping unit 503 is configured to divide each initial subtitle into multiple subtitle groups based on the similarity between the initial subtitles of adjacent image frames; the determination unit 504 is configured to, for a subtitle group containing multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

[0090] In this embodiment, the specific processing of the splitting unit 501, the extraction unit 502, the grouping unit 503, and the determination unit 504 of the video subtitle extraction device 500 can refer to Figure 2 Steps 201, 202, 203, and 204 in the corresponding embodiment.

[0091] In some alternative implementation manners of this embodiment, the determination unit 504 is further configured to: calculate the similarity between the target subtitles in adjacent subtitle groups; in response to detecting that the similarity between the target subtitles is greater than a predetermined value, use the target subtitle with the longest text length as the target subtitle for each subtitle group in the adjacent subtitle groups.

[0092] In some alternative implementation manners of this embodiment, the extraction unit 502 is further configured to: perform character recognition on each image frame through optical character recognition technology to obtain text information, text positions, and confidence levels; filter out invalid text information based on at least one of the text language, text length, confidence level, and text position to obtain the initial subtitles.

[0093] In some alternative implementation manners of this embodiment, the splitting unit 501 is further configured to: decode each frame of the middle part of the target video to obtain at least one intermediate video frame; input the at least one intermediate video frame into a subtitle area detection model and output the subtitle areas of each intermediate video frame; determine the subtitle area of the target video based on the subtitle areas of each intermediate video frame; crop the non-subtitle area of the target video according to the subtitle area of the target video to obtain a cropped video; decode the cropped video frame by frame to obtain multiple image frames.

[0094] In some alternative implementation manners of this embodiment, the grouping unit 503 is further configured to: set the initial caption of the first image frame as the comparison caption, set the initial caption of the second image frame as the current caption, create a caption group, and add the comparison caption to the caption group; perform the following grouping steps: in response to the difference between the number of the image frame where the current caption is located and the number of the image frame where the comparison caption is located being within a predetermined range, calculate the similarity between the current caption and the comparison caption; in response to the similarity being greater than a predetermined threshold, add the current caption to the caption group where the comparison caption is located, reset the trial count to zero, update the current caption based on the initial caption of the next image frame of the image frame where the current caption is located, and continue to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count not reaching the predetermined number of times, accumulate the trial count, update the current caption based on the initial caption of the next image frame of the image frame where the current caption is located, and continue to perform the grouping steps; in response to the similarity being less than or equal to the predetermined threshold and the trial count reaching the predetermined number of times, create a new caption group, use the current caption when the trial count is 1 as the new comparison caption, add it to the new caption group, use the initial caption of the next image frame of the image frame where the new comparison caption is located as the new current caption, and continue to perform the grouping steps.

[0095] In some alternative implementation manners of this embodiment, the grouping unit 503 is further configured to: in response to the length of the current caption being greater than the length of the comparison caption, update the comparison caption based on the current caption, and use the initial caption of the next image frame of the image frame where the current caption is located as the new current caption, and continue to perform the grouping steps.

[0096] In some alternative implementation manners of this embodiment, the grouping unit 503 is further configured to: in response to the difference between the number of the image frame where the current caption is located and the number of the image frame where the comparison caption is located not being within the predetermined range, detect whether the number of lines of the current caption is greater than 1 and equal to the number of lines of the comparison caption; in response to detecting that the number of lines of the current caption is greater than 1 and equal to the number of lines of the comparison caption, add the current caption to the caption group where the comparison caption is located.

[0097] In some alternative implementation manners of this embodiment, the determining unit 504 is further configured to: select a predetermined number of initial captions in descending order of the number of repeated occurrences of the initial captions in the caption group to obtain a candidate caption set; select the candidate caption with the longest text length from the candidate caption set as the target caption in the caption group.

[0098] In some alternative implementation manners of this embodiment, the device further includes an editing unit (not shown in the drawings), which is configured to: obtain the translation of the target caption in each caption group; add the translation to the target video according to the start time, end time, and text position of the target caption in each caption group.

[0099] It should be noted that in the technical solution of the present disclosure, in aspects such as the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information, it complies with the provisions of relevant laws and regulations, is used for legal purposes, and does not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data, and to safeguard the security of user personal information, network security, and national security.

[0100] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0101] An electronic device includes: one or more processors; a storage device on which one or more computer programs are stored, and when the one or more computer programs are executed by the one or more processors, the one or more processors implement the method described in process 200 or 400.

[0102] A computer-readable medium stores a computer program, wherein when the computer program is executed by a processor, it implements the method described in process 200 or 400.

[0103] Figure 6 A schematic block diagram of an exemplary electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0104] As Figure 6 shown, the device 600 includes a computing unit 601, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0105] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as a keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as a disk, optical disc, etc.; and communication unit 609, such as a network card, modem, wireless communication transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0106] Computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 601 executes the various methods and processes described above, such as the cell planning method. For example, in some embodiments, the cell planning method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the cell planning method described above can be executed. Alternatively, in other embodiments, computing unit 601 can be configured to execute the cell planning method in any other suitable way (e.g., by means of firmware).

[0107] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0108] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0111] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0112] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology. The server can be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0113] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0114] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for extracting video subtitles, comprising: Decompose a target video into multiple image frames and number each image frame in chronological order; Determine the initial subtitles of each image frame respectively; Based on the similarity between the initial subtitles of adjacent image frames, divide each initial subtitle into multiple subtitle groups; For a subtitle group containing multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and the text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

2. The method according to claim 1, wherein, The method further comprises: Calculate the similarity between the target subtitles in adjacent subtitle groups; In response to detecting that the similarity between the target subtitles is greater than a predetermined value, use the target subtitle with the longest text length as the target subtitle for each subtitle group in the adjacent subtitle groups.

3. The method according to claim 1, wherein The step of respectively determining the initial subtitles of each image frame includes: Perform optical character recognition on each image frame to obtain text information, text positions, and confidence levels; Filter out invalid text information based on at least one of the language of the text information, the text length of the text information, the confidence level, and the text position to obtain the initial subtitles.

4. The method according to claim 1, wherein, The step of decomposing the target video into multiple image frames includes: Decode the middle part of the target video frame by frame to obtain at least one intermediate video frame; Input the at least one intermediate video frame into a subtitle area detection model to output the subtitle areas of each intermediate video frame; Determine the subtitle area of the target video based on the subtitle areas of each intermediate video frame; Crop the non-subtitle area of the target video according to the subtitle area of the target video to obtain a cropped video; Decode the cropped video frame by frame to obtain multiple image frames.

5. The method according to claim 1, wherein The step of dividing each initial subtitle into multiple subtitle groups based on the similarity between the initial subtitles of adjacent image frames includes: Set the initial subtitle of the first image frame as the comparison subtitle, set the initial subtitle of the second image frame as the current subtitle, create a subtitle group, and add the comparison subtitle to the subtitle group; Execute the following grouping steps: in response to the difference between the number of the image frame where the current subtitle is located and the number of the image frame where the comparison subtitle is located being within a predetermined range, calculate the similarity between the current subtitle and the comparison subtitle; in response to the similarity being greater than a predetermined threshold, add the current subtitle to the subtitle group where the comparison subtitle is located, reset the trial count to zero, and update the current subtitle based on the initial subtitle of the next image frame after the image frame where the current subtitle is located, and continue to execute the grouping steps; In response to the similarity being less than or equal to the predetermined threshold and the trial count not reaching a predetermined number, accumulate the trial count, and update the current subtitle based on the initial subtitle of the next image frame after the image frame where the current subtitle is located, and continue to execute the grouping steps; In response to the similarity being less than or equal to the predetermined threshold and the number of attempts reaching the predetermined number, create a new subtitle group, use the current subtitle when the number of attempts is 1 as the new comparison subtitle, and add it to the new subtitle group. Use the initial subtitle of the next image frame of the image frame where the new comparison subtitle is located as the new current subtitle, and continue to execute the grouping step.

6. The method according to claim 5, wherein Updating the current subtitle based on the initial subtitle of the next image frame of the image frame where the current subtitle is located and continuing to execute the grouping step includes: In response to the length of the current subtitle being greater than the length of the comparison subtitle, update the comparison subtitle based on the current subtitle, and use the initial subtitle of the next image frame of the image frame where the current subtitle is located as the new current subtitle, and continue to execute the grouping step.

7. The method according to claim 6, wherein, The method further includes: In response to the difference between the number of the image frame where the current subtitle is located and the number of the image frame where the comparison subtitle is located not being within the predetermined range, detect whether the number of lines of the current subtitle is greater than 1 and equal to the number of lines of the comparison subtitle; In response to detecting that the number of lines of the current subtitle is greater than 1 and equal to the number of lines of the comparison subtitle, add the current subtitle to the subtitle group where the comparison subtitle is located.

8. The method according to claim 1, wherein, Determining the target subtitle in the subtitle group based on the number of repeated occurrences and text length of each initial subtitle in the subtitle group includes: Select a predetermined number of initial subtitles in descending order of the number of repeated occurrences of the initial subtitles in the subtitle group to obtain a candidate subtitle set; Select the candidate subtitle with the longest text length from the candidate subtitle set as the target subtitle in the subtitle group.

9. The method according to claim 3, wherein The method further includes: Obtain the translation of the target subtitle in each subtitle group; Add the translation to the target video according to the start time, end time and text position of the target subtitle in each subtitle group.

10. A video subtitle extraction device, comprising: A splitting unit configured to disassemble a target video into multiple image frames and number each image frame in chronological order; An extraction unit configured to respectively determine the initial subtitles of each image frame; A grouping unit configured to divide the initial subtitles into multiple subtitle groups based on the similarity between the initial subtitles of adjacent image frames; A determining unit configured to, for a subtitle group containing multiple initial subtitles, determine the target subtitle in the subtitle group based on the number of repeated occurrences and text length of each initial subtitle in the subtitle group, and set the start time and end time of the target subtitle based on the number of the image frame corresponding to the subtitle group.

11. An electronic device, comprising: One or more processors; A storage device having one or more computer programs stored thereon, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

12. A computer-readable medium having a computer program stored thereon, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Video caption recognition method and device, equipment and storage medium

    CN111582241A

  • Subtitle segmentation method, device and equipment and storage medium

    CN111709342A

  • Video subtitle identification method and device, medium and electronic equipment

    CN113052169A

  • Video subtitle extraction method and device and computer readable storage medium

    CN114363535A

  • Method and system for influencing digital content or access to content

    WO2017149321A1