Method and apparatus for generating marked video, method and apparatus for detecting video marking

By embedding watermarks at the video subtitle level and generating tagged videos using identifier sequences and subtitle offset types, the problem of insufficient robustness of watermarks in existing technologies is solved, and effective protection of video copyright is achieved.

CN116527965BActive Publication Date: 2026-05-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2022-01-24
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video digital watermarking technologies are not robust enough against attacks such as scaling, cropping, degrading, and editing, making it difficult to effectively protect video copyrights.

Method used

By embedding hidden watermarks in the subtitle dimension of the video, a tagged video is generated using identifier sequences and subtitle offset types, and the hidden watermark is embedded during playback. The watermark is detected using video encoding/decoding and computer vision techniques.

Benefits of technology

The robustness of the watermark has been improved, effectively resisting attacks such as scaling, cropping, degrading, and editing, ensuring the detection rate and accuracy during source tracing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527965B_ABST
    Figure CN116527965B_ABST
Patent Text Reader

Abstract

The application relates to a marked video generation method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining an object identifier and mapping the object identifier into an identifier sequence; determining the subtitle offset type corresponding to each identifier in the identifier sequence, and determining the clip information corresponding to each identifier; for each identifier, obtaining the subtitle offset segment pointed by the clip information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed by the clip information is obtained by performing subtitle offset on the original video segment with the same clip information in the original video; and based on the subtitle offset segment, a marked video corresponding to the original video and corresponding to the object identifier is generated. The method can realize strong robustness of video watermark embedding in the video field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network media technology, and in particular to a method, apparatus, computer device, storage medium, and computer program product for generating tagged videos, as well as a method, apparatus, computer device, storage medium, and computer program product for detecting video tags. Background Technology

[0002] In recent years, with the increasing awareness of copyright among the public, the importance of protecting the copyright of film and television works has become increasingly apparent. Due to its advantages of concealment, ease of traceability, and convenient operation, digital watermarking technology is being increasingly introduced into video copyright protection scenarios.

[0003] Existing video digital watermarking technologies primarily involve adding image watermarks to video frames to identify and display the video's source. However, under strong attacks such as scaling, cropping, degrading, and editing, image watermarks added in this way are easily compromised, making it difficult to effectively protect video copyright. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating labeled videos that can improve the robustness of digital watermarks, in response to the above-mentioned technical problems.

[0005] On one hand, this application provides a method for generating labeled videos. The method includes:

[0006] Obtain the object identifier and map the object identifier to a sequence of identifiers;

[0007] Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier;

[0008] For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier, which belongs to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by performing subtitle offset on the original video segments with the same segment information in the original video;

[0009] Based on the subtitle offset segment, a tagged video corresponding to the original video and the object identifier is generated.

[0010] On the other hand, this application also provides an apparatus for generating labeled videos. The apparatus includes:

[0011] The acquisition module is used to acquire the object identifier and map the object identifier to an identifier sequence;

[0012] The determination module is used to determine the subtitle offset type corresponding to each identifier in the identifier sequence, and to determine the segment information corresponding to each identifier.

[0013] The offset module is used to obtain, for each identifier, the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by offsetting the subtitles of the original video segments with the same segment information in the original video.

[0014] The generation module is used to generate a marked video that corresponds to the original video and the object identifier based on the subtitle offset segment.

[0015] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0016] Obtain the object identifier and map the object identifier to a sequence of identifiers;

[0017] Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier;

[0018] For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier, which belongs to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by performing subtitle offset on the original video segments with the same segment information in the original video;

[0019] Based on the subtitle offset segment, a tagged video corresponding to the original video and the object identifier is generated.

[0020] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0021] Obtain the object identifier and map the object identifier to a sequence of identifiers;

[0022] Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier;

[0023] For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier, which belongs to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by performing subtitle offset on the original video segments with the same segment information in the original video;

[0024] Based on the subtitle offset segment, a tagged video corresponding to the original video and the object identifier is generated.

[0025] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0026] Obtain the object identifier and map the object identifier to a sequence of identifiers;

[0027] Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier;

[0028] For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier, which belongs to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by performing subtitle offset on the original video segments with the same segment information in the original video;

[0029] Based on the subtitle offset segment, a tagged video corresponding to the original video and the object identifier is generated.

[0030] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for generating marked videos, through a certain logical mapping of the object identifier of the video playback party into a sequence of identifiers, and based on the subtitle offset type corresponding to each identifier in the identifier sequence and the segment information corresponding to each identifier, obtains the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier, and then combines the various subtitle offset segments to generate a marked video that corresponds to both the original video and the object identifier. In this way, the object identifier is used as a mark or watermark information and is implicitly embedded in the video in the form of subtitle offsets. Compared with the method of embedding in the image dimension, it is more robust and difficult to be destroyed by attacks such as scaling, cropping, degrading, and editing, thereby ensuring the detection rate and accuracy during source tracing.

[0031] On the other hand, this application also provides a method for detecting video tags. The method includes:

[0032] Acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected;

[0033] Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information respectively;

[0034] For each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong;

[0035] Based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined, and based on the identifier sequence, the object identifier marked in the video to be detected is determined.

[0036] On the other hand, this application also provides a video tag detection device. The device includes:

[0037] The acquisition module is used to acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected.

[0038] The determination module is used to determine the video segments in the video to be detected that correspond to each segment information according to the preset segment information;

[0039] The determining module is further configured to, for each video segment, determine an identifier corresponding to the video segment based on the subtitle offset type to which the multiple video frames included in the video segment belong;

[0040] The determining module is further configured to determine an identifier sequence based on the identifiers corresponding to each video segment in the video to be detected, and to determine the object identifier marked in the video to be detected based on the identifier sequence.

[0041] On the other hand, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0042] Acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected;

[0043] Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information respectively;

[0044] For each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong;

[0045] Based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined, and based on the identifier sequence, the object identifier marked in the video to be detected is determined.

[0046] On the other hand, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0047] Acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected;

[0048] Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information respectively;

[0049] For each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong;

[0050] Based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined, and based on the identifier sequence, the object identifier marked in the video to be detected is determined.

[0051] On the other hand, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0052] Acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected;

[0053] Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information respectively;

[0054] For each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong;

[0055] Based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined, and based on the identifier sequence, the object identifier marked in the video to be detected is determined.

[0056] The aforementioned video tag detection method, apparatus, computer equipment, storage medium, and computer program product determine the subtitle offset type corresponding to each video frame in the video to be detected, and determine the video segments in the video to be detected that correspond to each segment information according to preset segment information. For each video segment, based on the subtitle offset type to which the multiple video frames included in the video segment belong, an identifier corresponding to the video segment is determined. Then, based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined. Thus, the object identifier marked in the video to be detected is obtained by inverse mapping based on the identifier sequence. This realizes the detection of watermark information indirectly embedded in the marked video through subtitle offset, and can determine the video player or leakage source based on the marked object identifier, ensuring the detection rate and accuracy during source tracing. Attached Figure Description

[0057] Figure 1 This is an application environment diagram of a method for generating labeled videos in one embodiment;

[0058] Figure 2 This is a flowchart illustrating a method for generating labeled videos in one embodiment;

[0059] Figure 3 This is a flowchart illustrating the steps of a server acquiring a subtitle offset segment in one embodiment;

[0060] Figure 4 This is a flowchart illustrating a video tag detection method in one embodiment;

[0061] Figure 5A This is a schematic diagram of a subtitle watermark embedding system in one embodiment;

[0062] Figure 5B This is a schematic diagram of a subtitle watermark detection system in one embodiment;

[0063] Figure 6 This is a schematic diagram of the process for embedding watermarks on the A / B sides of a video in one embodiment;

[0064] Figure 7 This is a schematic diagram of the A / B side watermark video distribution process in one embodiment;

[0065] Figure 8 This is a schematic diagram of the video alignment process in one embodiment;

[0066] Figure 9 This is a flowchart illustrating the subtitle position detection process in one embodiment;

[0067] Figure 10 This is a flowchart illustrating the A / B sequence mapping process in one embodiment;

[0068] Figure 11 This is a structural block diagram of a device for generating labeled videos in one embodiment;

[0069] Figure 12 This is a structural block diagram of a video marker detection device in one embodiment;

[0070] Figure 13 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0072] In recent years, with the increasing awareness of copyright among the public, the importance of protecting the copyright of film and television works has become increasingly apparent. Due to its advantages of concealment, ease of traceability, and convenient operation, digital watermarking technology is being increasingly introduced into video copyright protection scenarios.

[0073] Digital watermarking is an information hiding technology that utilizes the limitations of human senses to tightly combine and hide digital signals (such as images, text, symbols, or numbers that can serve as markers) with the original data (such as image, audio, and video data). Digital watermarking can provide complete and reliable evidence of ownership for copyrighted information products.

[0074] Traditional video digital watermarking technology primarily involves overlaying watermarks at the video frame level, which means performing domain transformation on the image and embedding the watermark. However, this method of overlaying watermarks at the frame level has drawbacks such as high computational cost, impact on image quality, and unstable robustness. Under strong attacks such as scaling, cropping, degrading, and editing, the watermark is easily destroyed.

[0075] In view of this, embodiments of this application provide a method for generating marked videos and a corresponding method for detecting video marks, creatively embedding hidden watermarks at the subtitle level. By utilizing video encoding / decoding and computer vision technologies, embodiments of this application propose a method for embedding hidden watermarks during subtitle compression and playback, as well as for detecting watermarks in compressed videos. When a video is leaked or stolen, the source of the leak can be traced based on the hidden watermark, protecting video copyright. Compared to existing methods that embed watermarks at the image level, embodiments of this application embed hidden watermarks at the subtitle level, which not only improves the efficiency of watermark embedding but also has strong resistance to attacks and robustness.

[0076] To facilitate a better understanding of the technical content of this application, the relevant technical terms involved in the embodiments of this application are explained below.

[0077] Subtitles can generally be categorized into several forms, including hard subtitles, soft subtitles, and external subtitles. Hard subtitles are subtitles embedded within the video frame and become part of the image; they are visible as long as the video can be played. Soft subtitles are subtitles and video frame packaged in a single container; the subtitles can be displayed selectively during video playback or separated from the video frame; the subtitles and video frame are separate within the container. External subtitles are subtitles separated from the video container as a separate file; they can be loaded into the video container for playback using a playback tool. The subtitles involved in the embodiments of this application can specifically be soft subtitles or external subtitles, but are not limited to these.

[0078] Subtitles have formats, including but not limited to SRT (SubRipper Text), SSA (SubStation Alpha), and ASS (Advanced SubStation Alpha). Taking SRT format as an example, it consists of: a line of subtitle number, a line of time code, and a line of subtitle data.

[0079] For example: 45

[0081] 00:02:52,184-->00:02:53,617

[0082] A

[0083] This indicates the 45th subtitle, displayed from 2 minutes 52.184 seconds to 2 minutes 53.617 seconds into the video. The subtitle reads: A.

[0084] The solution of this application will be described in detail below. The method for generating marked videos provided in the embodiments of this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process, such as video data and subtitle data. The data storage system can be integrated onto server 104 or located in the cloud or on another network server. Terminal 102 requests video playback from server 104. Server 104 then obtains the object identifier corresponding to terminal 102 and generates a tagged video corresponding to the object identifier based on the identifier sequence mapped to the object identifier. Server 104 then sends the generated tagged video to terminal 102 for playback.

[0085] Terminal 102 can be, but is not limited to, one or more of various desktop computers, laptops, smartphones, tablets, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, etc. For example, terminal 102 can be a smart device capable of providing OTT (Over-The-Top) services, including but not limited to smart TVs and set-top boxes. OTT refers to providing various application services via the internet; typical OTT services include internet TV services and app stores. Terminal 102 can have applications installed, such as video playback applications, browsers, email applications, instant messaging applications, etc., without limitation. Specifically, applications can be standalone applications installed via an installation package, or mini-program applications that can be used without downloading and installation. The terminal can play videos through the installed applications.

[0086] Among them, server 104 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0087] In some embodiments, such as Figure 2 As shown, a method for generating tagged videos is provided. This method can be executed by a server or a terminal, or by both a server and a terminal. This application embodiment applies this method to... Figure 1 Taking the server in the example, the following steps are included:

[0088] Step S202: Obtain the object identifier and map the object identifier to an identifier sequence.

[0089] The object identifier refers to a unique identifier used to distinguish different objects or the terminals used by those objects. It can be, but is not limited to, one or more of the following: the object's account information (including but not limited to account name, object ID, etc.), the unique encoding information of the video playback application, or the IP (Internet Protocol Address) or MAC (Media Access Control Address) of the terminal with the video playback application installed. For example, if an object's registered account name on a video playback platform is "abc", then the corresponding object identifier is "abc". By obtaining the object identifier, it can be mapped and embedded as a watermark in the subtitles of the video, facilitating future tracing of leaked videos.

[0090] Because object identifiers vary, converting them into a unified identifier sequence through preset mapping rules enables computer devices to identify them faster and more accurately, thereby improving the efficiency of generating tagged videos. Typically, the identifier sequence has a fixed length and consists of a preset number of identifiers. The identifier specifies the subtitle offset type; one identifier uniquely corresponds to one subtitle offset type. Identifiers can be letters, numbers, or other characters. For example, if there are two preset subtitle offset types, they can be represented by binary digits "0" and "1," or by two different letters "A" and "B." Similarly, if there are more than two subtitle offset types, they can be represented by numbers 0-9 or letters A-Z, or by a combination of letters and numbers.

[0091] The term "subtitle offset type" refers to the method by which subtitles are offset. This offset type can include, but is not limited to, shifting the position of the subtitles, changing the character spacing, altering the font, color, and size, or not offsetting the subtitles at all (i.e., maintaining the original subtitle style). To minimize impact on the viewing experience, the subtitle offset should be subtle, so that viewers cannot or barely perceive the change. For example, one of the two subtitle offset types might be shifting all subtitles upwards by one pixel, while the other might be shifting all subtitles downwards by one pixel.

[0092] Specifically, a terminal can send a video playback request to the server through a running video playback application to obtain a specific video for playback. The video playback request includes video information (indicating which video it is, including but not limited to one or more of video name, video number, etc.) and an object identifier. In response to the terminal's video playback request, the server determines the corresponding original video and extracts the object identifier carried in the video playback request. The server converts the object identifier into an identifier sequence so that it can subsequently determine the tagged video corresponding to the original video based on this identifier sequence.

[0093] For example, if the object ID is "abc", the server can map it to AAAAAAABAABA, where AAAA represents "a", AAAB represents "b", and AABA represents "c". Alternatively, the server can map it to 000110, where 00 represents "a", 01 represents "b", and 10 represents "c". When the identifier sequence obtained by mapping the object identifier is not equal to the preset length, the server can process the identifier sequence according to preset logic, such as truncating or padding.

[0094] Step S204: Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier.

[0095] Segmentation refers to dividing a video into segments (or regions) along the time dimension. The corresponding segmentation information can be one or more of specific time information or segment sequence information. This segmentation information indicates the original video segments obtained by segmenting the original video. Similarly, this segmentation information also indicates the subtitle offset segments obtained by segmenting the subtitle offset video. For example, if the server divides the timeline into segments (or chunks) every 10 seconds, the segmentation information could indicate seconds 1-10, seconds 11-20, and so on. It's understandable that the duration of each segment can be uneven; for example, the segmentation information could also indicate seconds 1-10, seconds 11-30, etc. Alternatively, the segmentation information can also be a segment sequence number, indicating the sequence number of the segment obtained, such as segment 1, segment 2, etc.

[0096] Specifically, the server determines the subtitle offset type corresponding to each identifier in the identifier sequence obtained by mapping object identifiers. For example, the server sequentially determines which of several preset subtitle offset types corresponds to each identifier in the identifier sequence. Simultaneously, the server also needs to determine the segment information corresponding to each identifier to determine the subtitle offset type for each time period.

[0097] For example, the server pre-stores the correspondence between various identifiers and subtitle offset types. For instance, identifier "0" corresponds to the subtitle offset type that shifts the subtitle upwards, identifier "1" corresponds to the subtitle offset type that shifts the subtitle downwards, identifier "A" corresponds to the subtitle offset type that increases the subtitle character spacing, identifier "B" corresponds to the subtitle offset type that decreases the subtitle character spacing, and so on.

[0098] Step S206: For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by subtitle offsetting the original video segments with the same segment information in the original video.

[0099] Each identifier corresponds to at least one video segment. For example, in the identifier sequence, the first identifier indicates the first video segment, or the video segment from second 1 to second 10; the second identifier indicates the second video segment, or the video segment from second 11 to second 20; and so on.

[0100] It's understandable that, with long videos, each identifier can correspond to multiple video segments. For example, in a 1-hour original video, assuming the identifier sequence is 010, the first identifier "0" indicates the video segment from second 1 to second 10, the second identifier "1" indicates the video segment from second 11 to second 20, the third identifier "0" indicates the video segment from second 21 to third 30, and after iterating through all identifiers in the sequence, some subsequent video segments still lack corresponding identifiers. In this case, we can start again from the first identifier in the sequence and iterate through the sequence anew. Thus, each identifier can correspond to multiple video segments, such as the first identifier "0" indicating video segments from second 1 to second 10, second 31 to second 40, and so on.

[0101] Each segment information points to a subtitle offset segment, which refers to a video segment containing offset subtitles after the original video segment with the same segment information has been offset. Each subtitle offset segment includes at least one offset subtitle. In some embodiments, at least one subtitle offset segment includes more than one offset subtitle. That is, the offset subtitles included in a subtitle offset segment can be one, two, or more, etc., and this application embodiment does not limit this.

[0102] The subtitle offset segment can be obtained by pre-segmenting the original video with subtitle offset and storing it in a storage medium (storage space includes but is not limited to the local storage space of the server, database, etc.), or it can be obtained by pre-segmenting the original video segments with the same segment information with subtitle offset and storing them in a storage medium, or it can be obtained by directly extracting the original video segments with the same segment information from the original video and then performing subtitle offset.

[0103] Subtitle offset processing refers to offsetting subtitles. Methods of offsetting subtitles include, but are not limited to, shifting the position of the subtitles, changing the character spacing, changing the font, color, or size of the subtitles, or not offsetting the subtitles (i.e., maintaining the original subtitle style), one or more of these methods. The subtitle offset processing mentioned in this application embodiment can be adding offset subtitles to the original video (or original video segment), or offsetting the original subtitles in the original video (or original video segment).

[0104] It should be noted that subtitle offset processing is performed on a per-segment basis, based on the video segments obtained from the segmentation. For the same subtitle offset segment, the offset processing performed on the subtitles within it belongs to the same type. For example, in the first video segment (or the video segment from second 1 to second 10) containing 10 subtitles, the server performs type A subtitle offset processing on all the subtitles, such as shifting the position of each of the 10 subtitles upwards by 1 pixel, and so on.

[0105] For example, the server pre-applies various subtitle offset types to the original video. For instance, the server uniformly shifts the subtitles upwards by a certain number of pixels to obtain type A video; simultaneously, the server uniformly shifts the subtitles downwards by a certain number of pixels to obtain type B video. The server then segments both type A and type B videos according to a preset segmentation rule (e.g., segmenting every 10 seconds), resulting in multiple subtitle offset segments of type A and multiple subtitle offset segments of type B. Therefore, the server can directly extract subtitle offset segments of a specific subtitle offset type.

[0106] For example, the server pre-segments the original video according to preset segmentation rules, resulting in multiple video segments. The server then applies different subtitle offsets (Type A, Type B, Type C, etc.) to each video segment, thus creating a subtitle offset segment for each type of segment. Subsequently, the server can directly extract a subtitle offset segment of a specific type.

[0107] For example, the server may not pre-segment the video, but instead directly extract the video segment from the 10th to the 20th second of the original video according to the time period indicated by the segmentation information (e.g., indicating the 10th to the 20th second), and then perform a specific type of subtitle offset on the subtitles in the video segment to obtain a subtitle offset segment of a certain subtitle offset type.

[0108] Of course, this is not the only option. Those skilled in the art will understand that any method of subtitle offset can be applied without departing from the inventive concept and idea disclosed in this application. For example, the server can perform a specific type of subtitle offset on the subtitles in the video segment according to the time period indicated by the segmentation information, and then extract the video segment within that time period from the original video, and so on.

[0109] Specifically, for each identifier in the identifier sequence, the server determines the segment information corresponding to each identifier, and determines the subtitle offset segment that belongs to the subtitle offset type indicated by each identifier under the same segment information.

[0110] For example, suppose the identifier sequence obtained by mapping object identifiers is AAAAAAABAABA, where identifier A represents the subtitle offset type that shifts the subtitle upwards, and identifier B represents the subtitle offset type that shifts the subtitle downwards. For each identifier, the server sequentially determines its corresponding segment information. For example, the first identifier "A" corresponds to the video segment indicated by the segment information from second 1 to second 10, the eighth identifier "B" corresponds to the video segment indicated by the segment information from second 71 to second 80, and so on, the twelfth identifier "A" corresponds to the video segment indicated by the segment information from second 111 to second 120. Taking the first identifier as an example, the server determines which video segment its corresponding segment information points to, determines the subtitle offset type represented by the identifier, and determines the subtitle offset segment that corresponds to both the identifier and the segment information. The server can extract the corresponding subtitle offset segment from multiple pre-stored subtitle offset segments based on the segment information and identifier. Alternatively, it can perform subtitle offset on the subtitles in the video segment according to the time period indicated by the segment information, and then extract the video segment within that time period from the original video to obtain the corresponding subtitle offset segment.

[0111] Step S208: Based on the subtitle offset segment, generate a marked video that corresponds to the original video and the object identifier.

[0112] Specifically, the server reassembles the subtitle offset segments in chronological order based on each segment's information and corresponding identifier, thereby generating a tagged video that corresponds to the original video. Since the generated tagged video indirectly embeds object identifiers in the form of subtitle offsets, it facilitates the extraction of object identifiers from leaked videos by detecting subtitles, thus allowing for the tracing of the leaker.

[0113] In the aforementioned method for generating labeled videos, the object identifier of the video playback party is logically mapped into a sequence of identifiers. Based on the subtitle offset type corresponding to each identifier in the identifier sequence and the segment information corresponding to each identifier, the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier is obtained. These subtitle offset segments are then combined to generate a labeled video that corresponds to both the original video and the object identifier. Thus, the object identifier is implicitly embedded in the video as a type of label or watermark information in the form of subtitle offsets. Compared to embedding at the image dimension, this method has the advantages of low computational cost and high generation efficiency. Furthermore, since this application only slightly modifies the subtitles, the impact on the video quality is minimal and imperceptible, ensuring a good viewing experience. In addition, the watermarking method at the subtitle dimension is robust and difficult to be destroyed by attacks such as scaling, cropping, degrading, and editing, thus ensuring the detection rate and accuracy during source tracing.

[0114] Typically, an object identifier consists of multiple characters, including but not limited to one or more of numbers, letters, and special symbols. The server pre-defines the mapping relationship between each character and the identifier. For example, the character "a" corresponds to the identifier "0000", or the character "*" corresponds to the identifier "AAAB", etc. In some embodiments, mapping an object identifier to an identifier sequence includes: starting from the first character among the multiple characters, determining the identifier corresponding to each character sequentially; arranging the identifiers corresponding to each character according to a preset format to obtain an identifier sequence of a preset length.

[0115] Specifically, among the multiple characters constituting the object identifier, the server searches for the corresponding identifier sequentially, starting from the first character, according to a preset reading order, thus determining which identifier each character corresponds to. The reading order is not limited; it can be from the first character to the last or vice versa. However, for computer processing efficiency, it is usually set to read from the first character to the last. The server then arranges the identifiers according to a preset format, resulting in an identifier sequence of a certain length. Typically, the determined identifiers are arranged sequentially according to the order of their corresponding characters to form the final identifier sequence.

[0116] In the above embodiments, by mapping the object identifier to a fixed-length identifier sequence, different types of subtitle offset segments can be determined based on the identifier sequence, thereby obtaining a labeled video with indirectly embedded object identifiers. This achieves the embedding of subtitle watermarks into the video, facilitating subsequent source tracing.

[0117] In some embodiments, arranging the identifiers corresponding to each character according to a preset format to obtain an identifier sequence of a preset length includes: arranging the identifiers corresponding to each character according to a preset format; if the number of identifiers corresponding to all characters in the object identifier is less than a preset number, then padding is performed at the end of the arrangement using a pre-set padding identifier to obtain an identifier sequence of a preset length.

[0118] The padding identifier is a pre-defined identifier. To distinguish it from the identifiers corresponding to each subtitle offset type, it is usually set to a different character. For example, the identifiers corresponding to each subtitle offset type can be set to letters or numbers, while the padding identifier can be set to a special symbol, such as... Underscores "_", or unused letters or numbers, etc. Therefore, when the server reads the padding identifier, it can determine that the identifier does not correspond to a certain subtitle offset type, and the server can choose not to offset the subtitle. Of course, the padding identifier can also correspond to a subtitle offset type. For example, a pre-set mapping between padding identifiers and certain subtitle offset types can be used. When the server reads the padding identifier, it can offset the subtitle according to the subtitle offset type corresponding to that identifier.

[0119] Specifically, after arranging the identifiers corresponding to each character according to a preset format, if the resulting identifier sequence is less than a preset length, in other words, the number of identifiers corresponding to all characters in the object identifier is less than a preset number, the server pads the end of the arrangement (i.e., the identifier sequence of a certain length) with a certain number of padding identifiers to make the length of the identifier sequence reach the preset length.

[0120] In the above embodiments, by mapping the object identifier to a fixed-length identifier sequence, it is easier for the server to extract and determine the object identifier during subsequent tracing, thereby improving the accuracy of tracing.

[0121] As mentioned above, there are multiple ways for a server to obtain subtitle offset segments. For example, the server can pre-process subtitle offsets to obtain subtitle offset segments, and then only needs to extract several subtitle offset segments. Therefore, in some embodiments, for each identifier, obtaining the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier includes: obtaining pre-processed subtitle offset videos corresponding to each subtitle offset type, wherein the subtitle offset videos are obtained by performing corresponding subtitle offset processing on the original video; for each identifier in the identifier sequence, according to the segment information corresponding to the corresponding identifier, extracting subtitle offset segments with the same segment information from the subtitle offset videos belonging to the subtitle offset type corresponding to the corresponding identifier.

[0122] Specifically, the server performs corresponding subtitle offset processing on the original video segments, such as performing subtitle offset of type A, type B, type C, etc., to obtain subtitle offset videos of multiple types. When processing each identifier in the identifier sequence, the server determines the video segment of which time period or segment number it corresponds to based on the segment information, and then extracts the subtitle offset segment of a certain subtitle offset type corresponding to that time period or segment number. In some examples, the server can pre-process the subtitle offset videos of each type, obtain the subtitle offset segments corresponding to each segment information, and store them. This allows for direct retrieval of the corresponding subtitle offset segment from the storage medium based on the time period or sequence number indicated by the segment information when retrieving subtitle offset segments later. Alternatively, in some examples, the server can also store subtitle offset videos of each type and then extract the corresponding subtitle offset segments from the subtitle offset videos later based on the time period or sequence number indicated by the segment information.

[0123] In the above embodiments, by preprocessing to obtain subtitle offset videos corresponding to each subtitle offset type, the corresponding subtitle offset segments can be directly extracted, thereby improving the efficiency of obtaining subtitle offset segments.

[0124] In some embodiments, the subtitle offset video corresponding to each of the multiple subtitle offset types is generated by the following steps: decoding the original video to obtain multiple video frames, and adding offset subtitles of the same subtitle offset type to each frame to obtain multiple video frames of the corresponding subtitle offset type; encoding the multiple video frames with offset subtitles of the corresponding subtitle offset type to obtain the subtitle offset video corresponding to the subtitle offset type.

[0125] Specifically, the server first decodes the original video, obtaining multiple video frames. For example, the server can use FFMPEG (Fast Forward MPEG, a multimedia video processing tool) for video encoding and decoding. FFMPEG is open-source free software capable of recording, converting, and streaming various audio and video formats. For each decoded video frame, the server adds a subtitle of a specific offset type frame-by-frame, resulting in multiple video frames corresponding to that subtitle offset type. The server then encodes these multiple video frames with the corresponding subtitle offset type, reassembling them into a new video, thus obtaining a subtitle-offset video of that subtitle offset type. This process is repeated for each subtitle offset type, resulting in various types of subtitle-offset videos.

[0126] In the above embodiments, by decoding the original video and adding offset subtitles frame by frame, and then encoding and recombining them into a video, various types of subtitle offset videos can be obtained quickly, and the embedding efficiency of subtitle watermarks is high.

[0127] Another way for the server to obtain subtitle offset segments is to pre-segment the original video according to a preset segmentation rule, obtaining multiple video segments; then, subtitle offset processing is performed on the video segments according to identifiers, thereby obtaining subtitle offset segments corresponding to the corresponding identifiers. Therefore, in some embodiments, such as... Figure 3 As shown, for each identifier, the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier is obtained, including:

[0128] Step S302: Obtain the original video and perform segmentation processing on the original video to obtain multiple original video segments.

[0129] Step S304: Determine the original video segment pointed to by the segment information corresponding to each identifier in the identifier sequence.

[0130] Step S306: For the original video segments corresponding to each identifier, perform subtitle offset according to the subtitle offset type corresponding to the corresponding identifier to obtain the subtitle offset segment corresponding to the corresponding identifier.

[0131] Specifically, the server acquires the original video and segments it according to preset segmentation rules to obtain multiple original video segments. For example, the server can extract the specific video to be played based on a video playback request sent by the terminal.

[0132] For each identifier in the identifier sequence obtained by mapping the acquired object identifier, the server determines the segment information corresponding to that identifier to determine which specific video segment needs to be processed for subtitle offsetting. After determining the video segment, the server performs subtitle offsetting on the original video segment corresponding to the identifier according to the subtitle offset type, resulting in a subtitle offset segment corresponding to the corresponding identifier.

[0133] For example, the server can make slight changes to the position, character spacing, size, etc., of the original subtitles in the original video clip, thereby achieving subtitle offsetting of the original video clip. Alternatively, the server can add offset subtitles to an original video clip without subtitles, thus achieving subtitle offsetting of the original video clip. In this way, the server can obtain the subtitle offset segment corresponding to each identifier.

[0134] In the above embodiments, subtitle offset is performed on the original video segments obtained by segmentation based on the identifiers in the mapped identifier sequence, which is more efficient than subtitle embedding at the screen dimension; at the same time, the server does not need to pre-store various types of subtitle offset segments, saving storage resources.

[0135] After acquiring multiple subtitle offset segments, the server needs to recombine these segments into a complete video for transmission to the terminal. To this end, in some embodiments, a tagged video corresponding to both the original video and the object identifier is generated based on the subtitle offset segments. This includes: splicing the subtitle offset segments corresponding to each segment information according to their respective time sequences to generate a tagged video that corresponds to both the original video and the object identifier.

[0136] Specifically, the server splices together the subtitle offset segments corresponding to each segment according to their respective time sequences, thereby generating a tagged video that corresponds to the original video. Since the generated tagged video indirectly embeds object identifiers in the form of subtitle offsets, it facilitates the extraction of object identifiers from leaked videos by detecting subtitles, thus allowing for the tracing of the leaker.

[0137] For example, if the segment information is such as 1 second to 10 seconds, 11 seconds to 20 seconds, etc., then the server will sequentially splice the subtitle offset segments corresponding to 1 second to 10 seconds, 11 seconds to 20 seconds, etc., according to the time sequence, and encode them using tools such as FFMPEG, thereby converting the spliced ​​subtitle offset segments into a complete marked video.

[0138] For example, if the segment information is such as the first segment, the second segment, etc., the server will sequentially splice the subtitle offset segments corresponding to the first segment, the second segment, etc., according to the order of the segment numbers, and then convert the spliced ​​subtitle offset segments into a complete marked video.

[0139] In the above embodiments, by splicing the subtitle offset segments to generate a marked video with an indirect embedded object identifier, it is convenient to extract the object identifier by detecting the subtitles in the leaked video in the future, thereby tracing the source of the leaked video.

[0140] As mentioned earlier, in some scenarios, there may be a situation where the identifiers corresponding to various video segments in a complete video are repeating cyclical sequences of identifiers. When the server detects tagged videos to determine the source of the leak, it needs to detect the identifier sequence to identify the object. For this purpose, start and end identifiers are also set in the identifier sequence, which are used by computer devices to identify and determine the start and / or end of the identifier sequence.

[0141] The start and end identifiers can be set to one or more characters that are different from other identifiers (including identifiers corresponding to subtitle offset types and padding identifiers). For example, they can be set to a single special character or special string, or unused letters and numbers. The start and end identifiers can be set at the beginning or end of the identifier sequence, or at both the beginning and end.

[0142] For example, suppose the identifier sequence is "@10001010". The first character is the start and end identifier "@", indicating that the identifier starting from the next character corresponds to the subtitle offset type. The part up to the next start and end identifier "@", i.e., the "10001010" portion, is the part of the identifier sequence corresponding to the subtitle offset type. Starting from the next start and end identifier "@", the identifiers in the identifier sequence are traversed again.

[0143] Therefore, in some embodiments, the tagged video includes multiple tagged units, each tagged unit including multiple subtitle offset segments corresponding to a single identifier sequence, and the multiple tagged units are distinguished by the video segments corresponding to the start and end identifiers in the identifier sequence, the start and end identifiers being set at at least one position at the beginning and end of the identifier sequence.

[0144] Specifically, the server marks multiple subtitle offset segments corresponding to a single identifier sequence as a single marker unit, and a complete video includes multiple marker units. Each marker unit is distinguished by start and end identifiers in the identifier sequence.

[0145] For example, assuming the identifier sequence is "@ABC", for seconds 1-10, identifier "A" corresponds to subtitle offset segment X; for seconds 11-20, identifier "B" corresponds to subtitle offset segment Y; and for seconds 21-30, identifier "C" corresponds to subtitle offset segment Z. After traversing the identifier sequence, the three subtitle offset segments from seconds 1-30 constitute a marker unit. For seconds 21-30, the server retraces the identifier sequence, continuing the process of determining the subtitle offset segment corresponding to each identifier. Therefore, when the server detects the complete marked video, it can determine the identifier sequence as "@ABC" based on the start and end identifiers by extracting multiple identifiers "@ABC@ABC@ABC…". For example, for the multiple identifiers "!1001010!!1001010!!10001010!..." extracted by the server from a complete tagged video, the identifier sequence can be extracted as "!1001010!", where a start and end identifier "!" is set at both the beginning and end of the identifier sequence.

[0146] In the above embodiments, by setting start and end identifiers and grouping multiple subtitle offset segments corresponding to a single identifier sequence into a single tagging unit, the server can quickly extract the identifier sequence when detecting the video, thereby obtaining the object identifier.

[0147] In a specific example, the specific process of the method for generating tagged videos provided in this application includes: the terminal sends a video playback request to the server, the video playback request containing the video information to be played and the object identifier. The server, starting from the first character of the object identifier, sequentially determines the identifier corresponding to each character in the object identifier, and arranges the identifiers corresponding to each character according to a preset format, thereby mapping the object identifier to an identifier sequence.

[0148] Then, the server performs subtitle offset processing on the video segments under the segment information corresponding to each identifier in the identifier sequence, based on the subtitle offset type, to obtain the subtitle offset segment. Alternatively, the server can directly obtain the pre-processed subtitle offset segment stored on the storage medium.

[0149] After obtaining the subtitle offset segments from each segment, the server concatenates these segments to obtain the marked video. The server then returns the generated marked video to the terminal, allowing the terminal to play the marked video to an object.

[0150] Because the generated tagged video indirectly embeds the object identifier in the form of subtitle offset, it is convenient to extract the object identifier by detecting the subtitles in the leaked video in the future, thereby tracing the source of the video leak.

[0151] This application also provides an application scenario in which the above-described method for generating marked videos is applied. Specifically, the method for generating marked videos in this scenario is applied as follows: When a target user browses a video list on a video platform, they select a video they wish to watch via a terminal. In response to the target user's selection, the terminal sends a video playback request to the server. The server extracts video information, such as the video name, contained in the playback request and locates the corresponding original video in its database. Simultaneously, the server extracts the object identifier carried in the playback request and, based on the identifier sequence obtained from the object identifier, offsets the subtitles in the original video to obtain the marked video. The server returns the marked video to the terminal, which then plays the marked video to the target user. The video can be a pre-stored complete video, such as entertainment videos, educational videos, TV dramas, or short videos, without limitation. However, it should be noted that the video length should be at least sufficient to embed subtitles with a single identifier sequence.

[0152] Based on the same inventive concept, embodiments of this application also provide a method for detecting video tags. In some embodiments, such as Figure 4 As shown, a method for detecting video tags is provided. The method is illustrated using a computer device as an example. Specifically, the computer device can be... Figure 1 The method for detecting video tags in a terminal or server includes the following steps:

[0153] Step S402: Obtain the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected.

[0154] Specifically, computer equipment can obtain the video to be tested by searching the Internet. The computer equipment then detects the subtitles in each frame of the video to determine the specific subtitle offset type corresponding to each frame. For example, the computer equipment can use tools such as FFMPEG to decode the video to obtain all its frames, and then determine the subtitle offset type for each frame.

[0155] Step S404: Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information.

[0156] Specifically, the computer device stores pre-set segmentation information. Based on the time period indicated by each segmentation information, the computer device determines the video segments in the video to be detected that correspond to each segmentation information. For example, the computer device can segment the video to be detected according to the segmentation information to obtain video segments corresponding to each segmentation information. For instance, the video segments from second 1 to second 10 of the video to be detected can be extracted as the video segments corresponding to the segmentation information.

[0157] Step S406: For each video segment, determine the identifier corresponding to the video segment based on the subtitle offset type to which the multiple video frames included in the video segment belong.

[0158] Each video segment resulting from the segmentation comprises multiple video frames. Specifically, for a given video segment, the computer device identifies all the video frames contained within that segment and determines the subtitle offset type corresponding to each video frame. The computer device then performs statistical analysis on the subtitle offset types corresponding to each video frame and determines an identifier corresponding to that video segment based on the statistical results.

[0159] In some embodiments, for each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong, including: for each video segment, determining the number of video frames in the video segment that correspond to each subtitle offset type, and using the identifier corresponding to the subtitle offset type with the most frames as the identifier corresponding to the video segment.

[0160] Specifically, for a video segment, the computer device determines the subtitle offset type corresponding to all the video frames it contains, and counts the number of video frames under each subtitle offset type. The computer device then determines the subtitle offset type with the most frames and uses the identifier corresponding to that subtitle offset type as the identifier corresponding to that video segment.

[0161] For example, a video clip from second 1 to second 10 contains 20 video frames, each corresponding to either type A or type B subtitle offset. The server determines that 16 video frames correspond to type A and 4 to type B. Therefore, the server determines that type A corresponds to the maximum number of subtitle offset frames. Since type A corresponds to identifier "1", the server assigns identifier "1" as the identifier for the video clip from second 1 to second 10.

[0162] In the above embodiments, by statistically analyzing the number of frames with different subtitle offset types to determine the identifier corresponding to the video segment, the position detection of video frames has a certain fault tolerance rate and improves the accuracy of video tagging.

[0163] Step S408: Based on the identifiers corresponding to each video segment in the video to be detected, determine the identifier sequence, and determine the object identifier marked in the video to be detected based on the identifier sequence.

[0164] Specifically, for each video segment of the video to be detected, the computer device performs the above processing to obtain the identifier corresponding to each video segment. Based on the multiple identifiers corresponding to all video segments, the computer device extracts an identifier sequence from the multiple identifiers, and obtains the object identifier marked in the form of subtitle offset in the video to be detected by performing an inverse mapping on the identifier sequence.

[0165] For example, in chronological order, the computer device determines that all video segments of the video to be detected correspond to multiple identifiers as "100101010010101001010...". Since the identifier sequence has a preset length and contains a preset number of identifiers, the computer device extracts the fixed-length, recurring identifier "1001010" and determines it as the identifier sequence. Then, it performs a reverse mapping using a preset mapping rule to obtain the object identifier corresponding to the identifier sequence.

[0166] As mentioned earlier, the identifier sequence can also include start and end identifiers. Therefore, when extracting the identifier sequence, the computer device can make judgments based on the start and end identifiers. For example, if the computer device determines that all video segments of the video to be detected correspond to multiple identifiers as "@ABAA@ABAA@ABAA…", then the computer device can extract the identifier "@ABAA" and determine it as the identifier sequence, and then obtain the object identifier according to a preset mapping rule.

[0167] Therefore, since the object identifier is indirectly embedded in the generated tagged video in the form of subtitle offset, when the video is detected, the embedded object identifier can be extracted by detecting the subtitle offset, thereby determining the party that leaked the video (or the party that requested playback), and realizing the tracing of the video tag.

[0168] In the above-mentioned video tag detection method, the subtitle offset type corresponding to each video frame in the video to be detected is determined, and the video segments corresponding to each segment information in the video to be detected are determined according to the preset segment information. For each video segment, based on the subtitle offset type to which the multiple video frames included in the video segment belong, an identifier corresponding to the video segment is determined. Then, based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined. Thus, the object identifier marked in the video to be detected is obtained by inverse mapping based on the identifier sequence. This realizes the detection of watermark information indirectly embedded in the marked video through subtitle offset, and can determine the video player or leakage source based on the marked object identifier, ensuring the detection rate and accuracy during source tracing.

[0169] In some embodiments, determining the subtitle offset type corresponding to each video frame in the video to be detected includes: obtaining the original video corresponding to the video to be detected; and, under the same video frame dimension, determining the subtitle offset type corresponding to each video frame in the video to be detected based on the positional relationship between the subtitle to be detected in each video frame in the video to be detected and the original subtitle in the corresponding video frame in the original video.

[0170] Specifically, computer devices can retrieve the original video corresponding to the video to be detected from a copyright database. For example, a computer device can use video fingerprinting technology to search within a copyright video database to find the corresponding original copyrighted video. Video fingerprinting technology is a technique that uses computer vision, audio processing, and other technologies to reduce the dimensionality of video content into vectors, and can be used in scenarios such as video retrieval, video deduplication, and video recommendation.

[0171] Since the transmitted video to be tested may have been edited, scaled, stretched, etc., in order to ensure the accuracy of the detection, the computer equipment detects the position of the subtitle to be tested in each video frame of the video to be tested and the position of the original subtitle in the corresponding video frame of the original video under the same video frame dimension. Based on the positional relationship between the subtitle to be tested and the original subtitle, the subtitle offset type corresponding to each video frame in the video to be tested is determined.

[0172] For example, a computer device uses OCR technology to detect subtitles in video frames and determines the specific offset of the subtitle to be detected relative to the original subtitle. For instance, if the computer device detects that the subtitle to be detected has moved upward by 1 pixel compared to the original subtitle, then the computer device determines that the subtitle offset type corresponding to that video frame is type A; or, if the computer device detects that the subtitle to be detected has moved downward by 1 pixel compared to the original subtitle, then the computer device determines that the subtitle offset type corresponding to that video frame is type B. OCR (Optical Character Recognition) refers to the process of analyzing and recognizing textual materials in image files to obtain text and layout information. Therefore, by detecting the subtitle position in each video frame, the subtitle offset type corresponding to each video frame can be obtained.

[0173] The video frame dimension includes both temporal and spatial dimensions. Accordingly, after obtaining the original video corresponding to the video to be detected, the method further includes aligning the video to be detected and the original video in both the temporal and spatial dimensions. Specifically, the computer device determines the video frames of the video to be detected and the original video corresponding to the same time along the same time axis for temporal alignment. Furthermore, the computer device also aligns the video frames of the video to be detected and the original video in the spatial dimension according to the same pixel coordinate system, for example, with the top-left corner as the origin.

[0174] In the above embodiments, by aligning the video to be detected and the original video, it is ensured that the video to be detected and the original video are in the same video frame dimension, thus avoiding errors in the position of the detected subtitles.

[0175] As mentioned above, the identifier sequence may also include start and end identifiers. Accordingly, in some embodiments, determining the identifier sequence based on the identifiers corresponding to each video segment in the video to be detected includes: determining the start and end identifiers from the identifiers corresponding to each video segment in the video to be detected, and extracting multiple identifiers between two adjacent start and end identifiers into an identifier sequence.

[0176] Specifically, the computer device extracts multiple identifiers between two adjacent start and end identifiers from each of the multiple identifiers corresponding to all video segments of the video to be detected. For example, if the computer device determines that each of the multiple identifiers corresponding to all video segments of the video to be detected is "@AB@@AB…", then the computer device can extract the identifier "AB" from it and determine it as an identifier sequence.

[0177] In the above embodiments, by setting start and end identifiers and grouping multiple subtitle offset segments corresponding to a single identifier sequence into a single tagging unit, the server can quickly extract the identifier sequence based on the start and end identifiers when detecting the video, thereby quickly obtaining the object identifier.

[0178] In a specific example, the video tag detection method provided in this application includes the following steps: A server obtains the video to be detected, either reported by a terminal or obtained through other means, and searches for the original video corresponding to the video to be detected in a video copyright library. A computer device uses a decoder to split both the video to be detected and the original video into multiple video frames, and compares the subtitle positions in each video frame of the video to be detected with the subtitle positions in each corresponding video frame of the original video. Based on the positional relationship between the two, the specific subtitle offset type corresponding to each video frame is determined. Then, for each video segment, the computer device counts the subtitle offset types corresponding to each video frame under that video segment. For example, if type A is the most common, the subtitle offset type corresponding to that video segment is determined to be type A, and thus the identifier corresponding to that video segment is determined to be A or an identifier corresponding to A. Based on the multiple identifiers corresponding to all video segments of the video to be detected, the computer device extracts an identifier sequence and converts the identifier sequence into an object identifier, thereby determining the source of the video leak.

[0179] This application also provides an application scenario in which the above-described video tag detection method is applied. Specifically, the video tag detection method is applied in this scenario as follows: When a terminal plays a tagged video, the target object may perform operations such as video recording, video caching, video downloading, and video forwarding through the terminal. Since the object identifier is indirectly embedded in the tagged video through subtitle offset, the video recorded, cached, downloaded, and forwarded by the object also contains the object identifier.

[0180] Therefore, when a computer device obtains a leaked video, such as when it finds a video already in a copyright library online, the computer device can detect the video and extract the identifier sequence based on the positional relationship between the subtitles and the original subtitles in the copyright library, and convert it into an object identifier. This allows the source of the leaked video on the internet to be determined.

[0181] To better understand this application, a specific product example is provided below. Figure 5A and Figure 5B As shown, in a specific example, the video tag generation method in this embodiment can be executed and implemented by a subtitle watermark embedding system, and the video tag detection method can be executed and implemented by a subtitle watermark detection system. The subtitle watermark embedding system and the subtitle watermark detection system can be integrated on a single server or set up separately on different servers.

[0182] To simplify the explanation, we will use two types of subtitle offsets, A and B, as examples. The video (or video clip) corresponding to type A will be called the A-side video (or video clip). Similarly, the video (or video clip) corresponding to type B will be called the B-side video (or video clip).

[0183] The subtitle watermarking embedding system mainly includes two functions / modules: A / B side watermark embedding and A / B side watermarked video distribution. Subtitle watermark embedding is deeply integrated into the video production process (video production, encoding / decoding, subtitle compression, video distribution, etc.). The core logic of the subtitle watermark embedding system is as follows: by modifying the subtitles, A / B type subtitle watermarks are embedded into the input video (i.e., the original video mentioned in the aforementioned embodiments); then, through video distribution, several segments are selected from each of the A / B side videos and combined, thereby sending different watermarked videos (i.e., the marked videos mentioned in the aforementioned embodiments) to different objects. For example: for object 1, a marked video with the combination AAAABBBB (each character A or B represents a subtitle offset segment) is distributed; for object 2, a marked video with the combination BBBBAAAA is distributed.

[0184] For example, the process of embedding watermarks on both sides of a video (A / B) can be as follows: Figure 6 As shown, the server acquires subtitle files (such as SRT format soft subtitles) and slightly modifies their positions during subtitle compression to embed watermarks on both sides (A and B). For example, side A of the video has all subtitles shifted upwards by N pixels, while side B has all subtitles shifted downwards by N pixels. Specifically, the server offsets the acquired subtitle files to obtain type A and type B subtitles. Simultaneously, the server decodes the input video using a decoder to obtain video frames, adds subtitles (type A and type B subtitles respectively) to the video frames, and re-encodes them using an encoder to generate watermarked videos on both sides (A and B). The offset subtitles then become the video watermarks. Subtitle compression can be achieved using the decoder and encoder in FFMPEG, while subtitle position adjustment is achieved by modifying the subtitle source file.

[0185] For example, the distribution process for A / B side watermarked videos can be as follows: Figure 7 As shown. When an object requests video playback on a terminal (or a video playback client installed on the terminal), the server receives the video playback request sent by the terminal, obtains the object identifier contained therein, and maps the object identifier into an identifier sequence corresponding to the subtitle offset type (A / B) according to a preset mapping rule. For example, the user ID "abc" is mapped to AAAA AAABAABA or 000000010010, where AAAA (or 0000) represents "a", AAAB (or 0001) represents "b", and AABA (or 0010) represents "c". For the generated A-side watermarked video and B-side watermarked video, the server segments them according to the preset segmentation rules to obtain subtitle offset segments of type A and type B. Combining the identifier sequence and the pre-generated A / B type subtitle offset segments, the server can generate a complete marked video for the server to send to the terminal for playback. For example, the server generates a specified M3U8 file for the terminal to play using appropriate tools. An M3U8 file is essentially a playlist / sequence, which may be a media playlist or a master playlist. Regardless of the type, the text within the playlist is encoded in UTF-8. When an M3U8 file is used as a media playlist, it records a series of media segments; playing these segments sequentially displays the complete multimedia resource.

[0186] Thus, the server completes the subtitle watermark embedding of the original video, and by making slight modifications to the subtitles, embeds the object identifier into the video, which facilitates subsequent source tracing.

[0187] The subtitle watermark detection system mainly includes a video alignment module, a subtitle position detection module, and an A / B sequence mapping module. The core logic of the subtitle watermark detection system is as follows: The video to be detected is searched in the copyright database using video fingerprinting technology to find the corresponding copyrighted original video, i.e., the original video without subtitle watermarks and with normally compressed subtitles. Then, the video to be detected is aligned with the original video in both time and spatial dimensions. Time-dimensional alignment refers to alignment along the timeline. For example, the video alignment process can be as follows: Figure 8 As shown, the server first uses video fingerprinting technology to retrieve the corresponding original video from the video copyright library, and then aligns the original video in the time dimension (i.e., on the timeline). Since the video to be detected and the original video are in the same time dimension, the server aligns them in the spatial dimension to obtain the alignment result.

[0188] After alignment, the server sends the same frame from the video to be detected and the original video into the OCR module for subtitle detection. It then compares the positional relationship of the detection boxes (top / bottom / unchanged) to determine the positional relationship between the offset subtitle and the original subtitle. Based on this positional relationship, the server can determine whether the frame contains a subtitle watermark and the corresponding subtitle offset type. For example, the subtitle position detection process can be as follows: Figure 9 As shown, combining the alignment results obtained in the previous step, the server inputs the aligned video to be detected and the original video frames from the same time (i.e., the video frame to be detected and the original video frame) into the subtitle position detection module (i.e., the OCR module in the figure) for subtitle position detection. By comparing the detected subtitle position information (i.e., the subtitle position of the frame to be detected and the subtitle position of the original frame), it can be determined whether the subtitle offset type of the video frame is type A or type B. The final subtitle position detection output is the detection result for each video frame of the video to be detected. For example, if the video to be detected has 1000 frames, the server outputs a 1000-bit detection frame sequence ABAAA...

[0189] For example, the process of A / B sequence mapping can be as follows: Figure 10 As shown, the A / B sequence mapping module maps the detection results to the final detection result (i.e., the object ID). Specifically, the server divides the detected frame sequence ABAAA... obtained in the previous step into segments using the same time unit as the original video, and performs segment voting to obtain an A / B sequence composed of characters A and B. Then, the A / B sequence is reverse-mapped to obtain the object identifier. For example, when a video is distributed, the specific time unit for segmentation is 10 seconds per segment. Therefore, the server divides the detection results into 10-second units, and each segment is voted on (if there are more A frames, the detection result of that segment is A; otherwise, it is B). This results in a sequence composed of A / B frames, which is then reverse-mapped according to a preset mapping rule to obtain the detection result of the A / B sequence mapping module, i.e., the object identifier.

[0190] Therefore, by detecting the position of subtitles and extracting object identifiers from marked videos, the process of playing, disseminating, and leaking marked videos can be traced, which is beneficial to the protection of video copyright.

[0191] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0192] Based on the same inventive concept, this application also provides a labeled video generation apparatus for implementing the labeled video generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more of the labeled video generation apparatus embodiments provided below can be found in the limitations of the labeled video generation method described above, and will not be repeated here.

[0193] In some embodiments, such as Figure 11 As shown, a device for generating marked videos is provided. This device can be a software module, a hardware module, or a combination of both integrated into a computer device. Specifically, the device includes: an acquisition module 1101, a determination module 1102, an offset module 1103, and a generation module 1104, wherein:

[0194] The acquisition module 1101 is used to acquire the object identifier and map the object identifier to an identifier sequence.

[0195] The determination module 1102 is used to determine the subtitle offset type corresponding to each identifier in the identifier sequence, and to determine the segment information corresponding to each identifier.

[0196] The offset module 1103 is used to obtain, for each identifier, the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by subtitle offsetting the original video segments with the same segment information in the original video.

[0197] The generation module 1104 is used to generate a marked video that corresponds to the original video and the object identifier based on the subtitle offset segment.

[0198] In some embodiments, the object identifier consists of multiple characters, and the acquisition module is further configured to determine the identifier corresponding to each character in sequence, starting from the first and second characters; and arrange the identifiers corresponding to each character according to a preset format to obtain an identifier sequence of a preset length.

[0199] In some embodiments, the acquisition module is further configured to arrange the identifiers corresponding to each character according to a preset format. If the number of identifiers corresponding to all characters in the object identifier is less than the preset number, the identifiers are padded at the end of the arrangement by a pre-set padding identifier to obtain an identifier sequence of a preset length.

[0200] In some embodiments, the offset module is further configured to acquire the original video and perform segmentation processing on the original video to obtain multiple original video segments; determine the original video segment pointed to by the segmentation information corresponding to each identifier in the identifier sequence; and perform subtitle offset on the original video segments corresponding to each identifier according to the subtitle offset type corresponding to the corresponding identifier to obtain the subtitle offset segment corresponding to the corresponding identifier.

[0201] In some embodiments, the offset module is further configured to obtain pre-processed subtitle offset videos corresponding to each subtitle offset type, wherein the subtitle offset videos are obtained by performing corresponding subtitle offset processing on the original video; for each identifier in the identifier sequence, according to the segment information corresponding to the corresponding identifier, subtitle offset segments with the same segment information are extracted from the subtitle offset videos belonging to the subtitle offset type corresponding to the corresponding identifier.

[0202] In some embodiments, the offset module is further configured to decode the original video to obtain multiple video frames, and add offset subtitles of the same subtitle offset type to each frame to obtain multiple video frames of the corresponding subtitle offset type; and encode the multiple video frames with offset subtitles of the corresponding subtitle offset type to obtain a subtitle offset video corresponding to the subtitle offset type.

[0203] In some embodiments, the generation module is further configured to splice the subtitle offset segments corresponding to each segment information according to the time sequence of each segment information to generate a marked video that corresponds to the original video and the object identifier.

[0204] In some embodiments, the tagged video includes multiple tagged units, each tagged unit including multiple subtitle offset segments corresponding to a single identifier sequence, and the multiple tagged units are distinguished by the video segments corresponding to the start and end identifiers in the identifier sequence, the start and end identifiers being set at at least one position at the beginning and end of the identifier sequence.

[0205] Specific limitations regarding the device for generating labeled videos can be found in the limitations on the method for generating labeled videos described above, and will not be repeated here. Each module in the aforementioned device for generating labeled videos can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the operations corresponding to each module.

[0206] Based on the same inventive concept, this application also provides a video tag detection apparatus for implementing the video tag detection method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more video tag detection apparatus embodiments provided below can be found in the limitations of the video tag detection method described above, and will not be repeated here.

[0207] In some embodiments, such as Figure 12 As shown, a video marker detection device is provided. This device can be a software module, a hardware module, or a combination of both integrated into a computer device. Specifically, the device includes: an acquisition module 1201 and a determination module 1202, wherein:

[0208] The acquisition module 1201 is used to acquire the video to be detected and determine the subtitle offset type corresponding to each video frame in the video to be detected.

[0209] The determination module 1202 is used to determine the video segments in the video to be detected that correspond to each segment information according to the preset segment information.

[0210] The determination module 1202 is also used to determine, for each video segment, an identifier corresponding to the video segment based on the subtitle offset type to which the multiple video frames included in the video segment belong.

[0211] The determination module 1202 is also used to determine the identifier sequence based on the identifiers corresponding to each video segment in the video to be detected, and to determine the object identifier marked in the video to be detected based on the identifier sequence.

[0212] In some embodiments, the acquisition module is further configured to acquire the original video corresponding to the video to be detected; and, under the same video frame dimension, determine the subtitle offset type corresponding to each video frame in the video to be detected based on the positional relationship between the subtitle to be detected in each video frame in the video to be detected and the original subtitle in the corresponding video frame in the original video.

[0213] In some embodiments, the video frame dimensions include a temporal dimension and a spatial dimension. The apparatus further includes an alignment module for aligning the video to be detected with the original video in the temporal dimension and the spatial dimension, respectively.

[0214] In some embodiments, the determining module is further configured to, for each video segment, determine the number of video frames in the video segment corresponding to each subtitle offset type, and use the identifier corresponding to the subtitle offset type with the most frames as the identifier corresponding to the video segment.

[0215] In some embodiments, the determining module is further configured to determine the start and end identifiers from the identifiers corresponding to each video segment in the video to be detected, and extract multiple identifiers between two adjacent start and end identifiers into an identifier sequence.

[0216] Specific limitations regarding the video marker detection device can be found in the limitations of the video marker detection method described above, and will not be repeated here. Each module in the aforementioned video marker detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the corresponding operations of each module.

[0217] In some embodiments, a computer device is provided, which may be a terminal or a server, and its internal structure diagram may be as follows: Figure 13 As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface connected via a system bus. The processor, memory, and I / O interfaces are connected via the system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores video data and / or caption data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a video tag detection method.

[0218] Those skilled in the art will understand that Figure 13The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0219] In some embodiments, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the embodiments corresponding to the above-described method for generating marked videos, or to implement the steps in the embodiments corresponding to the above-described method for detecting video marks.

[0220] In some embodiments, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the embodiments corresponding to the above-described method for generating marked videos, or implements the steps in the embodiments corresponding to the above-described method for detecting video markers.

[0221] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the embodiments corresponding to the above-described method for generating marked videos, or implements the steps in the embodiments corresponding to the above-described method for detecting video marks.

[0222] It should be noted that the object information (including but not limited to account information, ID, coding information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the object or fully authorized by all parties.

[0223] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0224] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0225] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating labeled videos, characterized in that, The method includes: Obtain the object identifier, which consists of multiple characters; Starting from the first character among the plurality of characters, determine the identifier corresponding to each character in sequence; The identifiers corresponding to each character are arranged according to a preset format. If the number of identifiers corresponding to all characters in the object identifier is less than the preset number, the identifiers are padded at the end of the arrangement by a pre-set padding identifier to obtain an identifier sequence of a preset length. Determine the subtitle offset type corresponding to each identifier in the identifier sequence, and determine the segment information corresponding to each identifier; For each identifier, obtain the subtitle offset segment pointed to by the segment information corresponding to the identifier, which belongs to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by performing subtitle offset on the original video segments with the same segment information in the original video; Based on the subtitle offset segment, a tagged video corresponding to the original video and the object identifier is generated.

2. The method according to claim 1, characterized in that, For each identifier, obtaining the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier includes: The original video is acquired and then segmented to obtain multiple original video clips. Determine the original video segment pointed to by the segment information corresponding to each identifier in the identifier sequence; For each original video segment corresponding to an identifier, subtitle offset is performed according to the subtitle offset type corresponding to the corresponding identifier to obtain the subtitle offset segment corresponding to the corresponding identifier.

3. The method according to claim 1, characterized in that, For each identifier, obtaining the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier includes: Obtain the pre-processed subtitle offset video corresponding to each subtitle offset type, wherein the subtitle offset video is obtained by performing corresponding subtitle offset processing on the original video; For each identifier in the identifier sequence, according to the segment information corresponding to the corresponding identifier, extract the subtitle offset segments with the same segment information from the subtitle offset video belonging to the subtitle offset type corresponding to the corresponding identifier.

4. The method according to claim 3, characterized in that, The subtitle offset video corresponding to each of the various subtitle offset types is generated through the following steps: The original video is decoded to obtain multiple video frames, and offset subtitles of the same subtitle offset type are added to each frame to obtain multiple video frames of the corresponding subtitle offset type. Encode multiple video frames with offset subtitles of the corresponding subtitle offset type to obtain a subtitle offset video corresponding to the subtitle offset type.

5. The method according to claim 1, characterized in that, The step of generating a tagged video that corresponds to the original video and the object identifier based on the subtitle offset segment includes: According to the time sequence corresponding to each segment information, the subtitle offset segments corresponding to each segment information are spliced ​​together to generate a marked video that corresponds to the original video and the object identifier.

6. The method according to any one of claims 1 to 5, characterized in that, The tagged video includes multiple tagged units, each tagged unit includes multiple subtitle offset segments corresponding to a single identifier sequence, and the multiple tagged units are distinguished by the video segments corresponding to the start and end identifiers in the identifier sequence, the start and end identifiers being set at at least one position at the beginning and end of the identifier sequence.

7. A method for detecting video tags, characterized in that, The method includes: Obtain the video to be tested; Obtain the original video corresponding to the video to be detected; Under the same video frame dimension, based on the positional relationship between the subtitles to be detected in each video frame of the video to be detected and the original subtitles in the corresponding video frames of the original video, the subtitle offset type corresponding to each video frame of the video to be detected is determined; Based on the preset segmentation information, determine the video segments in the video to be detected that correspond to each segmentation information respectively; For each video segment, an identifier corresponding to the video segment is determined based on the subtitle offset type to which the multiple video frames included in the video segment belong; Based on the identifiers corresponding to each video segment in the video to be detected, an identifier sequence is determined, and based on the identifier sequence, the object identifier marked in the video to be detected is determined.

8. The method according to claim 7, characterized in that, The video frame dimensions include a temporal dimension and a spatial dimension. After obtaining the original video corresponding to the video to be detected, the method further includes: aligning the video to be detected with the original video in both the temporal and spatial dimensions.

9. The method according to claim 7, characterized in that, For each video segment, determining an identifier corresponding to the video segment based on the subtitle offset type to which the multiple video frames included in the video segment belong, includes: For each video segment, determine the number of video frames in the video segment that correspond to each subtitle offset type, and use the identifier corresponding to the subtitle offset type with the most frames as the identifier corresponding to the video segment.

10. The method according to any one of claims 7 to 9, characterized in that, The step of determining the identifier sequence based on the identifiers corresponding to each video segment in the video to be detected includes: The start and end identifiers are determined from the identifiers corresponding to each video segment in the video to be detected, and multiple identifiers between two adjacent start and end identifiers are extracted into an identifier sequence.

11. An apparatus for generating labeled videos, characterized in that, The device includes: The acquisition module is used to acquire an object identifier, which consists of multiple characters; starting from the first character, the identifier corresponding to each character is determined in sequence; the identifiers corresponding to each character are arranged according to a preset format; if the number of identifiers corresponding to all characters in the object identifier is less than a preset number, the identifiers are padded at the end of the arrangement by a pre-set padding identifier to obtain an identifier sequence of a preset length. The determination module is used to determine the subtitle offset type corresponding to each identifier in the identifier sequence, and to determine the segment information corresponding to each identifier. The offset module is used to obtain, for each identifier, the subtitle offset segment pointed to by the segment information corresponding to the identifier and belonging to the subtitle offset type corresponding to the corresponding identifier; the subtitle offset segment pointed to by the segment information is obtained by offsetting the subtitles of the original video segments with the same segment information in the original video. The generation module is used to generate a marked video that corresponds to the original video and the object identifier based on the subtitle offset segment.

12. The apparatus for generating marked video according to claim 11, characterized in that, The offset module is also used to acquire the original video and perform segmentation processing on the original video to obtain multiple original video segments; and to determine the original video segment pointed to by the segmentation information corresponding to each identifier in the identifier sequence. For each original video segment corresponding to an identifier, subtitle offset is performed according to the subtitle offset type corresponding to the corresponding identifier to obtain the subtitle offset segment corresponding to the corresponding identifier.

13. The apparatus for generating marked video according to claim 11, characterized in that, The offset module is also used to obtain pre-processed subtitle offset videos corresponding to each subtitle offset type, wherein the subtitle offset videos are obtained by performing corresponding subtitle offset processing on the original video; for each identifier in the identifier sequence, according to the segment information corresponding to the corresponding identifier, subtitle offset segments with the same segment information are extracted from the subtitle offset videos belonging to the subtitle offset type corresponding to the corresponding identifier.

14. The apparatus for generating marked video according to claim 13, characterized in that, The offset module is also used to decode the original video to obtain multiple video frames, and add offset subtitles of the same subtitle offset type to each frame to obtain multiple video frames of the corresponding subtitle offset type. Encode multiple video frames with offset subtitles of the corresponding subtitle offset type to obtain a subtitle offset video corresponding to the subtitle offset type.

15. The apparatus for generating marked video according to claim 11, characterized in that, The generation module is also used to splice the subtitle offset segments corresponding to each segment information according to the time sequence of each segment information to generate a marked video that corresponds to the original video and the object identifier.

16. The apparatus for generating marked video according to any one of claims 11 to 15, characterized in that, The tagged video includes multiple tagged units, each tagged unit includes multiple subtitle offset segments corresponding to a single identifier sequence, and the multiple tagged units are distinguished by the video segments corresponding to the start and end identifiers in the identifier sequence, the start and end identifiers being set at at least one position at the beginning and end of the identifier sequence.

17. A device for detecting video markers, characterized in that, The device includes: The acquisition module is used to acquire the video to be detected; acquire the original video corresponding to the video to be detected; and, under the same video frame dimension, determine the subtitle offset type corresponding to each video frame of the video to be detected based on the positional relationship between the subtitle to be detected in each video frame of the video to be detected and the original subtitle in the corresponding video frame of the original video. The determination module is used to determine the video segments in the video to be detected that correspond to each segment information according to the preset segment information; for each video segment, it determines the identifier corresponding to the video segment based on the subtitle offset type to which the multiple video frames included in the video segment belong; it determines the identifier sequence based on the identifiers corresponding to each video segment in the video to be detected, and determines the marked object identifier in the video to be detected based on the identifier sequence.

18. The video marker detection device according to claim 17, characterized in that, The video frame dimensions include a time dimension and a spatial dimension. The device also includes an alignment module for aligning the video to be detected with the original video in the time dimension and the spatial dimension, respectively.

19. The video marker detection device according to claim 17, characterized in that, The determining module is further configured to, for each video segment, determine the number of video frames in the video segment corresponding to each subtitle offset type, and use the identifier corresponding to the subtitle offset type with the most frames as the identifier corresponding to the video segment.

20. The video marker detection apparatus according to any one of claims 17 to 19, characterized in that, The determining module is further configured to determine the start and end identifiers from the identifiers corresponding to each video segment in the video to be detected, and extract multiple identifiers between two adjacent start and end identifiers into an identifier sequence.

21. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 10.

22. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.

23. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 10.