Short video infringement evidence collection method and device
By using semantic segmentation and multimodal fingerprinting technology, independently matchable fingerprint fragments of short videos are generated, and evidence is automatically collected across platforms. This solves the problem of infringement evidence collection in complex and varied cross-platform scenarios, and achieves efficient and complete evidence chain generation and infringement identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING POWER LAW INTELLIGENT TECH CO LTD
- Filing Date
- 2026-05-19
- Publication Date
- 2026-06-16
AI Technical Summary
Existing technologies cannot achieve fully automated infringement evidence collection for short videos across platforms and in complex, multi-variant scenarios. The robustness of the retrieval and matching process is insufficient, and the evidence collection process lacks unified timestamps and environmental metadata records, resulting in low retrieval recall rates, incomplete evidence chains, and difficulty in large-scale rights protection.
Original fingerprint fragments are generated through semantic segmentation, candidate videos are acquired across platforms and a unified feature space is constructed. Combined with multimodal fingerprint matching and automatic evidence collection, a complete chain of evidence is formed, including semantic segmentation, multimodal feature fusion, noise-resistant quantization and local sensitive hashing, generating fingerprint fragments that can be matched independently, and storing them with hash and trusted timestamps.
It significantly improves the robustness of matching in complex editing scenarios, enables accurate retrieval of infringing links across platforms and generation of structured case packages, ensures the integrity and traceability of evidence, and improves the accuracy of infringement identification and the efficiency of evidence collection.
Smart Images

Figure CN122223632A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and specifically to a method and apparatus for obtaining evidence of infringement in short videos. Background Technology
[0002] Currently, evidence collection for short video copyright infringement mainly relies on manual methods. Rights holders need to repeatedly search across multiple platforms, manually compare data, and manually take screenshots and record screens. The process for protecting rights for a single video can take 12-18 hours, making batch processing difficult. Existing technical solutions in the industry mainly fall into two categories: one is a combination of purely manual complaint tools that rely on the operator's experience to organize evidence; the other is a single-modal content recognition system deployed within the platform (such as image-aware hashing, audio fingerprinting, etc.). However, these systems typically only cover a single platform and have poor adaptability to complex edit variations.
[0003] Existing technologies suffer from the following drawbacks: In the retrieval and matching stage, the single-modal fingerprint algorithm lacks robustness to complex editing operations such as speed adjustment, cropping, occlusion, and background music replacement, resulting in low cross-platform retrieval recall. Furthermore, the evidence collection stage lacks unified timestamps, hash signatures, and environmental metadata records, making it difficult to form a complete chain of evidence. Therefore, it cannot achieve fully automated infringement evidence collection for short videos in cross-platform, complex, and multi-variant scenarios. Moreover, poor retrieval robustness, incomplete evidence chains, and fragmented processing procedures hinder large-scale rights protection. Summary of the Invention
[0004] This invention provides a method and apparatus for obtaining evidence of infringement in short videos, in order to solve the problem of the inability to achieve fully automated infringement evidence collection across the entire chain in cross-platform and complex multi-variant scenarios for short videos.
[0005] In a first aspect, the present invention provides a method for obtaining evidence of infringement in short videos, the method comprising: Acquire original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation; Candidate short videos are obtained from multiple short video platforms, and candidate fingerprint fragments with the same feature space as the original short video are generated for each candidate short video. Candidate fingerprint fragments are matched with original fingerprint fragments, and multiple similarity calculation indicators are combined to generate a case package containing multiple suspected infringing links; After automatically collecting evidence from suspected infringing links in the case package, evidence files are generated, and the evidence files are hashed and stored with a trusted timestamp to form a chain of evidence.
[0006] This invention provides a method for obtaining evidence of infringement in short videos. Through semantic segmentation and multimodal fingerprint construction, the video is decomposed into independently matchable fingerprint fragments, significantly improving the robustness of matching in complex editing scenarios such as editing, speed adjustment, and occlusion. By acquiring candidate videos across platforms and constructing fingerprints in a unified feature space, combined with multi-similarity fusion and automatic aggregation, accurate retrieval of cross-platform infringement links and generation of structured case packages are achieved, solving the efficiency problems of scattered retrieval across multiple platforms and manual classification. Through automatic evidence collection and hash timestamp storage, a complete, tamper-proof, and traceable chain of evidence is constructed, ensuring the validity of the evidence. This method automates the entire process from fingerprint construction and cross-platform retrieval to evidence fixation, improving the accuracy of infringement identification and the efficiency of evidence collection, and solving the problem of not being able to achieve fully automated infringement evidence collection for short videos in cross-platform and complex multi-variant scenarios.
[0007] In one optional implementation, the original short video is semantically segmented, and an original fingerprint fragment is generated based on the semantic segmentation, including: Original short videos are divided into multiple semantically continuous time segments by performing shot boundary detection through frame-level semantic segmentation. Each time segment is processed by variable speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate an original fingerprint segment that can be independently matched for the original short video.
[0008] In the above technical solution, the video is decomposed into multiple semantic segments through frame-level semantic segmentation, aligning the fingerprint unit with the semantic content of the video and improving the matching accuracy in local editing scenarios. Variable speed alignment is performed on each segment to eliminate interference from playback speed differences on timeline matching. Multimodal feature fusion is used to jointly encode visual, audio, and other multi-dimensional information, enhancing the noise resistance when a single modality is interfered with. Combined with noise-resistant quantization processing, the fingerprint is robust to changes such as brightness adjustment, image quality compression, and occlusion. Locality-sensitive hashing maps the fingerprint to an efficient retrieval identifier, achieving fast approximate matching in a large-scale fingerprint database. The fingerprint segments generated by this method have independent matching capabilities and maintain high recognition accuracy and retrieval efficiency even in complex editing scenarios.
[0009] In one optional implementation, each semantic segment undergoes variable-speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate an independently matchable original fingerprint segment for the original short video, including: Variable-speed alignment is performed on visual features, audio features, and subtitle semantic vectors within each time segment; The multimodal features aligned by variable speed are input into a bidirectional encoder network to generate a fused vector; The fused vector is subjected to noise-resistant quantization to generate discrete fingerprint codewords. The discrete fingerprint codewords are then mapped to hash buckets using the locality-sensitive hashing algorithm to generate original fingerprint fragments that can be independently matched for the original short video.
[0010] In the above technical solution, the video is accurately divided into multiple semantically continuous time segments by lens boundary detection, aligning the fingerprint unit with the natural semantic structure of the video and improving the matching accuracy in local editing scenarios. By performing variable-speed alignment on multi-dimensional features such as visual, audio, and subtitles, the interference of playback speed differences on timeline matching is eliminated, ensuring that temporal features remain consistent at the semantic level. The multimodal features after variable-speed alignment are input into a bidirectional encoder network, realizing deep interaction and fusion of different modal information and generating a more discriminative unified feature representation. The fusion vector is subjected to noise-resistant quantization processing, making the discrete fingerprint codewords tolerant to common transformations such as brightness adjustment, image quality compression, and local occlusion. By mapping the discrete fingerprint codewords to hash buckets through locality-sensitive hashing, an efficient approximate retrieval index structure is constructed, supporting fast positioning in large-scale fingerprint databases.
[0011] In one optional implementation, generating candidate fingerprint fragments with the same feature space as the original short video for each candidate short video includes: Based on preset keywords, preset topic tags, or preset account lists, the acquired candidate short videos are filtered to generate a record of candidate short videos to be analyzed. Extract the multimodal features of each candidate short video from the candidate short video record, and decompose each candidate short video into multiple semantic segments through frame-level semantic segmentation; Each semantic segment is processed by variable speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate candidate fingerprint segments.
[0012] In the above technical solution, a massive number of candidate videos are intelligently filtered through preset conditions, effectively focusing on high-value content and reducing data redundancy in subsequent processing. A multimodal feature extraction and semantic segmentation strategy completely consistent with the original video is adopted to ensure that the candidate fingerprint and the original fingerprint are in a unified feature space, laying the foundation for accurate comparison. By performing variable-speed alignment, multimodal fusion, and anti-interference processing on each semantic segment, the generated candidate fingerprint can effectively resist common editing interferences such as changes in playback speed, image quality compression, and partial occlusion. Even in complex scenarios, it can maintain a stable correspondence with the original fingerprint, thereby improving the accuracy and recall rate of cross-platform infringement identification.
[0013] In one optional implementation, candidate fingerprint fragments are matched with original fingerprint fragments, and multiple similarity calculation indicators are fused to generate a case package containing multiple suspected infringing links, including: Calculate the multimodal similarity between candidate fingerprint fragments and original fingerprint fragments; the multimodal similarity includes deep metric learning similarity, weighted cosine similarity, and rule confidence; By fusing deep metric learning similarity, weighted cosine similarity, and rule confidence using a preset fusion function, the probability of infringement of candidate short videos relative to original short videos is obtained. Variant types are determined based on the distribution of multimodal similarity across different modalities, screen layout patterns, and timeline variation characteristics. Links suspected of infringement with a probability of infringement exceeding a preset threshold are grouped and aggregated according to the original work they belong to, generating a case package containing multiple suspected infringing links, their infringement information, and variant types.
[0014] The above technical solution improves the accuracy and reliability of infringement determination by calculating and fusing multimodal similarity and rule confidence, comprehensively evaluating the multidimensional matching degree between candidate videos and original content; it also accurately identifies specific variant types by analyzing the distribution characteristics of similarity between modalities and the temporal changes of the images, providing a basis for subsequent differentiated processing; and it automatically aggregates high-probability infringing links according to the dimension of original works to form structured case packages, realizing the unified organization and management of cross-platform infringement clues and effectively improving the processing efficiency in batch rights protection scenarios.
[0015] In one alternative implementation, after generating a case package containing multiple suspected infringing links, the method further includes: Based on the forwarding relationships, secondary creation relationships, cross-platform synchronization relationships, account association relationships, and watermark information of each suspected infringing link across preset platforms, an infringement link graph is constructed with infringing videos and publishing accounts as nodes and infringement dissemination relationships as edges. The geographical distribution information and dissemination path length of each node are also statistically analyzed.
[0016] In one optional implementation, suspected infringing links in the case package are automatically collected for evidence, and the generated evidence files are hashed and stored with a trusted timestamp to form a chain of evidence, including: For each suspected infringing link in the case package, an automated browser component is used to access the page, perform screenshot and screen recording operations, collect page metadata, and generate evidence files; Record the browser fingerprint, collection node IP address and operating environment information during the evidence collection process, and generate an environment record list that can be used for verification; Hash calculations are performed on the evidence documents and environmental record list, and a trusted timestamp is obtained for evidence storage, forming an evidence chain containing hash values and evidence storage credentials.
[0017] In the above technical solution, an automated browser component is used to perform page access, screenshotting, screen recording, and metadata collection on suspected infringing links, thereby achieving fully automated execution of the evidence collection process; the evidence collection environment information is recorded simultaneously and an environment record list is generated, providing a traceable basis for subsequent verification of the authenticity and completeness of the evidence collection process; the evidence files and environment records are hashed and a trusted timestamp is attached for evidence storage, ensuring the integrity of the evidence content, the authority of the generation time, and the immutability afterward, thus forming a complete chain of evidence with technical verifiability.
[0018] In one optional implementation, the suspected infringing links in the case package are automatically collected for evidence, and the generated evidence files are hashed and stored with a trusted timestamp to form a chain of evidence. This also includes: The chain of evidence is scored for completeness, based on at least one of the following indicators: number of screenshots, screen recording duration, coverage of infringing segments, and degree of relevance of the authorization chain. If the integrity score is lower than the first threshold, a supplementary data collection task will be automatically triggered. If the integrity score is still lower than the second threshold after supplementary collection, the corresponding case package will be marked as requiring manual review.
[0019] In the above technical solution, by conducting multi-dimensional quantitative evaluation of the evidence chain, cases with insufficient evidence integrity are automatically identified and a supplementary collection process is triggered. If the requirements are still not met after supplementary collection, the case is promptly transferred to manual review to ensure that the evidence chain entering the subsequent procedures has sufficient credibility and completeness.
[0020] In one alternative implementation, the method further includes: The system receives case packages and corresponding evidence chains, automatically generates rights protection materials according to preset document template rules, and formats, packages, and submits the rights protection materials according to the interface specifications of the target platform.
[0021] In the above technical solution, based on case package and evidence chain information, rights protection materials adapted to different scenarios are automatically generated through preset template rules, and are uniformly formatted, packaged and automatically submitted according to the interface specifications of each platform. This realizes the standardization and automation of the rights protection document generation and delivery process, effectively reducing the operational complexity of cross-platform rights protection.
[0022] Secondly, the present invention provides a short video infringement evidence collection device, the device comprising: The original content acquisition and multimodal fingerprint construction module is used to acquire original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation. The cross-platform infringement retrieval and case aggregation module is used to obtain candidate short videos from multiple short video platforms and generate candidate fingerprint fragments with the same feature space as the original short video for each candidate short video; the candidate fingerprint fragments are matched with the original fingerprint fragments, and multiple similarity calculation indicators are integrated to generate a case package containing multiple suspected infringing links; The evidence collection and preservation module is used to automatically collect evidence from suspected infringing links in the case package, generate evidence files, and preserve the evidence files by hashing and using a trusted timestamp to form an evidence chain.
[0023] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the short video infringement evidence collection method of the first aspect or any corresponding embodiment described above.
[0024] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the short video infringement evidence collection method of the first aspect or any corresponding embodiment described above.
[0025] Fifthly, the present invention provides a computer program product, including computer instructions, which are used to cause a computer to execute the short video infringement evidence collection method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0026] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0027] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the first type of short video infringement evidence collection method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the second process of the short video infringement evidence collection method according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the third process of the short video infringement evidence collection method according to an embodiment of the present invention; Figure 5(a) is a flowchart illustrating the original video multimodal fingerprint construction and storage stage according to an embodiment of the present invention; Figure 5(b) is a flowchart illustrating the candidate short video acquisition and cross-platform infringement retrieval stages according to an embodiment of the present invention; Figure 5(c) is a flowchart illustrating the evidence collection and preservation stage according to an embodiment of the present invention; Figure 5(d) is a flowchart illustrating the generation and submission stage of rights protection documents according to an embodiment of the present invention; Figure 5(e) is a flowchart illustrating the revenue settlement, risk control, data governance, and disaster recovery stages according to an embodiment of the present invention. Figure 6 This is a schematic diagram of the multimodal fingerprint construction and infringement retrieval subsystem structure according to an embodiment of the present invention; Figure 7 This is a structural block diagram of a short video infringement evidence collection device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0030] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0031] As an optional application scenario of this invention, such as Figure 1 The diagram illustrates a short video infringement evidence collection system provided by this invention. This system is deployed in a computer environment including at least one server, a fingerprint database, an evidence storage database, and a revenue settlement system. The system includes an external environment and a cloud-based infringement evidence collection and automated rights protection platform. The external environment includes: A1: Rights holder's terminal / client, used to upload original videos and view the progress of rights protection.
[0032] A2: Operations personnel terminal, used to handle cases requiring manual review and manual submission.
[0033] A3: Short video platforms 1…N.
[0034] A4: Data service provider / content aggregation platform, providing candidate short video data sources.
[0035] A cloud-based platform for infringement evidence collection and automated rights protection, including: B1: Original Content Access and Authorization Management Module, which registers original videos and authorization certificates uploaded by rights holders.
[0036] B2: Multimodal feature extraction module, which extracts visual hashes, audio fingerprints, subtitle semantic vectors, motion trajectories, and watermark fingerprints from original videos.
[0037] B3: Fingerprint fragment construction and fusion module, which performs frame-level semantic segmentation, variable speed alignment, multimodal fusion, noise-resistant quantization and LSH (Locality-Sensitive Hashing) mapping to generate multimodal fingerprint fragments.
[0038] DB1: Original fingerprint database, storing multimodal fingerprint fragments and trusted source scores of original videos.
[0039] B4: Candidate Acquisition and Management Module, which acquires candidate short videos and their metadata from multiple platforms and data service providers.
[0040] DB2: Candidate fingerprint database, storing multimodal fingerprint fragments of candidate videos.
[0041] B5: Infringement retrieval and similarity fusion module, which pre-screens based on Local Sensitive Hash (LSH) and fuses deep metric learning similarity, weighted cosine similarity, and rule confidence to output infringement probability and variant type.
[0042] B6: The case package generation and link graph module aggregates multiple suspected links into a case package and constructs an infringement link graph, which generates a case package containing multiple links and their infringement information and variant types.
[0043] C1: Evidence collection module, which automatically accesses, screenshots, records screens, and collects metadata for suspected infringing links.
[0044] C2: Evidence storage and timestamp module, calculates hashes for evidence and completes timestamp storage through time synchronization service and blockchain / TEE.
[0045] DB3: Evidence repository / exhibition record, used to store evidence files, hash values, timestamps, and signature information.
[0046] C3: Evidence integrity scoring and automatic supplementary collection module, which scores the chain of evidence in a case and triggers supplementary collection or manual review when the score is insufficient.
[0047] C4: Document generation and rule engine module, which automatically generates documents such as warning letters, complaint letters, and lawsuits based on case packages, evidence scores, and rights information, and records rule hit logs.
[0048] C5: Platform adaptation and delivery pipeline module, which submits documents according to the interface of each platform or the manual process, and realizes parallel delivery and retry queue.
[0049] D1: Profit Settlement and Profit Sharing Module, which automatically settles the profits of all parties based on the case outcome and preset rules.
[0050] D2: Risk control and compliance module, which verifies abnormal claims, authorization chains and rights certificates to prevent frivolous lawsuits and abnormal fund flows.
[0051] E1: Metadata and Data Governance Module, which uniformly manages data models and lineages such as fingerprints, cases, evidence, documents, and revenue.
[0052] E2: Monitoring and metrics bus module, which summarizes metrics such as retrieval recall rate, evidence credibility, delivery success rate, and payout release rate.
[0053] E3: Active-active + cold standby deployment and disaster recovery module to ensure high availability and fault recovery of key components such as fingerprint database and evidence database.
[0054] According to an embodiment of the present invention, a method for obtaining evidence of infringement in short videos is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0055] This embodiment provides a method for obtaining evidence of copyright infringement in short videos, which can be used in the aforementioned short video copyright infringement evidence obtaining system. Figure 2 This is a flowchart of a short video infringement evidence collection method according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Obtain original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation.
[0056] Among them, original short videos refer to original video works that are uploaded by the rights holder (original creator) and have legal copyright.
[0057] Semantic segmentation refers to using technical means to cut a continuous original short video into several relatively independent, complete time segments with clear start and end times, according to the natural turning points of the video content (such as camera transitions) and semantic logic (such as changes in the storyline).
[0058] An original fingerprint fragment refers to a set of digital features generated through a series of technical means for each of the aforementioned semantic segments, used to uniquely identify the content of that fragment.
[0059] Specifically, this step belongs to the original video multimodal fingerprint construction and storage stage. The rights holder uploads the original short video V and authorization certificate information M to the platform through the client. After obtaining the original short video, the semantic segmentation and shot boundary detection module decomposes it into multiple semantically continuous time segments through frame-level semantic segmentation. The multimodal feature extraction module extracts multimodal features such as visual, audio, and subtitles for each segment. The multimodal fingerprint segment construction and fusion module performs variable speed alignment and deep fusion. After generating the fusion vector, it is processed by noise-resistant quantization and local sensitive hashing to finally form an original fingerprint segment that can be independently matched.
[0060] Step S202: Obtain candidate short videos from multiple short video platforms, and generate candidate fingerprint fragments for each candidate short video that share the same feature space as the original short video.
[0061] Candidate short videos refer to videos that the system scrapes or obtains from various short video platforms or data service providers for analysis. In other words, candidate short videos originate from third parties and are not uploaded by the rights holders themselves. They may involve copyright infringement or be irrelevant videos, requiring comparison before a determination can be made.
[0062] Specifically, this step belongs to the candidate short video acquisition and post-suspended library storage stage. After the candidate acquisition and management module acquires candidate short videos from multiple short video platforms, it adopts the same semantic segmentation, multimodal feature extraction and fingerprint generation process as the original short video to construct a candidate fingerprint segment in the same feature space as the original video for each candidate video, ensuring that subsequent comparisons are carried out under the same benchmark.
[0063] Step S203: Match the candidate fingerprint fragments with the original fingerprint fragments and integrate multiple similarity calculation indicators to generate a case package containing multiple suspected infringing links.
[0064] Among them, the multiple similarity calculation indicators refer to three core similarities used to comprehensively evaluate the degree of matching between candidate short videos and original short videos, including: 1. Deep Metrics for Learning Similarity: Definition: The semantic similarity score (typically ranging from 0 to 1) is output after the fusion vector of the candidate fingerprint fragment and the original fingerprint fragment is input into a pre-trained metric learning network (such as a Siamese network or a contrastive learning network).
[0065] Function: It captures the similarity of fingerprint fragments at the high-level semantic level through deep learning models, and has a good recognition ability for complex nonlinear transformations (such as content rearrangement and semantically preserved secondary creation).
[0066] 2. Weighted cosine similarity: Definition: The comprehensive score is obtained by calculating the cosine similarity of the feature subspaces of different modalities such as visual subvectors, audio subvectors, and subtitle semantic subvectors, and then summing them according to preset or learned weights.
[0067] Function: To measure the degree of matching between different modalities at the underlying feature level, and to reflect the differences in importance of different modalities in infringement determination through weight allocation.
[0068] 3. Rule confidence: Definition: Match confidence score calculated based on preset business rules. Business rules include: Time overlap ratio: The degree of overlap between candidate segments and original segments on the timeline.
[0069] Contextual similarity: The matching of scene features before and after the segment.
[0070] Watermark fingerprint matching: Whether it contains the same embedded watermark identifier.
[0071] Topic consistency: Whether it appears under the same or similar topic tags.
[0072] Function: To introduce prior business knowledge to supplement and verify the similarity calculated by the model, thereby improving the reliability and interpretability of the judgment results.
[0073] Specifically, the infringement retrieval and similarity fusion module matches candidate fingerprint fragments with original fingerprint fragments, integrates multiple indicators such as deep metric learning similarity, weighted cosine similarity, and rule confidence, comprehensively assesses the probability of infringement and determines the variant type, and finally automatically aggregates high-probability infringing links according to the original work dimension to generate a structured case package.
[0074] Step S204: After automatically collecting evidence from the suspected infringing links in the case package, generate evidence files, and store the evidence files by hashing and using a trusted timestamp to form an evidence chain.
[0075] Specifically, this step is the evidence collection and preservation process. The evidence collection module automatically accesses suspected infringing links in the case package, performs screenshots, screen recordings, and metadata collection to generate original evidence files; it simultaneously records browser fingerprints, IP addresses, and operating environment information during the evidence collection process to form an environment record list; it performs hash calculations on the evidence files and environment record list, and attaches a trusted timestamp for preservation through the preservation and timestamp module, generating an evidence chain containing hash values and preservation certificates to ensure the integrity of the evidence content, the authority of the generation time, and its verifiability after the fact.
[0076] The short video infringement evidence collection method provided in this embodiment organically connects the various stages of retrieval, evidence collection, rights protection and settlement by constructing and fusing multimodal fingerprint fragments, fusing cross-platform multidimensional similarity and case aggregation, and a score-driven evidence supplementation and document linkage mechanism. It solves the technical problems of "difficult retrieval, difficult evidence collection and difficult large-scale processing" in short video multi-variant infringement scenarios, and has significant technical effects in terms of matching robustness, evidence chain reliability and large-scale automated processing capabilities.
[0077] This embodiment provides a method for obtaining evidence of copyright infringement in short videos, which can be used in the aforementioned short video copyright infringement evidence obtaining system. Figure 3 This is a flowchart of a short video infringement evidence collection method according to an embodiment of the present invention, such as... Figure 3 As shown, the process includes the following steps: Step S301: Obtain original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation.
[0078] Specifically, as shown in Figure 5(a), in the stage of constructing and storing original video multimodal fingerprints, the above step S301 includes: Step S3011: Perform shot boundary detection on the original short video through frame-level semantic segmentation and divide it into multiple semantically continuous time segments.
[0079] This step involves semantic segmentation and shot boundary detection submodules performing frame-level analysis on video V based on fv and fs to detect shot transitions and divide the video into multiple semantically continuous time segments seg_i, each corresponding to a start and end time interval [ti_start, ti_end]. Simultaneously, statistical features of the preceding and following neighboring segments are calculated for each time segment to obtain the scene context code fc_i, which represents the scene environment in which the segment resides. Specifically, this includes: First, the rights holder uploads the original short video (V) and authorization certificate information (M) to the platform via the client. The platform then performs the following operations in the original content access and authorization management module: a) Generate the original content identifier ID_orig; b) Record the upload time t, upload account, data acquisition terminal type, file identifier of the authorization certificate, and authorization scope information; c) The above content is combined into an "Original Content Record" and temporarily stored in the Original Content Registration Form.
[0080] This step provides foundational data for subsequent fingerprint construction and trusted source scoring.
[0081] Next, multimodal feature extraction and semantic segmentation are performed, including: In the multimodal feature extraction module, video V is processed as follows: a) Visual feature extraction: Sampling of keyframes in the video and generating visual fingerprints using existing perceptual hashing or depth-based hashing algorithms, specifically including: First, keyframes are extracted from the video at time intervals, such as 1 to 3 frames per second. Then, the visual fingerprints of the keyframes are calculated using existing perceptual hashing algorithms or hash networks based on convolutional neural networks to obtain the visual hash feature sequence fv.
[0082] b) Audio fingerprint extraction: Decode the audio track, extract frequency domain features (such as Mel-frequency cepstral coefficients), and generate an audio fingerprint using existing audio fingerprinting algorithms, specifically including: First, the audio track in the video is extracted, decoded, and resampled. Then, frame-level frequency domain features, such as Mel-frequency Cepstral Coefficients (MFCC), are calculated, and an audio fingerprint sequence fa is generated using an existing audio fingerprinting algorithm.
[0083] c) Subtitle semantic vector extraction: Subtitles provided manually or automatically recognized by speech are segmented into sentences and encoded into a semantic vector sequence using a pre-trained language model. This includes: First, the manually provided subtitles or the transcribed text obtained through automatic speech recognition are segmented into sentences. Then, a pre-trained language model (such as a bidirectional encoder structure) is used to map each sentence of text into a semantic vector, forming a subtitle semantic vector sequence fs.
[0084] d) Inter-frame motion trajectory feature extraction: Extracting motion trajectory vectors of the main body region based on optical flow estimation or keypoint tracking methods, specifically including: First, optical flow or key point tracking algorithms are used to estimate the motion trajectory of the main region in the video. Then, the trajectory coordinates and their changes are encoded into a motion feature sequence fm.
[0085] e) Watermark Fingerprint Extraction: Perform watermark detection on video frames or audio signals to extract embedded watermark identifiers, specifically including: First, a digital watermark detection algorithm is executed on the video frame or audio signal; if a recognizable embedded watermark is detected, the watermark identifier fw is extracted.
[0086] The visual hashing, audio fingerprinting, subtitle semantic encoding, motion trajectory and watermark detection mentioned above can all be implemented using relevant technologies. This embodiment does not limit the specific algorithms.
[0087] Frame-level semantic segmentation and shot boundary detection, such as Figure 6 As shown in steps F1 to F2, the semantic segmentation and shot boundary detection module analyzes the video based on fv and fs: a) First, the camera switching position is identified through a scene change detection algorithm, and the video is divided into several coarse-grained shot segments; b) Within each shot segment, based on the similarity change of the subtitle semantic vector fs, it is further divided into semantically continuous sub-segments seg_i; then, the start and end times [ti_start,ti_end] are determined for each seg_i. c) Then calculate the statistical features of adjacent segments before and after seg_i (e.g., subject category, scene type, etc. of adjacent segments) and encode them as scene context features fc_i.
[0088] This step breaks down the originally continuous video into a structured sequence of time segments, with each segment corresponding to a relatively complete semantic unit, laying the foundation for subsequent segment-level fingerprint construction.
[0089] Step S3012: Perform variable speed alignment, multimodal feature fusion, noise-resistant quantization, and local sensitive hashing on each semantic segment to generate an original fingerprint segment that can be independently matched for the original short video.
[0090] In an optional implementation, step S3012 includes: Step a1: Perform variable-speed alignment on the visual features, audio features, and subtitle semantic vectors within each time segment.
[0091] Specifically, considering the potential for accelerated or decelerated playback of candidate short videos, dynamic time warping and other variable speed alignment algorithms are used for time series features such as fv, fa, and fm to align feature sequences with the same semantic content in the time dimension, thus solving the problem of inconsistent sequence lengths caused by different playback speeds.
[0092] Step a2: Input the variable-speed aligned multimodal features into the bidirectional encoder network to generate a fused vector.
[0093] Specifically, the aligned visual features, audio features, subtitle semantic vectors, and motion trajectory features are input into the bidirectional encoder network to obtain a fused vector fi in a unified embedding space. The bidirectional encoder can be implemented using existing structures such as bidirectional recurrent neural networks and bidirectional Transformers.
[0094] Step a3: Perform noise-resistant quantization on the fused vector to generate discrete fingerprint codewords, and map the discrete fingerprint codewords to hash buckets using the locality-sensitive hashing algorithm to generate original fingerprint fragments that can be independently matched for the original short video.
[0095] Specifically, noise-resistant quantization and locality-sensitive hashing: noise-resistant quantization is performed on the fused vector fi by encoding continuous values into discrete fingerprint codewords qi, which makes it robust to disturbances such as brightness, color, and slight occlusion; then, the locality-sensitive hashing algorithm is used to map qi to hash bucket IDs for subsequent approximate proximity retrieval in a large-scale fingerprint database.
[0096] Through the above processing, a complete video is represented as several "fingerprint fragments" Fi, each containing a time interval [ti_start,ti_end], a fusion vector fi, a quantized fingerprint qi, a hash bucket ID, and scene context features fc_i.
[0097] Steps a1 to a3 above belong to the multimodal fingerprint fragment construction and fusion stage, such as Figure 6 Steps F3 to F5, as shown, perform the following operations on each segment seg_i: a) Variable speed alignment. For time series features such as fv, fa, and fm, dynamic time warping or equivalent variable speed alignment algorithms are used to stretch / compress the feature sequences within a segment along the time axis. The alignment goal is to align feature points representing the same semantic content in the time dimension when candidate videos have acceleration, deceleration, or slight time shifts.
[0098] b) Multimodal encoding. Aligned visual features, audio features, caption semantic vectors, and motion trajectory features are concatenated in chronological order; then, they are input into a bidirectional encoder network (e.g., bidirectional RNN or bidirectional Transformer) to obtain a fused vector fi in a unified embedding space. This network can be trained offline based on historical labeled data.
[0099] c) Noise-resistant quantization and locality-sensitive hashing. The fused vector fi is quantized, mapping continuous values to a set of discrete codewords (e.g., through vector quantization or segmented threshold encoding) to obtain the quantized fingerprint qi. Then, in the quantization design, a fault tolerance range is reserved for disturbances such as brightness changes, color adjustments, partial occlusion, and picture-in-picture overlay. Then, the locality-sensitive hashing algorithm is used to map qi to one or more hash bucket IDs for subsequent approximate nearest neighbor retrieval.
[0100] Finally, a fingerprint fragment record Fi is generated for each seg_i, which includes at least: original content identifier ID_orig, fragment number i, time interval [ti_start,ti_end], fusion vector fi, quantized fingerprint qi, hash bucket ID, and scene context feature fc_i.
[0101] Through the above steps, the inventive point of this invention is achieved: by using frame-level semantic segmentation and multimodal variable-speed alignment, a complete video is decomposed into multiple independently matchable fingerprint segments. Furthermore, robust recognition capabilities against complex variations such as mirroring, occlusion, color correction, AI redrawing, and local editing are enhanced through noise-resistant quantization and locality-sensitive hashing. This invention represents a complete video as a set of independently matchable multimodal fingerprint segments, solving the problem that "a single fingerprint segment is not robust to local editing and multiple variations."
[0102] It should be noted that this embodiment also provides trusted source scoring and fingerprint database writing. The trusted source scoring submodule calculates the trusted source score (score_src) for each original content record based on the following factors: a) Trust level of the data collection terminal (e.g., certified device, ordinary device); b) The completeness and legality of the authorization certificate (e.g., whether there is complete copyright registration information); c) The publication time and upload path of the original work (e.g., whether it was uploaded directly from the rights holder's account).
[0103] d) The calculated score_src, along with the original content records and fingerprint fragment set {Fi}, is written into the original fingerprint database DB1. Each record in the fingerprint database is associated with its corresponding metadata to facilitate subsequent retrieval and evidence credibility assessment.
[0104] The credible source scoring submodule calculates a credible source score (score_src) for each original content record based on the trust level of the collection terminal, the completeness of the authorization credential chain, and the content publication time and upload path. For example, original videos collected by an authenticated terminal with complete authorization credentials receive a higher initial score; videos with unknown sources or incomplete authorization chains receive a lower score. Finally, the "original content record," the fingerprint fragment set {Fi}, and its corresponding score_src are written into the fingerprint database for subsequent retrieval and evidence credibility assessment.
[0105] Step S302: Obtain candidate short videos from multiple short video platforms, and generate candidate fingerprint fragments for each candidate short video that share the same feature space as the original short video.
[0106] Specifically, as shown in Figure 5(b), step S302 above includes: Step S3021: Based on preset keywords, preset topic tags, or preset account lists, filter the acquired candidate short videos to generate a record of candidate short videos to be analyzed.
[0107] Specifically, the candidate acquisition module communicates with open interfaces of multiple short video platforms, content aggregation service interfaces, and interfaces of cooperative data service providers to obtain metadata and access addresses of candidate short videos from multiple data sources, including short video platforms. Based on preset criteria such as keywords, topic tags, target account lists, and posting time windows, candidate videos are filtered to form "candidate video records," which are then written to the candidate database.
[0108] Furthermore, the candidate acquisition and management module periodically obtains candidate video information through the open interfaces of various short video platforms, content aggregation service interfaces, and data service provider interfaces, including: a) Video unique identifier, URL, title, description, hashtags, author account, and publication time; b) Basic metrics such as views, likes, and comments.
[0109] c) The system can filter out candidate videos that are highly relevant to the rights holder's work based on preset keywords, a list of key accounts, a list of topics, and a time window, write them into the candidate video record table, and store them in the candidate database.
[0110] Step S3022: Extract the multimodal features of each candidate short video in the candidate short video record, and decompose each candidate short video into multiple semantic segments through frame-level semantic segmentation.
[0111] Step S3023: Perform variable speed alignment, multimodal feature fusion, noise-resistant quantization, and local sensitive hashing on each semantic segment to generate candidate fingerprint segments.
[0112] Specifically, for each candidate video in the candidate library, the system calls the multimodal feature extraction and fingerprint fragment construction process in step S301 to obtain a set of candidate fingerprint fragments {F′j}, and writes it into the candidate fingerprint library to ensure that the original fingerprint and the candidate fingerprint can be directly compared in the same feature space.
[0113] Step S303 involves matching the candidate fingerprint fragments with the original fingerprint fragments and fusing multiple similarity calculation indicators to generate a case package containing multiple suspected infringing links. For details, please refer to [link to relevant documentation]. Figure 2 Step S203 of the illustrated embodiment will not be described again here.
[0114] Step S304 involves automatically collecting evidence from suspected infringing links in the case package, generating evidence files, and then hashing and storing these evidence files with a trusted timestamp to form a chain of evidence. For details, please refer to [link to relevant documentation]. Figure 2Step S204 of the illustrated embodiment will not be described again here.
[0115] The short video infringement evidence collection method provided in this embodiment is based on the multimodal fingerprint fragment construction and robust fusion method of frame-level semantic segmentation. It decomposes the complete video into multiple fingerprint fragments that can be matched independently, and supports robust recognition of complex variants such as mirroring, occlusion, color correction, AI redrawing and local editing through variable speed alignment, bidirectional encoding and noise-resistant quantization.
[0116] This embodiment provides a method for obtaining evidence of copyright infringement in short videos, which can be used in the aforementioned short video copyright infringement evidence obtaining system. Figure 4 This is a flowchart of a short video infringement evidence collection method according to an embodiment of the present invention, such as... Figure 4 As shown, the process includes the following steps: Step S401: Obtain the original short video, perform semantic segmentation on the original short video, and generate an original fingerprint fragment based on the semantic segmentation. For details, please refer to [link to relevant documentation]. Figure 3 Step S301 of the illustrated embodiment will not be described again here.
[0117] Step S402: Obtain candidate short videos from multiple short video platforms, and generate candidate fingerprint fragments for each candidate short video that share the same feature space as the original short video. For details, please refer to [link to details]. Figure 3 Step S302 of the illustrated embodiment will not be described again here.
[0118] Step S403: Match the candidate fingerprint fragments with the original fingerprint fragments and integrate multiple similarity calculation indicators to generate a case package containing multiple suspected infringing links.
[0119] Specifically, such as Figure 6 As shown in Figure 5(b), the Local Sensitive Hash (LSH) bucket candidate filtering module searches for original fragments Fi that fall into the same or adjacent hash buckets in the original fingerprint database DB_O based on the hash bucket ID of the candidate fingerprint fragment F′j, forming a candidate matching set, narrowing the subsequent calculation range, and improving retrieval efficiency. As shown in Figure 5(b), the above step S403 includes: Step S4031: Calculate the multimodal similarity between the candidate fingerprint fragment and the original fingerprint fragment; the multimodal similarity includes deep metric learning similarity, weighted cosine similarity, and rule confidence.
[0120] Specifically, the infringement retrieval and similarity fusion module calculates multiple similarities for fragment pairs (F′j,Fi) in the candidate set: a) Deep metric learning similarity sd: Input the fusion vector of candidate fragment F′j and each original fragment Fi in the candidate set into a pre-trained metric learning network (such as Siamese network or contrastive learning network), and output the semantic similarity in the range [0,1]. b) Weighted cosine similarity sc: Calculate the cosine similarity for visual sub-vectors, audio sub-vectors, and subtitle semantic sub-vectors respectively, and sum them according to the weights obtained by experience or learning to obtain the comprehensive cosine similarity sc. c) Rule confidence score sr: Each pair of candidate fingerprint fragments and original fingerprint fragments is scored based on rules such as the fragment time overlap ratio, context scene similarity, whether they contain the same watermark fingerprint, and whether they come from the same topic, and the rule confidence score sr is obtained.
[0121] Step S4032: The deep metric learning similarity, weighted cosine similarity and rule confidence are fused by a preset fusion function to obtain the infringement probability of the candidate short video relative to the original short video.
[0122] Specifically, the system combines sd, sc, and sr using a preset fusion function, such as linear weighting or fusion based on a small classification model, to obtain the infringement probability P_infringe between the candidate video and a certain original video.
[0123] Step S4033: Determine the variant type based on the distribution of multimodal similarity across different modalities, the layout pattern of the screen, and the timeline change characteristics.
[0124] Specifically, based on the distribution of similarity across different modalities, screen layout patterns, and timeline changes, variant types (such as pure copying, mirror flipping, picture-in-picture, editing and compositing, and secondary creation with added subtitles) are determined. Multimodal similarity is fused with rule confidence to adapt to the complex variants resulting from the superposition of multiple editing techniques in short video scenarios.
[0125] Step S4034: Group and aggregate suspected infringing links with an infringement probability higher than a preset threshold according to the dimension of the original work to which they belong, and generate a case package containing multiple suspected infringing links and their infringement information and variant types.
[0126] Specifically, the case engine aggregates candidate videos with an infringement probability exceeding a threshold. The system groups suspected infringing links according to their original content records, combining multiple suspected links of the same original work into a "case package." Each case package records information such as the infringement probability, variant type, platform information, and initial trusted source score for each link. This includes: For candidate videos with an infringement probability exceeding a preset threshold, the case package generation module aggregates them according to the dimension of original works: a) Group all suspected links pointing to the same ID_orig into the same case package; b) Record the infringement probability, variant type, platform information, and corresponding fingerprint matching fragment for each link in the case package; It should be noted that for frequently accessed original content, the search engine maintains a hot fingerprint cache, storing its fingerprint fragments and recent search results in memory to achieve millisecond-level return. For long-tail original content, asynchronous search tasks are used, and the estimated completion time is returned to the user interface. The system dynamically adjusts the hot fingerprint set based on access statistics and cache hit rate, implementing a "cold start + hot cache" search strategy that balances search recall and computational cost.
[0127] Step S404: Based on the forwarding relationship, secondary creation relationship, cross-platform synchronization relationship, account association relationship and watermark information of each suspected infringing link across preset platforms, construct an infringement link graph with infringing videos and publishing accounts as nodes and infringement dissemination relationship as edges, and count the geographical distribution information and dissemination path length of each node.
[0128] Specifically, the construction of the infringement chain graph includes: Within the case package, the link graph module comprehensively utilizes the forwarding / co-promotion relationships, account follows, embedded watermark information, and the publication time of each suspected link provided by the platform to construct an infringement link graph: a) Use suspected infringing videos and the accounts that posted them as nodes in the graph; b) Treat the relationships between forwarding, secondary creation, and cross-platform synchronization as graph edges; c) Calculate the geographical information and propagation path length of each node.
[0129] This graph structure provides intuitive data support for the development of subsequent strategies for handling bulk complaints and lawsuits.
[0130] Step S405: Automatically collect evidence from suspected infringing links in the case package to generate evidence files, and then hash and store the evidence files with a trusted timestamp to form an evidence chain.
[0131] Specifically, step S405 above also includes: Step S4051: For each suspected infringing link in the case package, access the page through an automated browser component, perform screenshot and screen recording operations, collect page metadata, and generate evidence files.
[0132] Specifically, after receiving the case package, the evidence collection module automatically accesses each suspected infringing link: a) Load the target page via a built-in browser component or automated script; b) During video playback, capture key screenshots containing infringing segments according to a preset strategy; c) Record the entire playback process or key sections; d) Collect page metadata such as page title, URL, publication time, number of views, number of likes, number of comments, etc.
[0133] The above data were compiled to form an evidentiary document.
[0134] Step S4052: Record the browser fingerprint, collection node IP address and operating environment information during the evidence collection process, and generate an environment record list that can be used for verification.
[0135] Specifically, the collection script automatically records browser fingerprints (such as User-Agent, resolution, and plugin information), collection node IP addresses, and key environment variables during execution, and generates a collection list to reconstruct the collection environment during court testimony.
[0136] Step S4053: Perform hash calculation on the evidence documents and environmental record list, and obtain a trusted timestamp for evidence storage to form an evidence chain containing hash values and evidence storage certificates.
[0137] Specifically, for each screenshot file, screen recording file, and metadata file in the evidence files, the system uses algorithms such as SHA-256 or SM3 to calculate a hash digest and writes the association relationship of "file identifier - hash value - generation time" into the evidence metadata table.
[0138] The evidence storage and timestamp module performs the following operations: a) Access the national time service or other trusted time sources to obtain an authoritative timestamp; b) Encapsulate the evidence file hash value, timestamp, collector's identity, etc., into an evidence storage request; c) Obtain evidence by storing the evidence through a blockchain network or a Trusted Execution Environment (TEE); d) Associate the evidence storage certificate with the evidence document identifier and write it into the evidence database DB3.
[0139] This step allows for the verification of evidence documents during subsequent examination sessions, proving that the documents existed at a specific time and were not tampered with.
[0140] Step S4054: The evidence chain is scored for completeness. The completeness score is based on at least one of the following indicators: number of screenshots, screen recording duration, coverage of infringing segments, and degree of association of the authorization chain. If the completeness score is lower than the first threshold, a supplementary collection task is automatically triggered. If the completeness score is still lower than the second threshold after supplementary collection, the corresponding case package is marked as requiring manual review.
[0141] Specifically, the evidence evaluation module calculates an evidence completeness score (score_evi) for each case package. The scoring model comprehensively considers features including: the depth of data collection jumps, the number of screenshots, screen recording duration, whether the main infringing segments are covered, whether the copyright statement information displayed on the platform is collected, and the degree of correlation with the original authorization chain. Based on the score result, if the score_evi is below a first threshold, a supplementary data collection task is automatically triggered (e.g., increasing the screenshot angle, extending the screen recording time, supplementing the comment page, etc.); if the score_evi is still low after supplementary data collection, the case package is marked as requiring manual review and pushed to the manual processing queue. The specific process includes: The evidence integrity scoring module calculates score_evi for each case package. The scoring model could consider: a) Whether all key infringing segments were captured; b) Whether the number of screenshots is sufficient to reflect the infringement; c) Does the screen recording duration cover the main infringing content? d) Does it include copyright information or statements displayed on the platform? e) Collect data on jump depth and authorization chain coverage, etc.
[0142] When score_evi falls below the first threshold, the system automatically initiates a supplementary data collection task, such as increasing the screenshot angle, extending the screen recording time, or collecting information from the comment section. If the score still does not reach the second threshold after the supplementary data collection, the case package is marked as "requires manual review" and pushed to the operations staff's terminal, whereby the staff decides whether to continue the rights protection efforts or adjust the strategy.
[0143] Step S406: Receive the case package and the corresponding chain of evidence, automatically generate rights protection materials according to the preset document template rules, and format, package and submit the rights protection materials according to the interface specifications of the target platform.
[0144] This step is part of the process for generating and submitting materials for rights protection, and specifically includes: 1. Automatic generation of rights protection documents: The document generation and rule engine module receives the following inputs: a) A list of suspected infringing links, probability of infringement, and variant types in the case package; b) List of evidence and score_evi; c) Copyright certificate information for original works (registration number, name of the right holder, scope of rights, etc.); d) Claims set by the rights holder (such as platform removal, demand for a certain amount of compensation, etc.).
[0145] The system automatically generates, based on predefined document templates and rules, the following: relevant letters (such as warning letters), complaint letters from various platforms and other relevant paths, evidence catalog and evidence description, and compensation calculation table (based on the number of infringements, number of views, platform revenue, etc.).
[0146] Different types of infringement and strength of evidence can trigger different combinations of templates and clauses. For example, for cases that are "suspected to be legitimate derivative works with weak evidence", only a communication letter with mild wording is generated, while for cases that are "highly suspected to be malicious plagiarism with sufficient evidence", a formal complaint text is generated.
[0147] 2. Platform adaptation and format encapsulation: The platform adaptation and submission pipeline module provides structured encapsulation of documents: for platforms providing APIs, it converts document content into the platform-required JSON or form parameter structures; for scenarios that only support manual uploads, it exports documents as standardized PDFs or structured JSON files for manual copying or uploading. Simultaneously, the module generates a unique task identifier and status record for each submission task.
[0148] 3. Decision log recording: During document generation and platform adaptation, the rules engine records decision logs, including: the relevant clauses cited and their triggering conditions, the basis for selecting the compensation range (e.g., playback volume range, evidence strength level, etc.), whether to adopt a strategy of consolidation, multiple defendants, or separate lawsuits, and the rule hit status of selecting a particular template from multiple templates. This log information is used for internal auditing and interpreting the basis for automated decision-making in related procedures.
[0149] 4. Document Submission and Retry Mechanism: The task queue scheduling module uses parallel threads to submit tasks in batches for tasks that support interfaces. If an interface fails or rate limiting occurs, the task is placed in a retry queue according to the set backoff strategy. Tasks that fail after multiple retries are marked as "requires manual submission" and pushed to the operations personnel's terminal.
[0150] The above process ensures that automatic submission can still be completed to the greatest extent possible even when the third-party platform interface is unstable or the call frequency is limited.
[0151] It should be noted that this embodiment also provides revenue settlement, risk control and data governance, and disaster recovery processes, specifically including: 1. Profit Settlement and Sharing: After a case is resolved, the revenue settlement module automatically calculates the revenue due to each party based on: the actual compensation amount (settlement amount, judgment amount, or platform payment amount), the pre-agreed profit-sharing rules (proportion of rights holder, agency, rights protection platform, etc.), and case costs (including estimated values of evidence collection costs, litigation costs, etc.), generates revenue flow records, and provides an interface to the financial system for invoicing and tax declaration.
[0152] 2. Risk Control and Compliance: The risk control module monitors revenue flow and case data in real time: triggering alerts for significantly high amounts, unusually short periods, or frequent withdrawals by a single entity; cross-validating rights certificates and authorization chains to identify potential fraudulent claims. It also maintains an anti-frivolous litigation blacklist, restricting services to entities identified as frequently filing malicious complaints. When necessary, the risk control module can temporarily suspend fund releases, requiring supplementary supporting materials or manual review.
[0153] 3. Data governance and continuous optimization of models / rules: The data governance and observability module primarily manages the metadata and lineage of fingerprint data, case data, evidence data, document data, and revenue data. It uses an indicator bus to statistically analyze metrics such as retrieval recall rate, evidence credibility distribution, complaint / litigation success rate, and revenue release rate in real time. Simultaneously, it feeds back platform processing results (e.g., whether an app is removed, whether compensation is paid) and judgment results to the training dataset to update the metric learning model, evidence integrity scoring model, and risk control rules.
[0154] Through the aforementioned closed-loop mechanism, the system can continuously adjust its algorithms and rules based on actual results, thereby improving overall performance.
[0155] 4. Active-Active + Cold Backup Deployment and Disaster Recovery: At the infrastructure level, this embodiment adopts a "dual-active + cold backup" architecture: the core original fingerprint database, candidate fingerprint database, and evidence database are deployed in two availability zones for real-time data synchronization, while a cold backup is maintained in a remote data center, with periodic incremental data backups. New features are released in a canary manner, first verifying performance and stability on a small percentage of traffic, and then gradually expanding the scope. When one availability zone fails, the system can automatically switch to another availability zone to continue providing services, keeping the interruption time within a preset range and not affecting ongoing retrieval, evidence collection, and submission tasks.
[0156] The short video infringement evidence collection method provided in this embodiment achieves fully automated processing across the entire chain, from constructing multimodal fingerprints of original videos, cross-platform infringement retrieval and case package generation, evidence collection and preservation, generation and submission of rights protection materials, to revenue settlement and risk control. The above steps respectively address the robust matching problem in multi-variant short video scenarios, the aggregation and chain analysis of multi-platform, multi-variant infringement results, and ensuring evidence quality and justiciability in large-scale automated rights protection scenarios. The data input and output relationships between the above modules and steps are clear, and those skilled in the art can implement the technical solution of this invention accordingly.
[0157] As one or more specific application embodiments of the present invention, in conjunction with Figures 5(a) to 5(e) and Figure 6 The present invention provides a further detailed description of the short video infringement evidence collection method, and the improvements of the present invention include at least the following: 1: A multimodal fingerprint fragment construction and robust fusion method based on frame-level semantic segmentation decomposes the complete video into multiple independently matchable fingerprint fragments, and supports robust recognition of complex variants such as mirroring, occlusion, color correction, AI redrawing and local editing through variable speed alignment, bidirectional encoding and noise-resistant quantization.
[0158] 2: A multi-modal similarity fusion and automatic case package aggregation method for multiple platforms. By fusing three types of similarity—deep metric learning, weighted cosine similarity, and rule confidence—it automatically outputs the infringement probability, variant type, and infringement path, and aggregates multiple suspected links of the same original work into a case package and generates an infringement link graph.
[0159] 3: Based on the linkage mechanism of automatic supplementary collection and rights protection material generation based on credible source scoring and evidence integrity scoring, the quality of the evidence chain is quantitatively evaluated. When the score is insufficient, supplementary collection or manual review is automatically triggered, and the template engine is driven to generate litigation-worthy rights protection materials, so as to achieve a balance between evidence quality and automation.
[0160] Visual hashing, audio fingerprint extraction, subtitle semantic encoding, locality-sensitive hashing, blockchain / TEE (Trusted Execution Environment) notarization, and template engine document generation can be implemented using existing technologies. This invention does not limit the specific implementation method, but focuses on the combination method and data processing flow of the above modules.
[0161] in: As shown in Figure 5(a), steps S501 to S505 correspond to the "Original Video Multimodal Fingerprint Construction and Storage" stage, and complete the preparation of basic data.
[0162] As shown in Figure 5(b), steps S601 to S607 correspond to the "candidate short video acquisition and cross-platform infringement retrieval" stage.
[0163] As shown in Figure 5(c), steps S701 to S705 correspond to the "evidence collection and preservation" stage.
[0164] As shown in Figure 5(d), steps S801 to S804 correspond to the "generation and submission of rights protection documents" stage.
[0165] As shown in Figure 5(e), steps S901 to S904 correspond to the "revenue settlement, risk control and data governance and disaster recovery" stage, providing closed-loop support for the aforementioned process.
[0166] like Figure 6 The diagram shown is a schematic representation of the multimodal fingerprint construction and infringement retrieval subsystem structure in an embodiment of the present invention, wherein: Steps F1 to F5 and DB_O: These correspond to steps S502 to S505 above and constitute the fingerprint fragment construction subsystem.
[0167] C1, DB_C: The candidate video ends reuse the same multimodal fingerprint fragment construction process.
[0168] R1~R5: Corresponding to steps S603 to S605 in the instruction manual, to realize multimodal similarity fusion based on Local Sensitive Hash (LSH) pre-screening.
[0169] R6: Corresponds to steps S606 to S607 in the instruction manual, realizing case package aggregation and infringement link graph construction.
[0170] This embodiment provides the overall system structure as follows: Figure 1 As shown, the method flow is illustrated in Figures 5(a) to 5(e), and the structure of the multimodal fingerprint construction and infringement retrieval subsystem is as follows. Figure 6 As shown.
[0171] like Figure 1 As shown, the system provides services to multiple stakeholders, including: a) Rights holder terminal / client: Allows copyright holders to upload original videos and view case status and revenue; b) Operations personnel terminal: for operations or legal personnel to handle cases that require manual data collection or submission; c) Multiple short video platforms and data service providers: serving as candidate short video data sources and recipients of rights protection complaints; d) Cloud-based infringement evidence collection and automated rights protection platform: including fingerprint and retrieval layer, evidence and document layer, revenue and risk control layer, and data governance and operation and maintenance layer.
[0172] The main modules of the cloud-based infringement evidence collection and automated rights protection platform include: a) Original access and authorization management module, multimodal feature extraction module, fingerprint fragment construction and fusion module, original fingerprint database; b) Candidate collection and management module, candidate fingerprint database, infringement retrieval and similarity fusion module, case package generation and link graph module; c) Evidence collection module, evidence storage and timestamp module, evidence library, evidence integrity scoring and automatic re-collection module; d) Document generation and rule engine module, platform adaptation and delivery communication module; e) Revenue settlement and profit sharing module, risk control and compliance module; f) Metadata and data governance module, monitoring and indicator bus module, and dual-active + cold standby deployment and disaster recovery module.
[0173] Figures 5(a) to 5(e) constitute the overall flowchart of this embodiment. The method of this embodiment includes the following stages: a) Original video multimodal fingerprint construction and storage stage (steps S501 to S505 in Figure 5(a)); b) Candidate short video acquisition and cross-platform infringement retrieval stage (steps S601 to S607 in Figure 5(b)); c) Evidence collection and preservation stage (steps S701 to S705 in Figure 5(c)); d) Document generation and submission stage (steps S801 to S804 in Figure 5(d)); e) Revenue settlement, risk control and data governance and disaster recovery phase (steps S901 to S904 of Figure 5(e)).
[0174] I. Construction and Storage of Original Video Multimodal Fingerprints: Step S501: Original Video Reception and Authorization Information Registration: The rights holder uploads their original short video (V) and authorization certificate information (M) to the platform via the client. The platform manages this through the original content access and authorization module: a) Generate the original content identifier ID_orig; b) Record the upload time t, upload account, data acquisition terminal type, file identifier of the authorization certificate, and authorization scope information; c) The above content is combined into an "Original Content Record" and temporarily stored in the Original Content Registration Form.
[0175] This step provides foundational data for subsequent fingerprint construction and trusted source scoring.
[0176] Step S502: Multimodal feature extraction: In the multimodal feature extraction module, the original short video V is processed as follows: a) Visual feature extraction. First, keyframes are extracted from the video at time intervals, such as 1 to 3 frames per second. Then, the visual fingerprints of the keyframes are calculated using existing perceptual hashing algorithms or hash networks based on convolutional neural networks to obtain the visual hash feature sequence fv.
[0177] b) Audio fingerprint extraction. First, the audio track in the video is extracted, decoded, and resampled. Then, frame-level frequency domain features, such as Mel-frequency Cepstral Coefficients (MFCCs), are calculated, and an audio fingerprint sequence fa is generated using existing audio fingerprinting algorithms.
[0178] c) Subtitle semantic vector extraction. First, the manually provided subtitles or the transcribed text obtained through automatic speech recognition are segmented into sentences. Then, a pre-trained language model (e.g., a bidirectional encoder structure) is used to map each sentence of text into a semantic vector, forming a subtitle semantic vector sequence fs.
[0179] d) Inter-frame motion trajectory feature extraction. First, optical flow or keypoint tracking algorithms are used to estimate the motion trajectory of the main region in the video. Then, the trajectory coordinates and their changes are encoded into a motion feature sequence fm.
[0180] e) Watermark fingerprint extraction. First, perform a digital watermark detection algorithm on the video frame or audio signal; if a recognizable embedded watermark is detected, extract the watermark identifier fw.
[0181] The aforementioned visual hashing, audio fingerprinting, subtitle semantic encoding, motion trajectory and watermark detection can all be implemented using existing technologies, and this invention does not limit the specific algorithms.
[0182] Step S503: Frame-level semantic segmentation and shot boundary detection: like Figure 6 As shown, the semantic segmentation and shot boundary detection module analyzes the video based on fv and fs: a) First, the camera switching position is identified through a scene change detection algorithm, and the video is divided into several coarse-grained shot segments; b) Within each shot segment, based on the similarity change of the subtitle semantic vector fs, it is further divided into semantically continuous sub-segments seg_i; then, the start and end times [ti_start,ti_end] are determined for each seg_i. c) Then calculate the statistical features of adjacent segments before and after seg_i (e.g., subject category, scene type, etc. of adjacent segments) and encode them as scene context features fc_i.
[0183] This step breaks down the originally continuous video into a structured sequence of time segments, with each segment corresponding to a relatively complete semantic unit, laying the foundation for subsequent segment-level fingerprint construction.
[0184] Step S504: Construction and fusion of multimodal fingerprint fragments: In the fingerprint fragment construction and fusion module (see...) Figure 6 In modules F3 to F5 (including the variable speed alignment module, the multimodal bidirectional encoder module, and the noise-resistant quantization and LSH fingerprint fragment generation module), the following operation is performed on each fragment seg_i: a) Variable speed alignment. For time series features such as fv, fa, and fm, dynamic time warping or equivalent variable speed alignment algorithms are used to stretch / compress the feature sequences within a segment along the time axis. The alignment goal is to align feature points representing the same semantic content in the time dimension when candidate videos have acceleration, deceleration, or slight time shifts.
[0185] b) Multimodal encoding. Aligned visual features, audio features, caption semantic vectors, and motion trajectory features are concatenated in chronological order; then, they are input into a bidirectional encoder network (e.g., bidirectional RNN or bidirectional Transformer) to obtain a fused vector fi in a unified embedding space. This network can be trained offline based on historical labeled data.
[0186] c) Noise-resistant quantization and locality-sensitive hashing. The fused vector fi is quantized, mapping continuous values to a set of discrete codewords (e.g., through vector quantization or segmented threshold encoding) to obtain the quantized fingerprint qi. Then, in the quantization design, a fault tolerance range is reserved for disturbances such as brightness changes, color adjustments, partial occlusion, and picture-in-picture overlay. Then, the locality-sensitive hashing algorithm is used to map qi to one or more hash bucket IDs for subsequent approximate nearest neighbor retrieval.
[0187] Finally, a fingerprint fragment record Fi is generated for each seg_i, which includes at least: original content identifier ID_orig, fragment number i, time interval [ti_start,ti_end], fusion vector fi, quantized fingerprint qi, hash bucket ID, and scene context feature fc_i.
[0188] Through the above steps, the present invention represents a complete video as a set of independently matchable multimodal fingerprint segments, solving the problem that "a single fingerprint segment is not robust to local editing and multiple variants".
[0189] Step S505: Trusted Source Scoring and Fingerprint Database Writing: The Trusted Source Scoring submodule calculates a Trusted Source Score (score_src) for each original content record based on the following factors: a) Trust level of the data collection terminal (e.g., certified device, ordinary device); b) The completeness and legality of the authorization certificate (e.g., whether there is complete copyright registration information); c) The publication time and upload path of the original work (e.g., whether it was uploaded directly from the rights holder's account).
[0190] d) The calculated score_src, along with the original content records and fingerprint fragment set {Fi}, is written into the original fingerprint database DB1. Each record in the fingerprint database is associated with its corresponding metadata to facilitate subsequent retrieval and evidence credibility assessment.
[0191] II. Acquisition of candidate short videos and cross-platform copyright search: Step S601: Acquisition of candidate short videos: The candidate acquisition and management module regularly obtains candidate video information through the open interfaces of various short video platforms, content aggregation service interfaces, and data service provider interfaces, including: a) Video unique identifier, URL, title, description, hashtags, author account, and publication time; b) Basic metrics such as views, likes, and comments.
[0192] c) The system can filter out candidate videos that are highly relevant to the rights holder's work based on preset keywords, a list of key accounts, a list of topics, and a time window, write them into the candidate video record table, and store them in the candidate database.
[0193] Step S602: Construction of candidate video fingerprint fragments: For each candidate video record in the candidate database, the system calls the same multimodal feature extraction and fingerprint fragment construction process as the original fingerprint (steps S502 to S504) to obtain the candidate fingerprint fragment set {F′j}, and writes it into the candidate fingerprint database DB2. This ensures that the candidate fingerprints and the original fingerprints are in the same feature space.
[0194] Step S603: Candidate fragment selection based on Locality Sensitive Hash (LSH): like Figure 6 As shown, the Local Sensitive Hash (LSH) bucket candidate filtering module searches for original fragments Fi that fall into the same or adjacent hash buckets in the original fingerprint database DB_O based on the hash bucket ID of the candidate fingerprint fragment F′j, forming a candidate matching set, narrowing the subsequent calculation range, and improving retrieval efficiency.
[0195] Step S604: Multimodal similarity calculation: The infringement retrieval and similarity fusion module calculates multiple similarities for fragment pairs (F′j,Fi) in the candidate set: a) Deep metric learning similarity sd: Input the fused vectors of the two into a pre-trained metric learning network (such as Siamese network or contrastive learning network), and output the semantic similarity in the range [0,1]. b) Weighted cosine similarity sc: Calculate the cosine similarity for visual sub-vectors, audio sub-vectors, and subtitle semantic sub-vectors respectively, and sum them according to the weights obtained from experience or learning; c) Rule confidence score (sr): Scored based on rules such as the proportion of overlapping time segments, similarity of context and scene, whether they contain the same watermark fingerprint, and whether they come from the same topic.
[0196] Step S605: Similarity fusion and variant type determination: The system combines the aforementioned similarity metrics sd, sc, and sr using a preset fusion function, such as linear weighting or fusion based on a small classification model, to obtain the infringement probability P_infringe between the candidate video and a certain original video. Simultaneously, based on features such as the distribution of similarity across different modalities, screen layout patterns, and timeline changes, it determines the variant type (e.g., pure reposting, mirror flipping, picture-in-picture, editing and compositing, adding subtitles, etc.).
[0197] Step S606: Case package aggregation and hot / cold search strategy: For candidate videos with an infringement probability exceeding a preset threshold, the case package generation module aggregates them according to the dimension of original works: a) Group all suspected links pointing to the same ID_orig into the same case package; b) Record the infringement probability, variant type, platform information, and corresponding fingerprint matching fragment for each link in the case package; The search engine maintains a set of hot fingerprints for frequently accessed original works, caches their fingerprint fragments and recent search results in memory to achieve millisecond-level response; for infrequently accessed works, it adopts batch processing and asynchronous retrieval, and returns the estimated completion time on the front end, realizing a "cold start + hot caching" search strategy.
[0198] Step S607: Construction of the Infringement Link Map: Within the case package, the link graph module comprehensively utilizes the forwarding / co-promotion relationships, account follows, embedded watermark information, and the publication time of each suspected link provided by the platform to construct an infringement link graph: a) Use suspected infringing videos and the accounts that posted them as nodes in the graph; b) Treat the relationships between forwarding, secondary creation, and cross-platform synchronization as graph edges; c) Calculate the geographical information and propagation path length of each node.
[0199] This graph structure provides intuitive data support for the development of subsequent strategies for handling bulk complaints and lawsuits.
[0200] III. Evidence Collection and Preservation Process: Step S701: Automated forensic access, screenshotting, and screen recording: After receiving the case package, the evidence collection module automatically accesses each suspected infringing link: a) Load the target page via a built-in browser component or automated script; b) During video playback, capture key screenshots containing infringing segments according to a preset strategy; c) Record the entire playback process or key sections; d) Collect metadata such as page title, URL, publication time, number of views, number of likes, and number of comments.
[0201] Step S702: Evidence hash calculation: For each screenshot file, screen recording file, and metadata file in the evidence files, the system uses algorithms such as SHA-256 or SM3 to calculate a hash digest and writes the association relationship of "file identifier - hash value - generation time" into the evidence metadata table.
[0202] Step S703: Timestamp Evidence Preservation The evidence storage and timestamp module performs the following operations: a) Access the national time service or other trusted time sources to obtain an authoritative timestamp; b) Encapsulate the evidence file hash value, timestamp, collector's identity, etc., into an evidence storage request; c) Obtain evidence by storing the evidence through a blockchain network or a Trusted Execution Environment (TEE); d) Associate the evidence storage certificate with the evidence document identifier and write it into the evidence database DB3.
[0203] This step allows for the verification of evidence documents during subsequent examination sessions, proving that the documents existed at a specific time and were not tampered with.
[0204] Step S704: Collect environmental records and generate a list: The data collection script records additional information during execution: a) Browser fingerprint (User-Agent, window size, plugin status, etc.); b) Collect node IP address, region, and operating system version; c) The main operation steps in the data collection process (such as whether to log in, whether to scroll the page, click path, etc.).
[0205] The above information is stored in the form of a collection list for use in reconstructing the evidence collection environment during court presentations.
[0206] Step S705: Evidence integrity scoring and automatic re-collection / manual review: The evidence integrity scoring module calculates an integrity score (score_evi) for each case package. The scoring model may include: a) Whether all key infringing segments were captured; b) Whether the number of screenshots is sufficient to reflect the infringement; c) Does the screen recording duration cover the main infringing content? d) Does it include copyright information or statements displayed on the platform? e) Collect data on jump depth and authorization chain coverage, etc.
[0207] When the integrity score (score_evi) falls below the first threshold, the system automatically initiates a supplementary data collection task, such as increasing the screenshot angle, extending the screen recording time, or collecting information from the comment section. If the score still does not reach the second threshold after the supplementary data collection, the case package is marked as "requires manual review" and pushed to the operations staff's terminal, whereby the staff decides whether to continue the rights protection efforts or adjust the strategy.
[0208] IV. Process for Generating and Submitting Rights Protection Documents: Step S801: Automatic generation of rights protection documents: The document generation and rule engine module receives the following inputs: a) A list of suspected infringing links, probability of infringement, and variant types in the case package; b) List of evidence and completeness score (score_evi); c) Copyright certificate information for original works (registration number, name of the right holder, scope of rights, etc.); d) Claims set by the rights holder (such as platform removal, demand for a certain amount of compensation, etc.).
[0209] The system automatically generates, based on predefined document templates and rules, the following: relevant letters (such as warning letters), complaint letters from various platforms and other relevant paths, evidence catalog and evidence description, and compensation calculation table (based on the number of infringements, number of views, platform revenue, etc.).
[0210] Different types of infringement and strength of evidence can trigger different combinations of templates and clauses. For example, for cases that are "suspected to be legitimate derivative works with weak evidence", only a communication letter with mild wording is generated, while for cases that are "highly suspected to be malicious plagiarism with sufficient evidence", a formal complaint text is generated.
[0211] Step S802: Platform Adaptation and Format Encapsulation The platform adaptation and submission pipeline module provides structured encapsulation of documents: for platforms providing APIs, it converts document content into the platform-required JSON or form parameter structures; for scenarios that only support manual uploads, it exports documents as standardized PDFs or structured JSON files for manual copying or uploading. Simultaneously, the module generates a unique task identifier and status record for each submission task.
[0212] Step S803: Decision log recording: During document generation and platform adaptation, the rules engine records decision logs, including: the relevant clauses cited and their triggering conditions, the basis for selecting the compensation range (e.g., playback volume range, evidence strength level, etc.), whether to adopt a strategy of consolidation, multiple defendants, or separate lawsuits, and the rule hit status of selecting a particular template from multiple templates. This log information is used for internal auditing and interpreting the basis for automated decision-making in related procedures.
[0213] Step S804: Document Submission and Retry Mechanism The task queue scheduling module uses parallel threads to submit tasks in batches for tasks that support interfaces. If an interface fails or rate limiting occurs, the task is placed in a retry queue according to the set backoff strategy. Tasks that fail after multiple retries are marked as "requires manual submission" and pushed to the operations personnel's terminal.
[0214] The above process ensures that automatic submission can still be completed to the greatest extent possible even when the third-party platform interface is unstable or the call frequency is limited.
[0215] V. Revenue Settlement, Risk Control and Data Governance, and Disaster Recovery: Step S901: Revenue Settlement and Profit Sharing After a case is resolved, the revenue settlement module automatically calculates the revenue due to each party based on: the actual compensation amount (settlement amount, judgment amount, or platform payment amount), the pre-agreed profit-sharing rules (proportion of rights holder, agency, rights protection platform, etc.), and case costs (including estimated values of evidence collection costs, litigation costs, etc.), generates revenue flow records, and provides an interface to the financial system for invoicing and tax declaration.
[0216] Step S902: Risk Control and Compliance The risk control module monitors revenue flow and case data in real time: triggering alerts for significantly high amounts, unusually short periods, or frequent withdrawals by a single entity; cross-validating rights certificates and authorization chains to identify potential fraudulent claims. It also maintains an anti-frivolous litigation blacklist, restricting services to entities identified as frequently filing malicious complaints. When necessary, the risk control module can temporarily suspend fund releases, requiring supplementary supporting materials or manual review.
[0217] Step S903: Data Governance and Continuous Optimization of Models / Rules The data governance and observability module primarily manages the metadata and lineage of fingerprint data, case data, evidence data, document data, and revenue data. It uses an indicator bus to statistically analyze metrics such as retrieval recall rate, evidence credibility distribution, complaint / litigation success rate, and revenue release rate in real time. Simultaneously, it feeds back platform processing results (e.g., whether an app is removed, whether compensation is paid) and judgment results to the training dataset to update the metric learning model, evidence integrity scoring model, and risk control rules.
[0218] Through the aforementioned closed-loop mechanism, the system can continuously adjust its algorithms and rules based on actual results, thereby improving overall performance.
[0219] Step S904: Active-Active + Cold Standby Deployment and Disaster Recovery At the infrastructure level, this embodiment adopts a "dual-active + cold backup" architecture: the core original fingerprint database, candidate fingerprint database, and evidence database are deployed in two availability zones for real-time data synchronization, while a cold backup is maintained in a remote data center, with periodic incremental data backups. New features are released in a canary manner, first verifying performance and stability on a small percentage of traffic, and then gradually expanding the scope. When one availability zone fails, the system can automatically switch to another availability zone to continue providing services, keeping the interruption time within a preset range and not affecting ongoing retrieval, evidence collection, and submission tasks.
[0220] The short video copyright infringement evidence collection method provided in this embodiment has the following beneficial effects: 1. This invention decomposes a complete video into multiple semantically continuous fingerprint segments through frame-level semantic segmentation, shot boundary detection, and scene context encoding, and performs multimodal feature fusion and variable speed alignment at the segment level. Combined with bidirectional encoders, noise-resistant quantization, and locality-sensitive hashing, it can still stably match edited segments even in complex variations such as mirror flipping, occlusion, filter color correction, subtitle addition, variable speed playback, partial editing, and AI redrawing. This avoids the problem of traditional "single fingerprint for the entire segment" solutions failing due to local modifications, significantly improving the recall rate of complex variations under the same resource conditions and significantly enhancing the matching robustness in short video multi-variant scenarios.
[0221] 2. This invention reuses the fingerprint fragment construction process from the original source on the multi-platform candidate short video side. It pre-screens candidate fragments using local sensitive hashing, then integrates deep metric learning similarity, weighted cosine similarity, and rule confidence to output infringement probability and variant type labels. Compared to schemes relying solely on a single modality or simple threshold judgment, this invention can more accurately distinguish between legitimate citations and infringing uses. Through automatic case package aggregation and infringement link graph generation, this invention automatically merges multi-platform, multi-link results according to the original work dimension, forming structured case packages and propagation path data. This reduces the workload of manual deduplication, classification, and propagation link mapping, supports large-scale, cross-platform batch rights protection, and improves the efficiency of cross-platform multimodal infringement retrieval and case aggregation.
[0222] 3. This invention automatically completes page access, screenshotting, screen recording, and metadata collection during the evidence collection stage. It calculates a hash value for each evidence file and uses a trusted time source and blockchain / TEE for timestamping, forming a complete "collection → hashing → signing → on-chain" chain. Simultaneously, it records browser fingerprints, the collection node environment, and an operation list. Combined with an evidence integrity scoring model, it quantitatively evaluates indicators such as the number of screenshots, screen recording duration, key segment coverage, jump depth, and authorization chain coverage. If the score is insufficient, it automatically triggers supplementary collection or manual review, improving the completeness and verifiability of the evidence chain at the system level and reducing the risk of evidence being rejected due to flaws during complaint or litigation stages.
[0223] 4. Based on the infringement probability, variant types, evidence integrity scores, and rights certificate information in the case package, this invention automatically generates documents such as lawyer's letters, platform complaint letters, complaints, evidence catalogs, and compensation calculation tables through a template engine. It also automatically selects appropriate templates and clause combinations according to different infringement forms, significantly reducing the cost of preparing documents for each case. During document generation and submission, the rule engine records every legal clause reference, compensation range selection, and triggering conditions for case merging / splitting, forming a traceable decision log. Combined with interface submission, parallel delivery, and retry queue mechanisms, it maintains a high overall submission success rate even when the platform interface is abnormal or experiencing rate limits. This improves automation while retaining sufficient interpretability and auditability, achieving automatic generation and traceable decision-making for rights protection documents, and improving practical processing efficiency.
[0224] 5. This invention links case processing results with the revenue settlement module, automatically calculating and distributing revenue based on compensation amount, rights protection costs, and preset profit-sharing rules. A risk control module detects and freezes abnormal amounts, abnormal cycles, and frequent withdrawals. Combined with rights certificate verification, authorization chain auditing, and anti-frivolous litigation strategies, it reduces the risk of frivolous litigation and abnormal fund flows. The data governance and observability module uniformly manages the metadata and lineage of fingerprint, case, evidence, document, and revenue data, monitoring real-time retrieval recall rate, evidence credibility distribution, complaint / litigation success rate, and revenue release rate. It also feeds back platform processing and judgment results to model and rule updates, achieving continuous optimization of retrieval, evidence collection, and risk control effectiveness. Simultaneously, through "dual-active + cold backup" deployment and a canary release mechanism, it ensures high availability of the core fingerprint database and evidence storage, reducing the impact of system failures on ongoing retrieval and evidence collection tasks.
[0225] In summary, this invention organically connects the various stages of retrieval, evidence collection, rights protection, and settlement by constructing and fusing multimodal fingerprint fragments, fusing cross-platform multidimensional similarity and case aggregation, and using a scoring-driven evidence supplementation and document linkage mechanism. It solves the technical problems of "difficult retrieval, difficult evidence collection, and difficult large-scale processing" in short video multi-variant infringement scenarios, and has significant technical effects in terms of matching robustness, evidence chain reliability, and large-scale automated processing capabilities.
[0226] This embodiment also provides a short video infringement evidence collection device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated for details already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0227] This embodiment provides a short video copyright infringement evidence collection device, such as... Figure 7 As shown, it includes: The original content access and multimodal fingerprint construction module 1001 is used to acquire original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation.
[0228] The cross-platform infringement retrieval and case aggregation module 1002 is used to obtain candidate short videos from multiple short video platforms and generate candidate fingerprint fragments with the same feature space as the original short video for each candidate short video; it matches the candidate fingerprint fragments with the original fingerprint fragments and integrates multiple similarity calculation indicators to generate a case package containing multiple suspected infringing links.
[0229] 1003 is used to automatically collect evidence from suspected infringing links in the case package, generate evidence files, and store the evidence files by hashing and using a trusted timestamp to form an evidence chain.
[0230] In some optional implementations, the original access and multimodal fingerprint construction module 1001 includes: The semantic segmentation unit is used to perform shot boundary detection on original short videos through frame-level semantic segmentation, dividing them into multiple semantically continuous time segments.
[0231] The multimodal fingerprint fragment construction and fusion unit is used to perform variable speed alignment, multimodal feature fusion, noise-resistant quantization, and local sensitive hashing on each time segment to generate original fingerprint fragments that can be independently matched for the original short video.
[0232] In some optional implementations, the multimodal fingerprint fragment construction and fusion unit includes: The variable-speed alignment subunit is used to perform variable-speed alignment of visual features, audio features, and caption semantic vectors within each time segment.
[0233] The multimodal coding subunit is used to input the variable-speed aligned multimodal features into the bidirectional encoder network to generate a fused vector.
[0234] The noise-resistant quantization and locality-sensitive hashing subunit is used to perform noise-resistant quantization on the fused vector, generate discrete fingerprint codewords, and map the discrete fingerprint codewords to hash buckets through the locality-sensitive hashing algorithm to generate original fingerprint fragments that can be independently matched for the original short video.
[0235] In some optional implementations, the cross-platform infringement retrieval and case aggregation module 1002 includes: The candidate short video acquisition unit is used to filter the acquired candidate short videos based on preset keywords, preset topic tags, or preset account lists, and generate a record of candidate short videos to be analyzed.
[0236] The candidate fingerprint fragment construction unit is used to extract the multimodal features of each candidate short video in the candidate short video record, and decompose each candidate short video into multiple semantic fragments through frame-level semantic segmentation; perform variable speed alignment, multimodal feature fusion, noise-resistant quantization and local sensitive hashing on each semantic fragment to generate candidate fingerprint fragments.
[0237] In some optional implementations, the cross-platform infringement retrieval and case aggregation module 1002 further includes: The multimodal similarity calculation unit is used to calculate the multimodal similarity between candidate fingerprint fragments and original fingerprint fragments; the multimodal similarity includes deep metric learning similarity, weighted cosine similarity, and rule confidence.
[0238] The similarity fusion and variant type determination unit is used to fuse deep metric learning similarity, weighted cosine similarity, and rule confidence through a preset fusion function to obtain the infringement probability of the candidate short video relative to the original short video; and to determine the variant type based on the distribution of multimodal similarity in different modalities, the layout pattern, and the timeline change characteristics.
[0239] The case package aggregation and hot / cold search unit is used to group and aggregate suspected infringing links with an infringement probability higher than a preset threshold according to the dimension of the original work to which they belong, and generate a case package containing multiple suspected infringing links and their infringement information and variant types.
[0240] In some alternative embodiments, the device further includes: The infringement link graph construction module is used to construct an infringement link graph based on the forwarding relationship, secondary creation relationship, cross-platform synchronization relationship, account association relationship and watermark information of each suspected infringing link across preset platforms. The graph has infringing videos and publishing accounts as nodes and infringement dissemination relationship as edges, and the module also counts the geographical distribution information and dissemination path length of each node.
[0241] In some optional implementations, the evidence collection and preservation module 1003 includes: The evidence file generation unit is used to access each suspected infringing link in the case package through an automated browser component, perform screenshot and screen recording operations, collect page metadata, and generate evidence files.
[0242] The environment recording and inventory generation unit is used to record browser fingerprints, collection node IP addresses and operating environment information during the evidence collection process, and generate an environment record inventory that can be used for verification.
[0243] The evidence hash calculation and timestamp storage unit is used to perform hash calculations on evidence documents and environmental record lists, and obtain trusted timestamps for storage, forming an evidence chain containing hash values and storage certificates.
[0244] In some optional implementations, the evidence collection and preservation module 1003 further includes: The evidence integrity scoring unit is used to score the integrity of the evidence chain. The integrity score is based on at least one of the following indicators: number of screenshots, screen recording duration, coverage of infringing segments, and degree of association of the authorization chain. If the integrity score is lower than the first threshold, a supplementary collection task is automatically triggered. If the integrity score is still lower than the second threshold after supplementary collection, the corresponding case package is marked as requiring manual review.
[0245] In some alternative embodiments, the device further includes: The rights protection material generation and submission module is used to receive case packages and corresponding evidence chains, automatically generate rights protection materials according to preset document template rules, and format, package and submit the rights protection materials according to the interface specifications of the target platform.
[0246] The short video infringement evidence collection device provided in this embodiment of the invention can execute the short video infringement evidence collection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0247] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0248] The following is a detailed reference. Figure 8 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 1101, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1102 or a program loaded from memory 1108 into random access memory (RAM) 1103. The RAM 1103 also stores various programs and data required for the operation of the electronic device. The processor 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0249] Typically, the following devices can be connected to I / O interface 1105: input devices 1106 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1107 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 1108 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1109. Communication device 1109 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0250] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 1109, or installed from a memory 1108, or installed from a ROM 1102. When the computer program is executed by the processor 1101, it performs the functions defined in the short video infringement evidence collection method of the embodiments of the present invention.
[0251] Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0252] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the short video infringement evidence collection method shown in the above embodiments is implemented.
[0253] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0254] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for obtaining evidence of copyright infringement in short videos, characterized in that, The method includes: Obtain original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segments; Candidate short videos are obtained from multiple short video platforms, and candidate fingerprint fragments with the same feature space as the original short video are generated for each candidate short video. Candidate fingerprint fragments are matched with original fingerprint fragments, and multiple similarity calculation indicators are combined to generate a case package containing multiple suspected infringing links; After automatically collecting evidence from the suspected infringing links in the case package, evidence files are generated, and the evidence files are hashed and stored with a trusted timestamp to form an evidence chain.
2. The method according to claim 1, characterized in that, The step of semantically segmenting the original short video and generating an original fingerprint fragment of the original short video based on the semantic segmentation includes: Original short videos are divided into multiple semantically continuous time segments by performing shot boundary detection through frame-level semantic segmentation. Each time segment is processed by variable speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate an original fingerprint segment that can be independently matched for the original short video.
3. The method according to claim 2, characterized in that, Each semantic segment undergoes variable-speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate an original fingerprint segment that can be independently matched for the original short video, including: Variable-speed alignment is performed on visual features, audio features, and subtitle semantic vectors within each time segment; The multimodal features aligned by variable speed are input into a bidirectional encoder network to generate a fused vector; The fused vector is subjected to noise-resistant quantization to generate discrete fingerprint codewords; The discrete fingerprint codewords are mapped to hash buckets using the Locality Sensitive Hash (LSH) algorithm to generate original fingerprint fragments that can be independently matched for the original short video.
4. The method according to claim 1, characterized in that, For each candidate short video, generate candidate fingerprint fragments with the same feature space as the original short video, including: Based on preset keywords, preset topic tags, or preset account lists, the acquired candidate short videos are filtered to generate a record of candidate short videos to be analyzed. Extract the multimodal features of each candidate short video from the candidate short video record, and decompose each candidate short video into multiple semantic segments through frame-level semantic segmentation; Each semantic segment is processed by variable speed alignment, multimodal feature fusion, noise-resistant quantization, and locality-sensitive hashing to generate candidate fingerprint segments.
5. The method according to claim 1, characterized in that, The process of matching candidate fingerprint fragments with original fingerprint fragments and integrating multiple similarity calculation indicators to generate a case package containing multiple suspected infringing links includes: Calculate the multimodal similarity between candidate fingerprint fragments and original fingerprint fragments; the multimodal similarity includes deep metric learning similarity, weighted cosine similarity, and rule confidence; By fusing the deep metric learning similarity, weighted cosine similarity, and rule confidence using a preset fusion function, the probability of infringement of the candidate short video relative to the original short video is obtained. Variant types are determined based on the distribution of multimodal similarity across different modalities, screen layout patterns, and timeline variation characteristics. Links suspected of infringement with a probability of infringement exceeding a preset threshold are grouped and aggregated according to the original work they belong to, generating a case package containing multiple suspected infringing links, their infringement information, and variant types.
6. The method according to claim 1, characterized in that, After generating a case package containing multiple suspected infringing links, the method further includes: Based on the forwarding relationships, secondary creation relationships, cross-platform synchronization relationships, account association relationships, and watermark information of each suspected infringing link across preset platforms, an infringement link graph is constructed with infringing videos and publishing accounts as nodes and infringement dissemination relationships as edges. The geographical distribution information and dissemination path length of each node are also statistically analyzed.
7. The method according to claim 1, characterized in that, The system automatically collects evidence from suspected infringing links in the case package, and hashes and stores the generated evidence files with trusted timestamps to form a chain of evidence, including: For each suspected infringing link in the case package, an automated browser component is used to access the page, perform screenshot and screen recording operations, collect page metadata, and generate evidence files; Record the browser fingerprint, collection node IP address and operating environment information during the evidence collection process, and generate an environment record list that can be used for verification; Hash calculations are performed on the evidence documents and environmental record list, and a trusted timestamp is obtained for evidence storage, forming an evidence chain containing hash values and evidence storage credentials.
8. The method according to claim 1, characterized in that, The system automatically collects evidence from suspected infringing links in the case package, and hashes and stores the generated evidence files with trusted timestamps to form a chain of evidence. This also includes: The evidence chain is scored for completeness, and the completeness score is based on at least one of the following indicators: number of screenshots, screen recording duration, coverage of infringing segments, and degree of association of the authorization chain. If the integrity score is lower than the first threshold, a supplementary data collection task will be automatically triggered. If the integrity score is still lower than the second threshold after supplementary collection, the corresponding case package will be marked as requiring manual review.
9. The method according to claim 1, characterized in that, The method further includes: The system receives the case package and the corresponding chain of evidence, automatically generates rights protection materials according to preset document template rules, and formats, packages, and submits the rights protection materials according to the interface specifications of the target platform.
10. A device for collecting evidence of copyright infringement in short videos, characterized in that, The device includes: The original content access and multimodal fingerprint construction module is used to acquire original short videos, perform semantic segmentation on the original short videos, and generate original fingerprint fragments based on the semantic segmentation. The cross-platform infringement retrieval and case aggregation module is used to obtain candidate short videos from multiple short video platforms and generate candidate fingerprint fragments with the same feature space as the original short video for each candidate short video; the candidate fingerprint fragments are matched with the original fingerprint fragments, and multiple similarity calculation indicators are integrated to generate a case package containing multiple suspected infringing links; The evidence collection and preservation module is used to automatically collect evidence from suspected infringing links in the case package, generate evidence files, and preserve the evidence files by hashing and using a trusted timestamp to form an evidence chain.