Internet short video original protection monitoring identification method and system
By dynamically generating baseline digital fingerprints and watermarking strategies, combined with causal graph models and cross-domain infringement hotspot maps, the challenges of identifying infringement at the level of short video clips on the Internet and tracing the source of counterfeit goods in physical products have been solved. This has enabled efficient infringement identification and source tracing, and improved the adaptive capability and accuracy of the monitoring system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies are insufficient to effectively identify and trace infringing acts at the fragment level in short videos on the Internet, and cannot link virtual content infringement with the traceability of counterfeit physical goods, resulting in a decrease in the effectiveness of protection over time and a disconnect from offline commodity transactions.
By generating dynamic baseline digital fingerprints and digital watermarks, and combining causal graph models and cross-domain infringement hotspot maps, the fingerprint model and watermark embedding strategy are optimized to achieve fragment-level identification and source tracing of short video infringements. Furthermore, a two-way data channel between virtual content and physical goods is established by sharing specific identifiers.
It enables accurate identification and source tracing of short video infringements, improves recall and accuracy, quickly responds to emerging threats, ensures intelligent allocation and adaptive capabilities of monitoring resources, and achieves mutual verification and collaborative early warning of virtual content infringement and counterfeit physical goods.
Smart Images

Figure CN121865047A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital copyright protection technology, specifically to an identification method and system for monitoring and protecting the original content of short videos on the Internet. Background Technology
[0002] With the explosive growth of the short video industry on the Internet, the speed of content creation and dissemination is unprecedented. However, the resulting copyright infringement has become increasingly serious. Unauthorized editing, recording, imitation, and copying are rampant, exposing the inadequacy of the current technical protection system in the face of massive, high-frequency, and fragmented short video content.
[0003] Current mainstream protection schemes rely on content-based digital fingerprints (such as media DNA) for post-event comparison. However, the unique time limits, rapid editing, special effects overlays, and image cropping techniques of short videos severely interfere with the stability of traditional feature extraction algorithms. Research shows that at the fragment level, copy detection of short videos is more difficult than that of regular videos. Infringers can easily bypass fingerprint recognition based on a single modality (such as image-only) through simple technical manipulations, such as speed adjustment, color correction, adding borders, and adding noise.
[0004] Secondly, existing protection systems mostly rely on fixed feature libraries and comparison thresholds, making them unable to adapt to the rapid evolution of infringement methods. These systems cannot learn and adjust their strategies based on emerging infringement patterns (e.g., specific types of AI face-swapping, abuse of specific background music), resulting in diminishing monitoring effectiveness over time. Furthermore, many existing solutions focus on infringement detection, but the rights confirmation process is often simplistic and crude (e.g., relying solely on uploaded timestamp evidence), leading to limited probative value of the generated ownership certificates in judicial practice. While blockchain-based evidence preservation technology has been introduced to solidify evidence, seamlessly integrating the "identification" of infringing content with the "proof" of ownership and the subsequent "value realization" in the supply chain to form a closed loop remains a challenge.
[0005] Furthermore, existing protections are limited to the digital content itself and are disconnected from the transactions of physical goods that the content carries or drives. For example, the unauthorized broadcast of an original clothing review video may directly lead to the rise of a counterfeit supply chain, and current technology cannot link video infringement incidents with the tracing of counterfeit goods offline, thus failing to achieve full-chain protection from the source of the idea to the physical product. Summary of the Invention
[0006] This application aims to automatically optimize fingerprint models, watermark embedding strategies, and identification thresholds through continuous learning of new infringement samples and cross-domain correlation feedback. This enables the identification and source tracing of infringement in short videos at the fragment level. Furthermore, by sharing specific identifiers, it establishes a two-way data channel between the short video protection system and the anti-counterfeiting and traceability system for the supply chain, thereby achieving mutual verification and collaborative early warning of virtual content infringement and counterfeit goods clues. This provides an identification method and system for monitoring and protecting the original content of internet short videos.
[0007] To achieve this objective, the following technical solution is adopted in this application: A method for identifying content protection monitoring in short internet videos is provided, comprising the following steps: S1, in response to the request to add the original short video to the database, generate the baseline digital fingerprint of the original short video, and dynamically generate a digital watermark and embed it into the video bitstream; S2, the baseline digital fingerprint, digital watermark embedding key, original short video hash value, and video-associated supply chain commodity identifier generated in step S1 are stored together in the original feature library. The indexing strategy and content update frequency of the original feature library are dynamically adjusted by the cross-domain infringement hotspot map generated based on the output of the causal graph model in step S4. S3, the real-time video stream data of the target monitoring platform is segmented using a sliding window, and the first similar video segment is screened based on fingerprint features; the existing video data of the platform is screened for the second similar video segment through watermark blind detection, and the fingerprint feature screening and watermark blind detection screening share the original feature library; S4. For the similar video segments identified as suspected infringing videos in step S3, construct a causal graph model based on the time axis, analyze the frame-level correspondence and content modification type between the suspected infringing videos and the original videos, and calculate the comprehensive infringement confidence level. S5. Extract the product traceability code from the infringing video finally determined in step S4, synchronize it with the anti-counterfeiting traceability system of the supply chain in real time, and adjust the watermark embedding algorithm preference used in step S1 based on the offline high-quality counterfeit product feature information fed back by the system.
[0008] Preferably, in step S1, the method for generating the baseline digital fingerprint of the original short video includes the following steps: A1. Based on the embedding domain, embedding strength, and error correction coding scheme of the digital watermark dynamically generated for the current original short video, the robustness center of gravity targeted by the digital watermark is analyzed. A2, assess the vulnerability of the digital watermark to protecting the current original short video; A3. Based on the robustness centroid analyzed in step A1 and the vulnerability assessment results obtained in step A2, visual features, audio features, text features and time sequence features are dynamically allocated as weights when generating the baseline digital fingerprint using a preset strategy matrix. A4. Using the dynamic weight strategy output in step A3 as prior knowledge, the network is injected into the pre-trained attention fusion network to drive the network to autonomously extract and fuse features from the multimodal data of the original short video, and finally generate the benchmark digital fingerprint.
[0009] Preferably, in step A1, the method for analyzing the robustness centroid of the dynamically generated digital watermark includes the following steps: A11, read the watermark embedding domain parameters dynamically configured for the current original short video, then perform robust feature mapping, and parse out the primary robust features; A12, collaboratively analyze the embedding strength and error correction coding scheme of the embedding domain and combine the analysis results of the two to quantify the reliability of the digital watermark; A13, combining the comprehensive analysis results of the parsing results of steps A11-A12, the robust center of gravity targeted by the digital watermark is generated through the reverse attack mode.
[0010] Preferably, in step A13, the method for generating the robust centroid targeted by the digital watermark through a reverse attack mode includes the following steps: A131, Match possible attack scenarios from the operation library that can destroy the digital watermark with the watermark features recorded in the comprehensive analysis results; A132, perform feature sensitivity mapping on each attack mode deduced in step A131, and then weight and fuse the mapped sensitivity features to form the robustness center of gravity targeted by the digital watermark.
[0011] Preferably, in step A2, the vulnerability assessment method includes the following steps: A21 generates a list of weaknesses in content amplification through a constructed content and watermark interaction model, specifically: The original short video is analyzed for its content characteristics to obtain the content characteristics of the original short video. The content characteristics obtained through content characteristic analysis are compared with the premise of digital watermark generation to identify contradictions in watermark generation and synthesize them into a list of content amplification weaknesses. A22, retrieve the historical attack vector formed by the first n historical attacks against the digital watermark in attack order, and combine it with the content amplification vulnerability list to calculate the threat coefficient of each type of historical attack in the vector to the current original short video, and generate a quantitative threat mapping table. A23. Randomly select key frames from the original short video as sample frames, then apply the attack from the quantized threat mapping table to each sample frame, then run the watermark decoding process, and calculate the proportion of frames where the watermark is successfully decoded under different attacks. Finally, weight and fuse the proportions of each frame to obtain the average survival probability, which serves as an evaluation index value for the vulnerability of the digital watermark to the original short video.
[0012] Preferably, in step S2, the method for generating the cross-domain infringement hotspot map and dynamically adjusting the indexing strategy and content update frequency of the original feature library includes the following steps: S21. Extract the triplet information corresponding to each infringement identification case from the infringement identification cases analyzed by the causal graph model within a specified time period, including the infringing content, counterfeiting method and related products. S22, Based on the extracted triplet information, perform content clustering, counterfeiting method clustering, and related product clustering on each infringement identification case, and then generate the cross-domain infringement hotspot map after mining the relationship between each element in the triplet information. S23, extract original short videos associated with infringement hotspots that are stored in the original feature library from the cross-domain infringement hotspot map, and add a "high-risk" tag to the index of the baseline digital fingerprint corresponding to the extracted video. At the same time, add an index dimension to the original feature library; and increase the scanning frequency of feature updates for original short videos marked as "high-risk".
[0013] Preferably, in step S3, the method for screening the first similar video segment based on fingerprint features in the real-time video stream data of the target monitoring platform includes the following steps: B1, Extract basic fingerprint features from real-time segmented video captured using a sliding window, the basic fingerprint features including at least one weight dynamically assigned in step A3; B2, compare the basic fingerprint features extracted in step B1 with the benchmark digital fingerprints stored in the latest updated original feature library. If the comparison is successful, the real-time segmented video will be marked as the first similar video segment; If the comparison fails, it is determined that the screening of similar segments in the real-time segmented video has failed.
[0014] Preferably, in step S3, the method for screening second similar video segments in the platform's existing video data using watermark blind detection includes the following steps: C1. Decode the watermark features and basic fingerprint features from the existing video, and perform a reverse query in the latest updated original feature library. If the query is successful, proceed to step C2; otherwise, determine that the existing video is not the second similar video segment. C2: Extract the original short videos successfully retrieved in the reverse query of step C1, and parse the baseline digital fingerprint embedded in the extracted original short videos. Then, perform similarity matching with the baseline digital fingerprint of the existing videos decoded in step C1. If a match is successful, the existing video will be marked as the second similar video segment; If the match fails, it is determined that the existing video is not the second similar video segment.
[0015] Preferably, in step S4, the method for constructing the causal graph model includes the following steps: S41, after aligning the suspected infringing video with the original video for suspected infringing segments, extract key frame pairs within the alignment interval. The key frame pairs include the aligned key frames of the original video and the key frames of the suspected infringing video. S42, define the nodes of the causal graph as each keyframe extracted in step S41, and the node attributes include the multimodal feature vector of the keyframe and the inter-frame similarity score of the keyframe pair; define the edges of the causal graph as the inter-frame transition time order. S43, Analyze the properties of the edges in the causal graph to identify the types of content modifications made to the original video by the suspected infringing video; S44, For each type of content modification identified, calculate the malice weight, and then combine them into the total malice score of the suspected infringing video path; In step S4, the calculated comprehensive infringement confidence level is: the weighted sum of the digital fingerprint similarity between the suspected infringing video and the aligned original video, the watermark detection strength, the total malice of the path, the publisher's historical infringement score, and the supply chain risk value.
[0016] This application also provides an identification system for monitoring and protecting the original content of short videos on the Internet, implementing the aforementioned identification method, including: The fingerprint-watermark collaborative generation module, in response to the original short video's database entry request, generates a baseline digital fingerprint of the original short video and dynamically generates a digital watermark and embeds it into the video bitstream. The original feature library update module, connected to the fingerprint-watermark collaborative generation module, is used to store the generated baseline digital fingerprint, digital watermark embedding key, original short video hash value, and video-associated supply chain commodity identifier in the original feature library. The indexing strategy and content update frequency of the original feature library are dynamically adjusted according to the cross-domain infringement hotspot map. The suspected infringement video screening module is connected to the original feature library update module. It is used to segment the real-time video stream data of the target monitoring platform using a sliding window and screen the first similar video segment based on fingerprint features. For the platform's existing video data, it screens the second similar video segment through watermark blind detection. The fingerprint feature screening and the watermark blind detection screening share the original feature library. The causal graph model construction module is connected to the suspected infringement video screening module. It is used to construct a causal graph model based on the time axis for similar video segments that are screened as suspected infringement videos, and generate the cross-domain infringement hotspot map based on the output of the causal graph model. The infringement confidence calculation module is connected to the causal graph model construction module. It is used to analyze the frame-level correspondence and content modification type between the suspected infringing video and the original video based on the constructed causal graph model, and to calculate the comprehensive infringement confidence. The watermark embedding algorithm adjustment module, connected to the infringement confidence calculation module and the fingerprint-watermark collaborative generation module, is used to extract the product traceability code from infringing videos with a comprehensive infringement confidence score higher than the threshold, synchronize it to the anti-counterfeiting traceability system of the supply chain in real time, and adjust the watermark embedding algorithm preference according to the offline high-counterfeit product feature information fed back by the system, and then notify the fingerprint-watermark collaborative generation module.
[0017] This application has the following beneficial effects: 1. The baseline digital fingerprint generated in this application is controlled by a dynamically generated digital watermark. The resulting fingerprint incorporates a watermark protection strategy, making it more adversarial and laying the foundation for subsequent short video infringement and source tracing at the fragment level. By analyzing and integrating cross-domain infringement hotspot maps of online infringement patterns and offline counterfeit data, the priority of video infringement analysis is adjusted, and incremental updates or enhanced indexes of original video features are triggered, improving the recall rate, precision rate, and response speed for detecting emerging threats of high-risk videos. By allocating dynamic weights in step A3, a balance is found between the timeliness and accuracy of screening similar video fragments. Using the hotspot map as a dynamic control instruction for the original feature library, intelligent tilting of monitoring resources from uniform coverage to hotspot focus is achieved. Through repeated iterations of S1-S5, the robustness of the digital watermark becomes increasingly better, and the fingerprint becomes more adversarial.
[0018] 2. In step S1, the generated baseline digital fingerprint is not generated independently, but is controlled by the dynamically generated digital watermark. The system first analyzes the embedding domain and strength of the watermark to clarify its robust protection focus for the current original short video, and at the same time assesses the potential vulnerability of the watermark when facing specific content characteristics such as violent movement and complex textures. Finally, based on the dynamic weight strategy formulated by the above analysis, the attention fusion network is guided to prioritize the modal features that complement or need to be strengthened with the watermark. The resulting baseline digital fingerprint, which embeds the protection strategy of the digital watermark, becomes a more targeted and adversarial content identity identifier, laying the foundation for subsequent accurate identification of short video infringement at the fragment level and source tracing, that is, realizing mutual verification and collaborative early warning of virtual content infringement and counterfeit goods clues.
[0019] 3. In step S2, by analyzing and integrating the cross-domain infringement hotspot map of online infringement patterns and offline counterfeit data, an enhanced index is established for high-risk content and its priority in the identification queue is increased. At the same time, for newly emerging infringement methods, the incremental update or enhanced index of the corresponding original video features is triggered, which improves the recall rate, precision rate and response speed of high-risk videos in discovering emerging threats.
[0020] 4. In step B1, the low-cost fingerprint features extracted from the real-time segmented video (defined as basic fingerprint features) include at least one weight assigned in step A3. The weights assigned in step A3 are dynamically changing, thus the extracted basic fingerprint features possess randomness. Step A3's weight calculation considers factors such as the robustness of the digital watermark, which are related to accurately identifying short video infringement at the segment level and achieving source tracing. Therefore, screening for the first similar video segment based on the basic fingerprint features, including the weights calculated in step A3, is more targeted. Furthermore, by extracting basic fingerprint features instead of all fingerprint features, and using the weights calculated in step A3 as one of the basic fingerprint features, a balance is struck between improving the timeliness and accuracy of screening for the first similar video segment.
[0021] 5. In step C1, the blind detection result (reverse query) of the digital watermark is used as the initial screening result for the suspected second similar video segment. Then, the basic fingerprint features in the existing video are matched with the benchmark digital fingerprint of the original short video obtained by reverse query. The greater the similarity, the greater the probability that the existing video will be judged as an infringing video in the future. The mutual verification method of digital watermark + digital fingerprint reduces the probability of subsequent misjudgment.
[0022] 6. In step S3, a dynamically updated original feature library is used to ensure that the input and output standards of the two screening strategies are consistent. Whether it is the fingerprint indexing of real-time video streams or the bidirectional verification of watermarks combined with fingerprint features for existing videos, it ultimately returns to the feature library for ownership confirmation. This ensures the consistency of the system's judgment on whether real-time streams and existing videos are infringing. By quickly identifying the first similar video segment, a rapid warning of infringing videos is achieved. Then, the existing videos are screened for the second similar video segment. When the first similar video segment and the second similar video segment are the same video segment, complete evidence from initial discovery to judicial confirmation of ownership is formed.
[0023] 7. In step S4, the infringement cases (including infringing content, counterfeiting methods, and related products) output by the causal graph model serve as the data source for generating a cross-domain infringement hotspot map. This map depicts the infringement situation in real time and directly serves as a dynamic adjustment instruction for the original feature library in step S2. Based on this, the system increases the index priority of relevant original video features and accelerates the feature library update frequency for videos involving such methods or products. This achieves an intelligent tilt of monitoring resources from uniform coverage to hotspot focus, forming a feedback loop from analyzing infringement to optimizing protection, thereby improving the overall monitoring efficiency and adaptability of the system.
[0024] 8. In step S5, the watermark embedding algorithm adjusts the robustness center of the watermark generation based on the feedback feature information of the counterfeit goods, enhancing the adversarial nature of the baseline digital fingerprint controlled by the watermark. Through repeated iterative cycles of steps S1-S5, the robustness center of the digital watermark generated from the original short video becomes increasingly better, and the generated baseline digital fingerprint becomes increasingly adversarial. Attached Figure Description
[0025] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a diagram illustrating the implementation steps of the identification method for original content protection monitoring of short videos on the Internet provided in this application embodiment. Detailed Implementation
[0027] The technical solution of this application will be further described below with reference to the accompanying drawings and specific embodiments.
[0028] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this application. To better illustrate the embodiments of this application, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0029] In the accompanying drawings of the embodiments of this application, the same or similar reference numerals correspond to the same or similar components. In the description of this application, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting this application. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0030] In the description of this application, unless otherwise expressly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0031] The identification method for monitoring and protecting the original content of short videos on the Internet provided in this application embodiment, such as... Figure 1 As shown, the steps include: S1, in response to the request to add original short videos to the database, simultaneously performs multimodal feature extraction based on deep convolutional neural networks and temporal networks, generating a unified high-dimensional vector that integrates visual, audio, text and rhythm features as the baseline digital fingerprint of the video; at the same time, based on the video's metadata and its actively associated product traceability code information, dynamically generates a digital watermark resistant to geometric attacks and signal processing and embeds it into the video bitstream. Step S1, the method for multimodal feature extraction of original short videos includes the following steps: A1. Based on the embedding domain, embedding strength, and error correction coding scheme of the digital watermark dynamically generated for the current original short video, the robustness center of gravity targeted by the digital watermark is analyzed. For example, the system dynamically generates a digital watermark for a close-up video of a skincare product. The technical log shows that the watermark is embedded in the discrete wavelet transform domain of the video, and the algorithm parameters are optimized to resist contrast adjustment and Gaussian blur. After parsing the above scheme in step A1, the conclusion is output: the robustness of this digital watermark lies in maintaining the stability of the visual structure. Therefore, the system determines that deep structural features such as edges and textures in the visual modality are considered dimensions that need to be prioritized during subsequent feature fusion.
[0032] The purpose of step A1 is to transform the underlying technical parameters of digital watermarking into a high-level semantic understanding of its protection strategy focus through a series of mapping and inference rules. This provides a basis for the adaptive adjustment of feature weights when generating a baseline digital fingerprint based on multimodal feature fusion. Step A1 specifically includes the following steps: A11, read the watermark embedding domain parameters dynamically configured for the current original short video, then perform robust feature mapping, and parse out the primary robust features; The embedding domain determines the type of attack the watermark resists. In this embodiment, the precise embedding domain identifier is obtained from the configuration log of the digital watermark generation module. For example, the embedding domain may include the discrete wavelet transform domain, and the sub-bands may include horizontal intermediate frequency details (LH3) and vertical intermediate frequency details (HL3). Then, rule mapping is performed, including: For temporal embeddings, such as digital watermarks that directly modify pixel values, they are extremely sensitive to geometric attacks like video rotation, scaling, and cropping, but have some resistance to compression and filtering. Therefore, for this type of temporal embedding, the centroid rule for parsing could be: resist signal processing distortion, but be wary of geometric attacks such as image cropping.
[0033] For frequency domain embedding, such as the frequency coefficients after watermark embedding transformation, such as mid-to-low frequency coefficients, it is robust to lossy compression, low-pass filtering, and noise addition, but vulnerable to geometric attacks. Therefore, the analytical centroid rule for this frequency domain embedding is, for example, to resist conventional signal processing and compression.
[0034] For wavelet domain embedding, such as mid-frequency subband embedding in wavelet transform (e.g., LH3, HL3), this domain possesses both spatial and frequency domain characteristics. Embedding in a specific subband implies a design goal of resisting compression (frequency domain characteristics) while also considering a certain degree of local clipping or deformation (spatial locality). Therefore, the mapping rule for the robustness centroid of wavelet domain embedding (e.g., LH3) can be set, for example, to resist compression while also considering a certain degree of local clipping or deformation.
[0035] In step A11, by performing robust feature mapping on the read watermark embedding domain parameters, the primary robust features parsed are as follows: For the scheme of embedding in the LH3 horizontal intermediate frequency detail and HL3 vertical intermediate frequency detail subband, the system parses its primary robust features as follows: It focuses on protecting the structural information of object edges and textures in the video and is resistant to compression and mild deformation attacks that preserve these structures.
[0036] A12, collaboratively analyzes the embedding strength and error correction coding scheme of the embedding domain to quantify the reliability of digital watermarks; For example, the system dynamically generates a digital watermark for a skincare product advertisement video. After completing the task, the watermark generation module sends the following technical summary to the feature extraction module: Embedding strength factor: 0.08; Error correction coding scheme: Reed-Solomon code; Coding redundancy: 30%; Expected application scenario tags: e-commerce advertising, high-value product display.
[0037] Then, based on the above technical summary, the embedding strength of the embedding domain of the skincare product advertising video is analyzed. For example, after the system reads that the embedding strength factor is 0.08, it identifies 0.08 as a heavy embedding according to the definition in the knowledge base.
[0038] The definition of embedding strength in a knowledge base is, for example: An embedding strength factor of 0.01-0.03 indicates light embedding, ensuring the absolute invisibility of the digital watermark in original short videos. It is suitable for ultra-high-definition content that is extremely sensitive to image quality and is expected to resist only very slight compression. An embedding strength factor of 0.03-0.08 indicates medium embedding, which seeks a balance between visibility and robustness, is suitable for general scenarios, and is expected to resist standard transcoding and platform compression. An embedding strength factor of 0.08-0.1 indicates heavy embedding, which means that the watermark signal energy is significantly enhanced. When designing this digital watermark, we considered that the watermark signal can still be detected after severe attenuation, and we did not seek visual cloaking of the watermark.
[0039] When the embedding strength factor read is 0.08, based on the above definition in the knowledge base, the system infers that the digital watermark dynamic generation algorithm has pre-determined that the content of the current original short video faces a high risk of being maliciously processed.
[0040] In step A12, the purpose of the system analyzing the error correction coding scheme of the embedded field is to determine, by analyzing the type and redundancy of the error correction coding, what kind of data damage mode the digital watermark has specifically reinforced to deal with.
[0041] For example, the error correction coding scheme read by the system for the embedding field is a Reed-Solomon code with a redundancy of 30%. The process of the system analyzing this error correction coding scheme is briefly described as follows: The read Reed-Solomon code is interpreted as follows: Reed-Solomon code is a type of encoding that excels at correcting bursty, continuous data block errors. In digital video, this error pattern is not random noise, but typically manifests as the loss, replacement, or large-area continuous contamination of data blocks.
[0042] The 30% redundancy figure is interpreted as follows: 30% redundancy means that nearly one-third of the transmitted information is verification data used for error correction. This is a very high percentage, indicating that digital watermarking algorithms allow for a significant reduction in the amount of effective information carried by the watermark in exchange for stronger error correction capabilities.
[0043] Finally, the system fuses the embedding strength analysis results of the embedding domain and the parsing results of the error correction coding scheme of the embedding domain. The fusion yields the following conclusion: this digital watermarking scheme is specifically optimized for attack scenarios where localized areas suffer devastating damage. For example, to obscure a brand logo in a video, infringers might apply a mosaic overlay, dynamic sticker occlusion, or cropping to a specific area of the video for several seconds. These attacks completely invalidate data in continuous spatial regions within the video frame, and the highly redundant Reed-Solomon code is specifically designed for such infringement scenarios.
[0044] For example, the analysis results of the embedding strength (embedding strength factor of 0.08) indicate that the digital watermark is prepared to cope with overall signal attenuation. The analysis results of the error correction coding (Reed-Solomon code, 30% redundancy) indicate that the digital watermark focuses on defending against localized data destruction. After fusing the analysis results of these two methods in step A12, the system finally determines that the generation of this digital watermark adopts an aggressive strategy that combines offense and defense. It defends against area attacks that degrade overall image quality, such as overall video compression and blurring, by increasing the global signal energy, and defends against point attacks aimed at locally erasing information, such as image occlusion and smearing, by configuring error correction mechanisms. This indicates that the system is aware that for current original short videos, infringers may use a combination of methods ranging from overall processing to local destruction. Therefore, the system updates the description of the robustness focus of this digital watermark, for example, to: having been strengthened to resist complex attacks, and adept at surviving in harsh environments where overall signal degradation and local data block loss coexist.
[0045] A13, combining the comprehensive analysis results of the analysis results of steps A11-A12, generates the robust center of gravity targeted by the digital watermark through the reverse attack mode.
[0046] Suppose that, in step A11, the primary robustness features parsed from the watermark embedding domain parameters (such as the wavelet domain mid-frequency subband) read from the digital watermark in the current original short video are: the core of the defense is the visual results in the video, such as image edges and textures. Step A12 reveals that when dynamically generating a digital watermark for the current original short video, the ability to cope with local structural damage is strengthened. In step A12, the primary robustness features parsed in step A11 and the results of the collaborative analysis in step A12 are, for example, fused to form: this digital watermarking scheme is specifically optimized for attack scenarios where local areas suffer devastating damage.
[0047] At this point, in step A13, based on the fusion result of step A12, the method for generating the robust center of gravity targeted by the digital watermark through the reverse attack mode is described as follows: For example, the sub-band feature analysis in step A11 reveals that the LH3 and HL3 sub-bands mainly carry the edge and texture details of the image in the horizontal and vertical directions. Choosing to embed them here indicates that the digital watermarking dynamic generation algorithm considers these structural details to be key features of the original short video and requires focused protection. The embedding strength and error correction coding intent analysis in step A12 shows that a moderately strong embedding (embedding strength factor 0.08) aims to balance watermark invisibility and survivability to resist a certain degree of signal attenuation; the use of highly redundant (30%) Reed-Solomon codes explicitly targets the defense against sudden errors or loss of data blocks.
[0048] Based on the two examples above, the system performs the following causal inference: Since digital watermarks are placed in subbands representing edge texture and are specifically reinforced with burst-error-resistant coding, it is inferred that attacks the dynamic digital watermark generation algorithm is concerned with are those operations that disrupt the continuity of video edge textures or cause severe damage to local data areas. These operations are defined as attack scenarios. For example, for a fusion result stating that "this digital watermarking scheme is specifically optimized for attack scenarios where local areas suffer devastating damage," through causal inference, the most likely attack scenarios include lossy compression, blurring and noise processing, local occlusion and smearing, and slight geometric deformation. Lossy compression, for example, involves video platform transcoding introducing block artifacts and blurring, directly degrading the sharpness and texture of the image edges, which is a typical case of mid-frequency subband information loss. Blurring and noise attacks include operations such as Gaussian blurring and mean filtering, which smooth edges and weaken textures, with the attack effect directly affecting mid-to-high frequency components in the frequency domain. Local occlusion and smearing typically involve adding dynamic emojis, text banners, or performing local mosaic processing on the video, causing the direct loss or replacement of image block data, which is precisely the burst error that Reed-Solomon codes are designed to counter. Slight geometric deformations, which are small-scale rotations, scaling, or lens distortions, can change the spatial position and shape of edges and textures, affecting the accuracy of subband coefficients.
[0049] For example, in a skincare product advertisement video, the dynamic watermark generation algorithm is concerned about scenarios where infringers attempt to preserve the product's visibility for dissemination by using a combination of lossy compression and localized occlusion (such as price tags) to destroy the watermark. The watermark generation algorithm specifically defends against such attacks by reinforcing and embedding mid-frequency sub-bands that characterize the product bottle's outline and label texture, along with error correction.
[0050] After deducing possible attack scenarios (attack patterns), the system performs the following mapping for the deduced attacks: For example, in the attack scenario of lossy compression, the mapping relationship maintained in the attack-feature sensitivity mapping knowledge base is as follows: lossy compression mainly affects the texture details and edge sharpness of visual modalities; therefore, the attribute that needs to be protected is the deep semantic features of vision. As another example, in the attack scenario of partial occlusion, the mapping relationship maintained in the knowledge base is as follows: partial occlusion mainly affects visual information at specific spatial locations; therefore, the attribute that needs to be protected is the global contextual features and temporal features of the visual space. Furthermore, in the attack scenario of slight geometric deformation, the mapping relationship maintained in the knowledge base is as follows: slight geometric deformation mainly affects the spatial location of feature points; therefore, the attribute that needs to be protected is visual features that are relatively invariant to geometric changes.
[0051] The aforementioned protective attributes are the sensitive features obtained through mapping. Finally, the sensitive features mapped to each attack mode are weighted and fused to form the robustness center of gravity targeted by the digital watermark. For example, for the currently generated digital watermark from original short videos, the sensitive features mapped to through attack modes include visual deep semantic features, spatiotemporal context features, and global color statistical features. The priority of these three sensitive features in the visual modality is: visual deep semantic features > spatiotemporal context features > global color statistical features. Therefore, when fusing them into the robustness center of the digital watermark, the weight of visual deep semantic features is greater than the weight of spatiotemporal context features, which is greater than the weight of global color statistical features. The method for weighted fusion of the sensitive features can adopt existing methods, such as quantizing each sensitive feature into normalized feature values, and then weighting and summing the normalized feature values as the feature fusion result. Finally, the robustness characteristics of this digital watermark output are as follows: The focus of this digital watermark generation scheme is to resist the destruction of the integrity of the visual result. Therefore, when generating the basic digital fingerprint, priority is given to protecting and strengthening the feature dimensions extracted from the video that can represent the deep structure and spatiotemporal context of the main object.
[0052] After parsing the robustness center of gravity of the digital watermark assigned to the current original short video in step A1, the method for generating the baseline digital fingerprint of the original short video in step S1 proceeds to the following steps: A2, Evaluate the vulnerability of the dynamically generated digital watermark in step S1 to protect current original short videos. The following example, a short makeup tutorial video, illustrates how to conduct a vulnerability assessment in step A2: Assuming the beauty tutorial video has rich dynamic textures, its digital watermark uses the aforementioned wavelet domain mid-frequency subband embedding with medium embedding strength and a Reed-Solomon code error correction coding scheme. First, through the constructed content-watermark interaction model, a list of content amplification weaknesses is generated. The generation method is as follows: First, the original short videos undergo content characteristic analysis. The system extracts meta-features from the videos and then analyzes the content characteristics of these extracted meta-features. In this embodiment, the meta-features extracted from the original short videos include average motion vector amplitude (representing the intensity of movement of the subject in the shot or video), texture complexity, keyframe chroma distribution, and scene switching frequency. For example, the content characteristics of the meta-features are defined as follows: average motion vector amplitude is categorized into high, medium, and low. For instance, an average motion vector amplitude of "high" indicates intense movement of the subject in the shot or video. The method for defining high, medium, and low average motion vector amplitude is as follows: amplitudes above a first threshold are defined as "high," amplitudes between a second and first threshold are defined as "medium," and amplitudes below the second threshold are defined as "low," where the first threshold is greater than the second threshold. Similarly, texture complexity is defined as high, medium, and low to characterize the image refinement of makeup textures, highlight details, etc. Keyframe chroma distribution is defined as wide or narrow. Wide indicates rich and varied colors, while narrow indicates simple color variations. The definition principle is the same as that of texture complexity and will not be repeated here. Scene switching frequency is also defined as high, medium, and low.
[0053] Then, the system performs watermark parameter matching analysis on the original short videos, comparing the content characteristics obtained from the above content characteristic analysis with the premise of digital watermark generation to identify contradictions in watermark generation. For example, the embedding effect of wavelet domain watermarks in high-frequency motion regions is unstable. Severe motion causes the wavelet coefficients to change drastically on the time axis, which may cause the energy of the watermark embedded at a fixed position to be averaged out or difficult to synchronize during decoding, thereby reducing the signal-to-noise ratio. At this time, the system identifies the contradiction: there is an inherent conflict between high motion characteristics and static domain embedded watermarks (such as wavelet domain watermarks), resulting in a decrease in the inherent robustness of the watermark in motion segments. For another example, rich textures are both a cover and an interference. On the one hand, they are conducive to watermark hiding; on the other hand, high texture density itself will occupy a large amount of mid-frequency sub-band energy, which may compete with the watermark signal, forcing the actual effective embedding strength to be relatively reduced. At this time, the system identifies the contradiction: in the most complex texture region, the expected value of the effective embedding strength of the watermark needs to be lowered by one level.
[0054] For the example above, the system's final output is an enlarged list of weaknesses, such as: [First weakness: robustness attenuation of watermarks in high-motion frames; Second weakness: insufficient effective strength of watermarks in high-texture areas].
[0055] Subsequently, the system retrieves the historical attack vectors formed by the first n historical attacks against the digital watermark dynamically generated in step S1, arranged in attack order. Combined with the content amplification vulnerability list formed above, the system calculates the threat coefficient of each type of historical attack in the vector to the current original short video and generates a quantitative threat mapping table.
[0056] For example, if the digital watermark generated in step S1 is a wavelet domain intermediate frequency watermark, the system will retrieve all historical attack vectors targeting wavelet domain intermediate frequency watermarks from the attack database. For instance, historical attacks targeting wavelet domain intermediate frequency watermarks include: Attack mode p1: Gaussian blurring is applied to the video. By embedding a wavelet domain mid-frequency watermark in the video, the success rate of defense against p1 attacks is, for example, 35%.
[0057] Attack mode p2: Localized brightness or contrast adjustment of the video. By embedding a wavelet domain mid-frequency watermark in the video, the success rate of defense against p2 attacks is, for example, 25%.
[0058] Attack mode p3: Adaptive compression of the video. By embedding a wavelet domain intermediate frequency watermark in the video, the success rate of defense against p3 attacks is, for example, 40%.
[0059] The above attack patterns p1-p3 form a historical attack vector, for example, expressed as: [attack pattern p1, attack pattern p2, attack pattern p3].
[0060] Then, combining the content amplification vulnerability list formed in step A21, the threat coefficient of each type of historical attack in the historical attack vector to the current original short video is calculated, and a quantitative threat mapping table is generated.
[0061] For example, for attack mode p1, Gaussian blur directly attenuates the energy of embedded subbands (including LH3 and / or HL3) in the current digital watermark. Combined with the second weakness listed in the content amplification weakness list—insufficient effective strength of watermarks in high-texture areas—the system infers that the signal-to-noise ratio of watermarks in densely textured areas of the current original short video is at extremely high risk of dropping to the detection threshold edge when subjected to blur attacks. Therefore, the threat coefficient of the p1 attack to the current original short video is calculated to be: High.
[0062] For example, regarding attack mode p3, the current watermark error correction coding uses Reed-Solomon codes with a redundancy of 30%. While Reed-Solomon codes can counteract the block artifacts described above, high compression ratios (e.g., greater than 28) may exceed the error correction capability. Therefore, considering the first weakness listed in the content amplification weakness list—high motion frame watermark robustness attenuation—the system infers that under high compression, the watermark carrying capacity of motion frames is more easily flattened. Therefore, the calculated threat level of the p3 attack to current original short videos is medium to high.
[0063] Finally, the output quantitative threat mapping table is expressed as: [Threat Mode: Threat Coefficient] = [p1: High, p2: Medium, p3: Medium-High].
[0064] After calculating the quantized threat mapping table corresponding to the digital watermark dynamically generated in step S1, key frames in the original short video are randomly selected as sample frames. Then, attacks from the quantized threat mapping table are applied to each sample frame. The watermark decoding process is then run, and the proportion of frames where the watermark is successfully decoded under different attacks is calculated. Finally, the proportions of each frame are weighted and fused into an average survival probability, which is used as an evaluation index value of the vulnerability of the original short video protected by the digital watermark.
[0065] It should be noted that when calculating the proportion of frames where the watermark was successfully decoded, the system does not process the entire video, but randomly selects key frames from the video as sample frames. The preferred key frames are video frames marked in the list of content magnification weaknesses, such as high motion frames or video frames with high texture areas.
[0066] The attack from the quantized threat map is applied to the sample frame. The strength is set according to historical data. For example, if a Gaussian blur with a radius of 2 pixels is applied in a historical attack, then a Gaussian blur with a radius of 2 pixels will be applied in this attack as well.
[0067] Then, the watermark decoding process is run on the video frames after the simulated attack, and the proportion of frames where the watermark is successfully decoded under different attacks is statistically analyzed. For example, the watermark decoding success rate is 65% under a blur attack (blurring radius is, for example, 3 pixels) and 72% under a high compression ratio (e.g., a compression ratio of 25). Then, the weighted average survival probability is calculated by combining the threat coefficients of the blur attack and the high compression ratio attack. For example, if the threat coefficient of the blur attack is "high" and quantized as "3", and the threat coefficient of the high compression ratio attack is "low" and quantized as "1", then the weighted average survival probability of the two is: (65%×3+72%×1) / (3+1).
[0068] Assuming the calculated average survival probability is 70%, this falls within a range requiring vigilance. A vulnerability assessment report on the protection of current original short videos using this digital watermark shows: The expected survival probability of this digital watermark in the original short video is 70%. The main vulnerabilities are: high-speed motion segments and highly textured areas are weak points protected by this digital watermark; its resistance to Gaussian blur attacks is below average. Therefore, when generating the baseline digital fingerprint for the current original short video, it is recommended to increase the weighting of motion optical flow features (to compensate for the weaknesses of the watermark in moving frames) and deep texture semantic features (to compensate for the weaknesses of the watermark in textured areas) to ensure that the fingerprint can still achieve high-precision recognition even if the watermark partially fails.
[0069] After evaluating the vulnerability of the dynamically generated digital watermark to protect the current original short video in step A2, the method for generating the baseline digital fingerprint of the original short video in step S1 proceeds to the following steps: A3. Based on the robustness centroid analyzed in step A1 and the vulnerability assessment results obtained in step A2, visual features, audio features, text features and time-series features are dynamically assigned as weights when generating the baseline digital fingerprint through a preset strategy matrix. For example, the robustness center of gravity analyzed in step A1 is: resisting attacks against the integrity of visual structure (such as blurring and partial occlusion), with a focus on protecting deep semantic structural features of vision. After the system inputs this semantic into a preset policy matrix, it transforms the semantic into an initial score for the importance of each modality. For example, using a 5-point scale, a higher score indicates a greater robustness center of gravity. For example, the robustness center of gravity score for the deep structural visual features of the current original short video is "5", the robustness center of gravity score for global statistical visual features is "2", the robustness center of gravity score for audio features and text features is "1" for both, and the robustness center of gravity score for temporal features (because structural attacks may involve temporal editing) is "3".
[0070] After inputting the vulnerability assessment result from step A2 into the policy matrix, the policy matrix generates a compensation vector based on the input vulnerability assessment structure. For example, the vulnerability assessment result obtained in step A2 is: the digital watermark has weaknesses in high-motion frames and high-texture areas, and its resistance to Gaussian blur attacks is lower than the mean. The policy matrix calculates and generates compensation coefficients for each feature on which the baseline digital fingerprint depends based on a preset compensation logic.
[0071] The following is an example of the compensation logic preset in the strategy matrix: Digital watermarking has weak protective capabilities in moving frames, so temporal features (optical flow) are enhanced to track the consistency of object motion; If digital watermarks do not provide sufficient protection in textured areas, then enhance the deep semantic features of the visual image to strengthen the abstract representation of the texture. Digital watermarking is weak against blur attacks, so it suppresses visual global statistical features that are sensitive to blur (such as color histograms) and instead relies on other deeper features that are more resistant to blur.
[0072] Based on the vulnerability assessment results and the aforementioned compensation logic, the system adjusts the compensation coefficient for deep structural visual features from the base value of 1 to 1.3 (enhancement), the compensation coefficient for global statistical visual features from the base value of 1 to 0.7 (suppression), and the compensation coefficient for temporal features from the base value of 1 to 1.2 (enhancement) according to preset compensation rules. The compensation coefficients for audio and text features remain unchanged at the base value of 1. A larger compensation coefficient results in a greater weighting of the corresponding feature when generating the baseline digital fingerprint.
[0073] The policy matrix also includes a built-in weight generation function, expressed as follows: , These represent the robustness centroid score and compensation coefficient, respectively. This represents the basic weight of each modal feature, such as temporal features and audio features, in generating the baseline digital fingerprint.
[0074] Assuming the base weights are predefined as an even distribution, the even distribution of base weights for deep structure visual features, audio features, text features, and temporal features can be expressed as follows: First, calculate the center of gravity adjustment weight. The robustness center of gravity score Transforming into a probability distribution, the Softmax function ensures that the sum of all weights is 1. After calculation, for example... =[0.5 (visual features), 0.1 (audio features), 0.1 (text features), 0.3 (temporal features)]. Then, the compensation coefficients are... Acting on In the middle, the compensated weight is obtained. .For example, =[0.5×1.3 (visual features), 0.1×1.0 (audio features), 0.1×1.0 (text features), 0.3×1.2 (temporal features)]=[0.65,0.1,0.1,0.36]. Finally, for Normalize the sum so that it equals 1. A normalization method could be: 0.65 + 0.1 + 0.1 + 0.36 = 1.21. For example, normalizing 0.65 results in a normalized value of 0.65 / 1.21 = 0.537. This normalized value is used as the weight of the corresponding modal feature when generating the baseline digital fingerprint.
[0075] After generating the weights of each modal feature in the original short video in step A3 when generating the baseline digital fingerprint, in step S1, the method for generating the baseline digital fingerprint of the original short video is transferred to the following step: A4 uses the dynamic weight strategy output from step A3 as prior knowledge to inject into the pre-trained attention fusion network, driving the network to autonomously extract and fuse features from the current multimodal data of original short videos, and finally generate a benchmark digital fingerprint.
[0076] It should be noted that the process of generating the baseline digital fingerprint using the attention fusion network can employ existing methods, which will not be detailed here. The technical advantage of this application lies in the fact that the prior knowledge used by this network to generate the baseline digital fingerprint is the dynamic weight strategy output in step A3.
[0077] In step S1, the generated baseline digital fingerprint is not generated independently, but is controlled by the dynamically generated digital watermark. The system first analyzes the embedding domain and strength of the watermark to clarify its robust protection focus for the current original short videos, while assessing the potential vulnerability of the watermark to specific content characteristics such as violent movement and complex textures. Finally, based on the dynamic weight strategy formulated by the above analysis, the attention fusion network is guided to prioritize modal features that complement or need to be strengthened with the watermark. The resulting baseline digital fingerprint, which incorporates the protection strategy of the digital watermark, becomes a more targeted and adversarial content identity identifier. This lays the foundation for subsequent accurate identification of short video infringement at the fragment level, source tracing, and mutual verification and collaborative early warning of virtual content infringement and counterfeit goods clues.
[0078] The following describes the method for dynamically generating robust digital watermarks resistant to geometric attacks and signal processing in step S1, based on the metadata of the current original short video and its actively associated product traceability code information: For example, an original clothing brand released a short video showcasing a new down jacket. This down jacket has a built-in NFC chip, and its unique traceability code is SN12345. The method for dynamically generating a digital watermark for this video is as follows: First, extract the video's metadata, such as the video creator ID: userA, the unique work ID: Video789, the video publication timestamp: 20251213, and the product traceability code: SN12345.
[0079] Then, the above information is concatenated according to a preset specific format to generate the original payload plaintext Ptext= userA|Video789|SN12345|20251213.
[0080] Subsequently, the original payload text is signed and encrypted using the private key of an asymmetric encryption algorithm such as RSA, and then Base64 or binary encoded into the ciphertext, thereby binding the video copyright protection and the product anti-counterfeiting cryptographic evidence generated through encoding together.
[0081] Subsequently, the system preprocesses and analyzes the current original short videos to extract key features. For example, the analysis results of the video are as follows: the video combines static product display with dynamic model catwalk; by calculating the motion vector amplitude of the video frames, it is identified that the video alternates between low-motion (such as close-up shots) and high-motion (such as model catwalk) segments; by analyzing the video footage, it is identified that the down jacket fabric has rich texture and a relatively simple background.
[0082] Then, based on the analysis results of the video content, an embedding strategy is selected from the strategy library. For example, for static product display segments with low motion characteristics and rich textures, the embedding domain is preferably the discrete cosine transform domain, specifically embedding at the mid-frequency coefficients, because the discrete cosine transform domain has better robustness to lossy compression attacks and can effectively resist platform transcoding. The embedding strength is adaptively adjusted by lowering the strength factor according to the texture complexity. The adaptive adjustment process of the embedding strength factor has been partially described above and is not within the scope of the claims of this application; therefore, it will not be described in detail here.
[0083] For model catwalk segments with high motion characteristics, the preferred embedding domain is the spatiotemporal domain. This is because the performance of the discrete cosine transform domain degrades during rapid motion, while embedding the watermark information along with the motion vector allows the watermark information to move with the content, thus resisting geometric attacks such as inter-frame cropping and jitter. The embedding strength is also adaptively adjusted.
[0084] Finally, the encrypted watermark payload is divided into multiple data blocks, and an error correction scheme is added. These data blocks are then embedded into different spatiotemporal locations in the video, such as different frames or different frequency bands.
[0085] In step S1, a baseline digital fingerprint for the original short video is generated, and the dynamically generated digital watermark is embedded into the video stream, as follows: Figure 1 As shown, the identification method for internet short video original content protection monitoring provided in this application proceeds as follows: S2, the baseline digital fingerprint, digital watermark embedding key, original short video hash value, and supply chain commodity identifier associated with the video generated in step S1 are stored together in the original feature library; the indexing strategy and content update frequency of the original feature library are dynamically adjusted by the cross-domain infringement hotspot map generated based on the output of the causal graph model in step S4. A digital watermark embedding key is a set of confidential parameters that control the watermark embedding process and subsequent extraction and decoding processes. For example, if a spread spectrum watermarking algorithm based on discrete cosine transform is used, the key typically includes a pseudo-random noise sequence seed used to modulate the watermark information, a specific frequency coefficient position matrix for watermark embedding, and an error correction coding scheme.
[0086] The hash value of original short videos is calculated by applying a cryptographic hash function such as SHA-256 to the binary bitstream of the video file; details are omitted here. Supply chain product identifiers include product traceability codes, such as the unique code on the down jacket mentioned above.
[0087] The process of storing the baseline digital fingerprint, digital watermark embedded key, original short video hash value, and video-associated supply chain product identifier into the original feature library is briefly described below: The vector consisting of the video hash value, watermark embedding key, and product traceability code is treated as a data packet, generated into a transaction, and written to the blockchain. Then, in the original feature library, a record is created using the video hash value as the primary key, fully storing the video hash value, baseline digital fingerprint, watermark key, product traceability code, and video metadata. Simultaneously, a vector index is created for the baseline digital fingerprint to facilitate subsequent video similarity searches, and an inverted index is created for the product traceability code and metadata to facilitate subsequent keyword searches.
[0088] The following example illustrates the method for dynamically adjusting the indexing strategy and content update frequency of the original feature library based on the generated cross-domain infringement hotspot map in step S2: Suppose that within the last 7 days, the infringement identification of videos A, B, and C was processed using the causal graph model in step S4. Video A is, for example, a review video of a certain brand of handbags; video B is, for example, a video of identifying a certain brand of liquor; and video C is, for example, a trial video of a certain brand of skincare products.
[0089] First, from the infringement identification analysis results of the above three videos using the causal graph model, the triple information of (infringing content, counterfeiting method, related products) is extracted. For example, the triple information for video A is expressed as: (Video A: Review of a certain brand of handbags, counterfeiting method: re-shooting with labels removed, related products: high-quality counterfeit handbags of a certain brand, product number H001). The triple information for video B is expressed as: (Video B: Identification video of a certain brand of liquor, counterfeiting method: partial mosaic, related products: liquor of this brand, batch 2024 Spring). The triple information for video C is expressed as: (Video C: Trial of a certain brand of skincare products, counterfeiting method: re-shooting with labels removed, related products: essence of this brand, model S123)
[0090] Then, based on the extracted triplet information, video content clustering, counterfeit method clustering, and related product clustering are performed. For example, if video A's review of a brand's handbag and video C's trial of a brand's skincare products both fall under the category of luxury goods or high-end product displays, then the content of videos A and C will be categorized as luxury goods or high-end product displays. Counterfeit method clustering, for example, categorizes the counterfeit methods in videos A and C as "re-filming with labels removed," and the counterfeit methods in video B as "partial mosaic processing." Related product clustering, for example, categorizes the related products of video A as leather goods, video B as alcoholic beverages, and video C as skincare products.
[0091] Then, the system uses graph calculations to discover the relationships between the elements in the triplet information. For example, the appearance of label removal and re-shooting techniques in videos showcasing luxury or high-end goods is positively correlated with the probability of counterfeit leather goods and skincare products. Finally, the high-risk triangular relationships between content, counterfeiting techniques, and clustering results are plotted in the cross-domain infringement hotspot map. A high-risk triangular relationship is, for example, a correlation between a certain type of video content exhibiting a certain type of alteration technique and the appearance of a certain type of counterfeit physical goods. The purpose of the cross-domain infringement hotspot map is to present these clusters of relationships.
[0092] After generating the cross-domain infringement hotspot map, in step S2, the method for dynamically adjusting the index building strategy and content update frequency of the original feature library based on the map is as follows: First, original short videos with original features stored in the original feature library that are associated with infringement hotspots are extracted from the cross-domain infringement hotspot map. Then, a "high-risk" label is added to the index of the baseline digital fingerprint generated by the extracted videos. At the same time, the index dimension of the original feature library is increased. Furthermore, the scanning frequency of feature updates is increased for original short videos marked as "high-risk".
[0093] For example, if a particular infringement hotspot in the data map is related to the display of luxury goods and high-end products, then original short videos showcasing luxury goods and high-end products are extracted from the original short video library. These extracted videos also contain original features stored in the original feature library. Then, a "high-risk" tag is added to the index of the baseline digital fingerprint generated from the extracted videos. This way, during a full-network scan, the fingerprints of these videos will be prioritized for comparison, ensuring timely monitoring. Simultaneously, an index dimension is added to the original feature library. For example, if the infringement hotspot uses the method of removing tags and re-enacting videos, then an index for "anti-removal-tag re-enactment features" is added to the original feature library. When a second fingerprint calculation is needed for this original short video, this added index quickly matches suspected infringing videos that have undergone tag removal and re-enactment processing. Then, this infringing video is analyzed to develop a baseline digital fingerprint that is more resistant to tag removal and re-enactment methods such as cropping and blurring.
[0094] For original short videos marked as "high-risk," the system increases the frequency of scanning for feature updates. For example, ordinary videos are set to undergo a full review of their fingerprint validity once a month, while for videos marked as "high-risk," the feature update frequency is dynamically adjusted to trigger once a week, in order to promptly detect and learn new variant attack methods targeting such infringing videos.
[0095] After creating or updating the original feature database through step S2, the identification method for monitoring and protecting the original content of short videos on the Internet provided in this application, such as... Figure 1 As shown, proceed to the following steps: S3 uses a sliding window to segment the real-time video stream data of the target monitoring platform and screens the first similar video segment based on fingerprint features; for the platform's existing video data, it uses watermark blind screening to screen the second similar video segment. The fingerprint feature screening and the watermark blind screening share the same original feature library. The following example, using the copyright infringement of an original high-end skincare product review video, details the specific processes of fingerprint feature screening and watermark blind screening: Assuming the original short video has already been added to the original feature database, its baseline digital fingerprint and digital watermark containing product traceability codes, among other original features, are stored in the database. The purpose of screening for the first similar video segment in the real-time video stream data of the target monitoring platform is to quickly identify suspected infringing videos in live streaming or real-time video upload scenarios.
[0096] First, the system slides and captures live streams or newly uploaded video streams in windows of fixed duration (e.g., 5 seconds), with a step size of 2.5 seconds. This ensures that adjacent video segments have 50% overlap, guaranteeing that no video segment from any starting point is missed.
[0097] For real-time video streams, to balance screening speed and accuracy, low-cost fingerprint features are calculated for each 5-second segment, rather than complete deep fingerprint features. For example, for visual features fused into the baseline digital fingerprint, the perceptual hash of keyframes or the mainstream color distribution vector is extracted, rather than all visual features; for audio features integrated into the fingerprint, a statistical overview of Mel-frequency cepstral coefficients is extracted. To improve the randomness of low-cost fingerprint feature extraction and thus enhance the targeting of video segment screening in real-time video stream data, preferably, the low-cost fingerprint features extracted from real-time segmented video (defined as basic fingerprint features) include at least one weight assigned in step A3 above. The weights assigned in step A3 are dynamically changing, thus possessing randomness. The weight calculation in step A3 considers factors such as the robustness centroid of digital watermarks, which are related to accurately identifying short video infringement at the segment level and achieving source tracing. Therefore, screening the first similar video segment based on the basic fingerprint features, including the weights calculated in step A3, is more targeted.
[0098] For example, an infringer is illegally broadcasting a "product close-up display" segment of a skincare product review during a live stream. The system captures 5 seconds of this segment's stream data in real time and extracts the basic fingerprint features of this real-time segmented video, such as fingerprint features representing the main color tone of the screen and the frequency contour of the background music, and the weights calculated for these features in step A3.
[0099] Then, the basic fingerprint features extracted from this real-time segmented video will be compared with the benchmark digital fingerprints stored in the latest updated original feature library. If the comparison is successful, the real-time segmented video will be marked as the first similar video segment; If the comparison fails, the screening for similar segments of the real-time segmented video is deemed to have failed.
[0100] In this application, the method for comparing the basic fingerprint features with the benchmark digital fingerprints in the original feature library is briefly described as follows: For example, the basic fingerprint features extracted from real-time segmented video include music features m1 and a weight wm1 dynamically assigned to m1 in step A3. A baseline digital fingerprint exists in the original feature library, carrying music features m2 and a weight wm2 dynamically assigned to m2 in step A3. If the similarity value of the music features between m1 and m2 exceeds a preset first threshold, and the similarity value between wm1 and wm2 exceeds a preset second threshold, then the basic fingerprint features are determined to be similar to the baseline digital fingerprint, and the comparison is successful; otherwise, the comparison fails. There are many methods for comparing the similarity between m1 and m2, and between wm1 and wm2. After quantifying the features into values, the absolute value of the difference between the values of m1 and m2 is calculated. If the absolute value of this difference is less than a threshold, then the two are considered similar.
[0101] By extracting basic fingerprint features from real-time segmented videos instead of all fingerprint features using the above method, the efficiency of screening the first similar video segments is improved. By adding the weight of one of the basic fingerprint features calculated in step A3, the effectiveness of screening the first similar video segments is improved.
[0102] In step S3, the method for screening the second similar video segments in the platform's existing video data using watermark blind detection is as follows: A comprehensive infringement investigation is conducted on existing videos on the platform, especially identifying historically infringing or complexly processed infringing videos. The screening of second-similar video clips based on watermark-based blind detection does not have high timeliness requirements and is usually an offline or background task; the system scans all existing videos on the platform according to a schedule.
[0103] The system directly runs a blind digital watermarking algorithm on existing video files. This algorithm attempts to detect and decode the presence of a watermark signal conforming to a predetermined format in the video's frequency domain. For example, one month ago, a user uploaded a complete original skincare product review video without the intro and outro. When scanning existing videos, the system successfully detected and decoded a series of watermark features from the wavelet transform domain of the video file. The decoded result was a traceability code for a certain brand's product. Then, the system performs a reverse lookup of the decoded traceability code in the latest updated original feature database. The feature database returns a unique original short video ID and metadata corresponding to the traceability code, indicating a successful watermark feature match. Finally, the system marks the existing video as the second most similar video segment.
[0104] To verify the accuracy of screening for the second similar video segment, after a successful reverse query, the original short video that was successfully queried is extracted, and the baseline digital fingerprint embedded in the extracted original short video is parsed. At the same time, the basic fingerprint features are parsed from the existing video and matched with the baseline digital fingerprint of the parsed original short video. If the match is successful, the existing video is marked as the second similar video segment; otherwise, the existing video is determined not to be the second similar video segment.
[0105] After step S3, the first similar video segment is screened from the real-time video stream data, and the second similar video segment is screened from the existing video data, as follows: Figure 1 As shown, the identification method for original content protection monitoring of short videos on the Internet provided in this embodiment proceeds to the following steps: S4. For the similar video segments screened in step S3, construct a causal graph model based on the time axis, analyze the frame-level correspondence and content modification type between the suspected infringing video (including the first or second similar video segments screened) and the original video, and calculate the comprehensive infringement confidence level. For example, the original video (an original short video) is a 30-second close-up video of a high-end skincare product, showing the product bottle from multiple angles in a professional studio, accompanied by brand background music and voice-over narration. The suspected infringing video identified in step S3 is a 45-second video sharing kitchen essentials. In the segment from 10 to 25 seconds, a similar product bottle is shown placed on a kitchen stove. The background music and narration have been replaced with a cooking theme, but the sequence of movements showing the bottle rotating is highly similar to the original video.
[0106] For the two video segments mentioned above, in step S4, coarse-grained alignment is first performed to determine which part of the original video the suspected infringing video corresponds to. For example, the segment from second 10 to second 25 of the suspected infringing video corresponds to the segment from second 5 to second 20 of the original video. The video segment alignment method is an existing method and will not be described in detail. Subsequently, within the alignment interval, keyframe pairs are extracted at fixed time intervals. Each keyframe pair includes a keyframe from the original video and a keyframe from the suspected infringing video. The keyframe pair is expressed as (OFi, SFj), where OF5 and SFj represent the i-th keyframe from the original video and the j-th keyframe from the suspected infringing video, respectively. For example, OF5 represents the video frame at second 5 of the original video.
[0107] Then, the nodes of the causal graph are defined as each keyframe extracted above, and the node attributes include the multimodal feature vector of the keyframe and the inter-frame similarity score of the keyframe pair; the edges of the causal graph are defined as the inter-frame transition time order.
[0108] For example, each keyframe constitutes a node in a cause-effect graph, denoted as . (the first in the causal diagram) (Number of nodes). The node attribute is the multimodal feature vector of the keyframe, that is, the vector composed of the multimodal features used to generate the baseline digital fingerprint in step S1. The inter-frame similarity score is calculated based on the multimodal features of the suspected infringing video keyframe in the keyframe pair and the multimodal features of the original video keyframe. There are many existing methods for calculating the inter-frame similarity score, which will not be specifically described here. The edges in the causal graph are from the previous node. Point to the next node The side represents the causal flow in time and the derivation relationship in content.
[0109] Subsequently, the properties of the edges in the causal graph were analyzed to identify the types of content modifications made to the original video by the suspected infringing video. The analysis method was: comparing nodes. and The system identifies differences in the change patterns between the original video's features and those of the suspected infringing video. For example, in multiple consecutive nodes, the original video's background might appear as a solid-color photography studio, while the suspected infringing video's background might appear as kitchen tiles, even though the product's main features are highly similar. The system determines that background replacement has occurred at the "edges" between these consecutive nodes. Similarly, in multiple consecutive nodes, the original video's audio spectrum might match the brand's music, while the suspected infringing video's audio spectrum is completely different but matches the visual content, such as a new narration of the kitchen. In this case, the system determines that audio replacement has occurred at the edges between these consecutive nodes.
[0110] Finally, for each type of content modification identified, a malice weight is calculated, and these weights are then combined to obtain the total path malice score of the suspected infringing video. Preferably, the malice weight is predefined based on the degree of damage the modification causes to the value of the original video. For example, a modification that directly copies the main body of the product is assigned a malice weight of 0.9 (the maximum weight value is 1, between 0 and 1), because it steals the core creative work of the original video. A modification that replaces the background or audio is assigned a malice weight of 0.6; and modifications such as speed adjustment, color correction, and adding filters are assigned a malice weight of 0.3. Finally, the total path malice score of the video is obtained by weighted averaging of the malice weights of all edges in the same suspected infringing video.
[0111] In this application, the comprehensive infringement confidence score calculated for a suspected infringing video is: the weighted sum of the digital fingerprint similarity between the suspected infringing video and the aligned original video, the watermark detection strength, the total maliciousness of the path, the publisher's historical infringement score, and the supply chain risk value. Specifically, the digital fingerprint similarity is preferably the cosine similarity between the baseline digital fingerprint generated from the original video in step S1 and the fingerprint carried in the suspected infringing video; the watermark detection strength is the confidence score of detecting an original watermark in the suspected infringing video; the publisher's historical infringement score is a normalized score of the number and severity of past infringements by the account that published the suspected infringing video; and the supply chain risk value is the score, provided by the anti-counterfeiting and traceability system of the product supply chain, based on the received total maliciousness of the path, representing the surge in online counterfeit product complaints associated with the product.
[0112] In step S4, videos suspected of being infringing with a comprehensive infringement confidence level exceeding a preset threshold are ultimately determined to be infringing videos, and then... Figure 1 As shown, the identification method for internet short video original content protection monitoring provided in this application proceeds as follows: S5. Extract the product traceability code from the infringing video finally determined in step S4, synchronize it to the anti-counterfeiting traceability system of the supply chain in real time, and adjust the watermark embedding algorithm preference used in step S1 based on the offline high-counterfeit product feature information (common physical defect features of high-counterfeit products) fed back by the system.
[0113] Here is a brief explanation of the method for adjusting the watermark embedding algorithm preferences when creating original short videos in step S1: In step S5, after extracting the product traceability code from the finally determined infringing video, the anti-counterfeiting traceability system of the supply chain quickly matches the corresponding product from the product database based on this product traceability code. Then, it simultaneously matches all offline counterfeit products currently on the market (preferably high-quality counterfeits, which have richer features and higher feature similarity to the rights-protected product) from the counterfeit product database. The feature information (multimodal features) of these high-quality counterfeits is then input into the watermark embedding algorithm. The algorithm adjusts the robustness center of gravity of the watermark generation based on the input, enhancing the adversarial nature of the watermark-controlled baseline digital fingerprint. How the watermark embedding algorithm adjusts the robustness center of gravity and how the generation of the baseline digital fingerprint is controlled by the watermark to increase fingerprint adversarial nature has been explained in detail in step S1 and will not be repeated here.
[0114] In summary, the baseline digital fingerprint generated in this application is controlled by a dynamically generated digital watermark. This results in a fingerprint with embedded watermark protection strategies, making it more adversarial and laying the foundation for subsequent short-video infringement and source tracing at the fragment level. By analyzing and integrating cross-domain infringement hotspot maps of online infringement patterns and offline counterfeit data, the priority of video infringement analysis is adjusted, triggering incremental updates or enhanced indexes of original video features, thus improving the recall, precision, and response speed for detecting emerging threats in high-risk videos. By allocating dynamic weights in step A3, a balance is struck between the timeliness and accuracy of screening similar video fragments. Using the hotspot map as a dynamic control command for the original feature library, intelligent tilting of monitoring resources from uniform coverage to hotspot focus is achieved. Through repeated iterations of S1-S5, the robustness of the digital watermark becomes increasingly superior, making the fingerprint more adversarial.
[0115] This application also provides a method for intelligent anti-counterfeiting and traceability of the supply chain, including the following steps: L1, in response to the product production line off-line event, generates an encrypted unique traceability code and simultaneously controls the sensing device to collect multimodal physical features of the product packaging or the product itself under preset excitation conditions, forming the physical feature fingerprint of the product. Then, after binding the encrypted traceability code associated with the same product with the digest value of the physical feature fingerprint, it is uploaded to the blockchain node for evidence storage. In step L1, there are many existing methods for generating unique encrypted traceability codes for products, such as generating unique product QR codes, which will not be specifically explained. In this step, the collection of multimodal physical characteristics of newly produced products is performed under preset excitation conditions. These preset excitation conditions are not fixed but are optimized and adjusted based on the common physical defect characteristics of suspected counterfeit products fed back by the Internet short video original content protection monitoring system in step L4. The adjustment method includes the following steps: L111, obtain the latest updated list of infringing video source codes associated with this product from the Internet short video original protection monitoring and identification system; L112, analyze the common physical defect characteristics (multimodal characteristics) of the high-quality counterfeits in the analysis list. L113, based on the common physical defect characteristics of the analyzed high-quality counterfeit products, adjusts the preset excitation conditions for collecting multimodal physical characteristics of the new batch of products.
[0116] In step L112, the method for analyzing the common physical defect characteristics of high-quality counterfeit products associated with each traceability code in the list includes the following steps: L1121, retrieve the terminal verification records associated with each traceability code in the list, including historical verification images and / or videos uploaded by the user; For example, the list shows that the official product videos corresponding to the traceability codes PF1 and PF2 have been stolen in multiple short promotional videos. The traceability system searches its internal database for all terminal verification records associated with the PF1 and PF2 codes, focusing on records where verification failed or had low confidence levels, and also retrieves verification images and videos uploaded by the user at that time.
[0117] L1122 performs the same multimodal feature extraction as for genuine products on the retrieved terminal verification records, and calculates the feature deviation of each dimension. Taking serum as an example, we can extract the grid texture features of a designated area on the label of a suspected counterfeit product in an image, analyze the reflective intensity distribution of the laser logo on the bottle cap under designated side lighting, and measure the bending angle of the dropper head and the curvature of the bottle shoulder.
[0118] Then, the features extracted from these suspected counterfeit products are compared one by one with the physical fingerprint of the genuine batch of products recorded in the blockchain at the time of manufacture. The system calculates the deviation of each feature dimension. Feature comparison is feature similarity comparison, such as comparing the consistency of pixel grayscale values at the same location. Feature deviation is, for example, the ratio of the absolute value of the difference between the grayscale values of the two compared pixels to the grayscale value of that pixel in the genuine product. There are many existing methods for feature similarity comparison and feature deviation calculation, which will not be described in detail here.
[0119] L1123, the traceability system performs cluster analysis on the feature deviation vectors of all suspicious samples (feature deviation greater than the preset corresponding deviation threshold), and then dynamically defines common physical defect features based on the clustering results.
[0120] For example, suppose that approximately 80% of the samples are found to cluster in a significantly different region from the genuine product in the feature dimension of "distribution of reflective intensity of laser-engraved logo on bottle cap," while 60% of the samples show another specific deviation in the feature dimension of "bending angle of dropper head." Then, according to predefined rules, the system defines the common physical defect feature of the clustering result "80% of the samples cluster in a significantly different region from the genuine product in the feature dimension of 'distribution of reflective intensity of laser-engraved logo on bottle cap'" as follows: Under 75° side lighting, the reflective area of the bottle cap logo exhibits a diffuse halo, and the intensity of the core highlight point is less than 30% of the standard of the genuine product. This definition of common physical defect feature indicates that the laser engraving molds of these samples have insufficient precision or differences in material reflective properties. 75° side lighting is the excitation condition when collecting multimodal physical features from the genuine product. For the clustering result of "60% of the samples are in the feature dimension of "dropper head bending angle", the common physical defect features defined by the system according to predefined rules are, for example, the average angle between the silicone part of the dropper head and the glass tube is 142°, while the standard for genuine products is 135±1°, indicating that there are deviations in the assembly process or mold of these suspicious samples.
[0121] It should be noted that, since the suspicious samples are not fixed and the clustering results are not unique, the defined common physical defect characteristics are dynamic.
[0122] In step L113, the method of adjusting the preset incentive conditions for a new batch of goods produced according to dynamically defined common physical defect characteristics includes the following steps: L1131, for each dynamically defined common physical defect feature, a separate enhanced acquisition strategy is formulated. The formulation method is as follows: The historical excitation conditions carried in the common physical defect features are analyzed, and then adjusted according to a preset strategy to form the excitation correction conditions. For example, from the common physical defect feature mentioned above, "the bottle cap logo exhibits a diffuse halo in the reflective area under 75° side lighting, and the intensity of the core highlight point is less than 30% of the standard of the genuine product," the historical excitation condition "the bottle cap logo is sampled under 75° side lighting" is analyzed. Then, this historical excitation condition is adjusted according to a preset strategy. For example, in the strategy library, the preset adjustment strategy for the sampling angle of the bottle cap logo is: to add two additional angles ±5° to the original sampling angle for supplementary sampling. The adjusted excitation correction condition is to add two additional sampling angles of 70° and 80° to the original 75° sampling angle. Because the analysis shows that the defects, such as the halo, are more obvious when the angle changes slightly in high-quality counterfeits, the above adjustment strategy is preset.
[0123] Simultaneously, the discriminative features in the common physical defect characteristics are analyzed, and the corresponding acquisition parameters are adjusted according to the analysis results. Then, the adjusted excitation correction conditions and acquisition parameters are combined into an enhanced acquisition strategy as the adjusted preset excitation conditions.
[0124] For example, from the common physical defect characteristic that "the bottle cap logo exhibits a diffused halo in the reflective area under 75° side lighting, and the intensity of the core highlight is less than 30% of the standard for genuine products," the discriminant feature extracted is "the intensity of the core highlight is less than 30% of the standard for genuine products." Based on the preset discriminant feature and acquisition parameters, the mapping relationship is adjusted. For instance, when the discriminant feature is "the intensity of the core highlight is less than 40% of the standard for genuine products," the exposure time of the high-definition camera when capturing the reflective effect is shortened by n seconds, and the contrast is increased (adjusted by d1), to ensure that the intensity of the core highlight in the newly produced batch of products can be captured and quantified more strongly. Here, shortening the exposure time by n seconds and increasing the contrast adjustment by d1 are the acquisition parameters to be adjusted.
[0125] In step L1, the method for collecting multimodal physical features of the product packaging or the product itself to form the product's physical feature fingerprint, based on the adjusted preset excitation conditions, is briefly described below: For example, a macro camera is controlled to capture microscopic texture images of a specified area on the surface of a product packaging under a specified illumination angle in a preset excitation condition; a light source of a specified wavelength in a preset excitation condition is controlled to be excited, and then fluorescent or light-changing response images of the anti-counterfeiting area of the product packaging are captured, and then the physical features of the captured image area are extracted and formed into a physical feature vector, which serves as the physical feature fingerprint of the product.
[0126] After adjusting the preset incentive conditions for the new batch of products following step L1, the intelligent anti-counterfeiting and traceability method for the supply chain provided in this application is as follows: Figure 1 As shown, proceed to the following steps: L2 synchronizes the common physical defect features of counterfeit products dynamically defined during the preset excitation condition adjustment in step L1 to the identification system for monitoring the original protection of internet short videos. When generating original short videos for the new batch of products and storing them in the identification system, this serves as the basis for the identification system's preference for the watermark embedding algorithm that dynamically generates digital watermarks for the original short videos.
[0127] The specific details of how the recognition system adjusts the watermark embedding algorithm preferences have been explained above and will not be repeated here.
[0128] L3, update the original feature library, and then, based on the benchmark digital fingerprint and digital watermark of the original short videos of the new batch of products, screen the real-time data stream data and the platform's existing video data of the target monitoring platform for suspected infringing videos; It should be noted that after the watermark embedding algorithm adjusts the watermark embedding preferences and embeds the digital watermark into the original short video of the product in the new batch, the baseline digital fingerprint, digital watermark embedding key, hash value of the original short video, and supply chain product identifier associated with the video generated for the original short video are stored together in the original feature library. At this point, the update of the original feature library is completed.
[0129] The Internet short video original content protection and monitoring system uses the update of the original feature database as an instruction. Based on the benchmark digital fingerprint and digital watermark dynamically generated for the original short video, it screens suspected infringing videos. The screening method has been explained in detail above and will not be repeated here.
[0130] After identifying suspected infringing videos through step L3, the intelligent anti-counterfeiting and traceability method for the supply chain transitions to the following steps: L4 retrieves the terminal verification records associated with the traceability codes carried in each suspected infringing video that has been screened, and then performs the same multimodal feature extraction as the new batch of the product to form the physical feature fingerprint of the suspected infringing product. For example, the traceability codes carried in the suspected infringing videos a and b are sy123 and sy1234, respectively. The terminal verification records associated with sy123 include, for example, the product QR code image pa1 and the product outer packaging image pa2 taken by users a1 and a2 during their historical verification. The terminal verification records associated with sy1234 include, for example, the product QR code image pb1 and the product outer packaging image pb2 taken by users b1 and b2 during their historical verification.
[0131] In step L4, the traceability system performs the same multimodal feature extraction on pa1, pa2, pb1, and pb2 as on the newly produced product, and then forms the extracted multimodal features into vectors as the physical feature fingerprints of the corresponding suspected infringing products. For example, the multimodal feature vectors extracted from pa1 and pa2 constitute the physical feature fingerprints of the suspected infringing products associated with the suspected infringing video a.
[0132] L5, compare the physical feature fingerprint of the suspected infringing product with the physical feature fingerprint generated for the new batch of the same product in step L1. If the comparison is successful, the suspected infringing product is determined to be genuine. If the comparison fails, the suspected infringing product is determined to be an infringing product.
[0133] It should be noted that there are many existing methods for comparing the similarity of feature fingerprints, which will not be discussed in detail here.
[0134] In summary, the intelligent anti-counterfeiting and traceability system for the supply chain provided in this application synchronizes the common physical defect characteristics of counterfeit products dynamically defined when adjusting preset incentive conditions to the identification system for monitoring the original protection of internet short videos. The identification system, in turn, synchronizes the product traceability code extracted from each identified infringing video to the anti-counterfeiting and traceability system to dynamically define the common physical defect characteristics of counterfeit products. The two systems are linked and respond dynamically, realizing the proactive verification of the authenticity of the product and the dynamic updating of anti-counterfeiting strategies. This significantly increases the difficulty of counterfeiting products and can promptly and accurately detect counterfeit products in each batch of products that have been taken offline, greatly reducing the difficulty of combating counterfeits in short videos.
[0135] It should be stated that the above-described specific embodiments are merely preferred embodiments and technical principles applied in this application. Those skilled in the art should understand that various modifications, equivalent substitutions, and variations can be made to this application. However, such variations, as long as they do not depart from the spirit of this application, should be within the scope of protection of this application. Furthermore, some terminology used in this application's specification and claims is not limiting but merely for ease of description.
Claims
1. A method for identifying and protecting original content in short internet videos, characterized by the following steps: include: S1, in response to the request to add the original short video to the database, generate the baseline digital fingerprint of the original short video, and dynamically generate a digital watermark and embed it into the video bitstream; S2, the baseline digital fingerprint, digital watermark embedding key, original short video hash value, and video-associated supply chain commodity identifier generated in step S1 are stored together in the original feature library. The indexing strategy and content update frequency of the original feature library are dynamically adjusted by the cross-domain infringement hotspot map generated based on the output of the causal graph model in step S4. S3, the real-time video stream data of the target monitoring platform is segmented using a sliding window, and the first similar video segment is screened based on fingerprint features; the existing video data of the platform is screened for the second similar video segment through watermark blind detection, and the fingerprint feature screening and watermark blind detection screening share the original feature library; S4. For the similar video segments identified as suspected infringing videos in step S3, construct a causal graph model based on the time axis, analyze the frame-level correspondence and content modification type between the suspected infringing videos and the original videos, and calculate the comprehensive infringement confidence level. S5. Extract the product traceability code from the infringing video finally determined in step S4, synchronize it with the anti-counterfeiting traceability system of the supply chain in real time, and adjust the watermark embedding algorithm preference used in step S1 based on the offline high-quality counterfeit product feature information fed back by the system.
2. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 1, characterized in that, In step S1, the method for generating the baseline digital fingerprint of the original short video includes the following steps: A1. Based on the embedding domain, embedding strength, and error correction coding scheme of the digital watermark dynamically generated for the current original short video, the robustness center of gravity targeted by the digital watermark is analyzed. A2, assess the vulnerability of the digital watermark to protecting the current original short video; A3. Based on the robustness centroid analyzed in step A1 and the vulnerability assessment results obtained in step A2, visual features, audio features, text features and time-series features are dynamically allocated as weights when generating the baseline digital fingerprint using a preset strategy matrix. A4. Using the dynamic weight strategy output in step A3 as prior knowledge, the network is injected into the pre-trained attention fusion network to drive the network to autonomously extract and fuse features from the multimodal data of the original short video, and finally generate the benchmark digital fingerprint.
3. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 2, characterized in that, Step A1, the method for resolving the robustness centroid of the dynamically generated digital watermark, includes the following steps: A11, read the watermark embedding domain parameters dynamically configured for the current original short video, then perform robust feature mapping, and parse out the primary robust features; A12, collaboratively analyze the embedding strength and error correction coding scheme of the embedding domain and combine the analysis results of the two to quantify the reliability of the digital watermark; A13, combining the comprehensive analysis results of the parsing results of steps A11-A12, the robust center of gravity targeted by the digital watermark is generated through the reverse attack mode.
4. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 3, characterized in that, In step A13, the method for generating the robust centroid targeted by the digital watermark through a reverse attack mode includes the following steps: A131, Match possible attack scenarios from the operation library that can destroy the digital watermark with the watermark features recorded in the comprehensive analysis results; A132, perform feature sensitivity mapping on each attack mode deduced in step A131, and then weight and fuse the mapped sensitivity features to form the robustness center of gravity targeted by the digital watermark.
5. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 2, characterized in that, In step A2, the vulnerability assessment method includes the following steps: A21 generates a list of weaknesses in content amplification through a constructed content and watermark interaction model, specifically: The original short video is analyzed for its content characteristics to obtain the content characteristics of the original short video. The content characteristics obtained through content characteristic analysis are compared with the premise of digital watermark generation to identify contradictions in watermark generation and synthesize them into a list of content amplification weaknesses. A22, retrieve the historical attack vector formed by the first n historical attacks against the digital watermark in attack order, and combine it with the content amplification vulnerability list to calculate the threat coefficient of each type of historical attack in the vector to the current original short video, and generate a quantitative threat mapping table. A23. Randomly select key frames from the original short video as sample frames, then apply the attack from the quantized threat mapping table to each sample frame, then run the watermark decoding process, and calculate the proportion of frames where the watermark is successfully decoded under different attacks. Finally, weight and fuse the proportions of each frame to obtain the average survival probability, which serves as an evaluation index value for the vulnerability of the digital watermark to protect the original short video.
6. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 1, characterized in that, Step S2, the method for generating the cross-domain infringement hotspot map and dynamically adjusting the indexing strategy and content update frequency of the original feature library includes the following steps: S21. Extract the triplet information corresponding to each infringement identification case from the infringement identification cases analyzed by the causal graph model within a specified time period, including the infringing content, counterfeiting method and related products. S22, Based on the extracted triplet information, perform content clustering, counterfeiting method clustering, and related product clustering on each infringement identification case, and then generate the cross-domain infringement hotspot map after mining the relationship between each element in the triplet information. S23, extract original short videos associated with infringement hotspots that are stored in the original feature library from the cross-domain infringement hotspot map, and add a "high-risk" tag to the index of the baseline digital fingerprint corresponding to the extracted video. At the same time, add an index dimension to the original feature library; and increase the scanning frequency of feature updates for original short videos marked as "high-risk".
7. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 2, characterized in that, In step S3, the method for screening the first similar video segment based on fingerprint features in the real-time video stream data of the target monitoring platform includes the following steps: B1, Extract basic fingerprint features from real-time segmented video captured using a sliding window, the basic fingerprint features including at least one weight dynamically assigned in step A3; B2, compare the basic fingerprint features extracted in step B1 with the benchmark digital fingerprints stored in the latest updated original feature library. If the comparison is successful, the real-time segmented video will be marked as the first similar video segment; If the comparison fails, it is determined that the screening of similar segments in the real-time segmented video has failed.
8. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 1, characterized in that, Step S3, the method for screening second similar video segments in the platform's existing video data through watermark blind detection, includes the following steps: C1. Decode the watermark features and basic fingerprint features from the existing video, and perform a reverse query in the latest updated original feature library. If the query is successful, proceed to step C2; otherwise, determine that the existing video is not the second similar video segment. C2: Extract the original short videos successfully retrieved in the reverse query of step C1, and parse the baseline digital fingerprint embedded in the extracted original short videos. Then, perform similarity matching with the baseline digital fingerprint of the existing videos decoded in step C1. If a match is successful, the existing video will be marked as the second similar video segment; If the match fails, it is determined that the existing video is not the second similar video segment.
9. The identification method for monitoring and protecting the original content of short videos on the Internet according to claim 1, characterized in that, In step S4, the method for constructing the causal graph model includes the following steps: S41, after aligning the suspected infringing video with the original video for suspected infringing segments, extract key frame pairs within the alignment interval. The key frame pairs include the aligned key frames of the original video and the key frames of the suspected infringing video. S42, define the nodes of the causal graph as each keyframe extracted in step S41, and the node attributes include the multimodal feature vector of the keyframe and the inter-frame similarity score of the keyframe pair; define the edges of the causal graph as the inter-frame transition time order. S43, Analyze the properties of the edges in the causal graph to identify the types of content modifications made to the original video by the suspected infringing video; S44, For each type of content modification identified, calculate the malice weight, and then combine them into the total malice score of the suspected infringing video path; In step S4, the calculated comprehensive infringement confidence level is: the weighted sum of the digital fingerprint similarity between the suspected infringing video and the aligned original video, the watermark detection strength, the total malice of the path, the publisher's historical infringement score, and the supply chain risk value.
10. A recognition system for monitoring and protecting the original content of short videos on the Internet, implementing the recognition method as described in any one of claims 1-9, characterized in that, include: The fingerprint-watermark collaborative generation module, in response to the original short video's database entry request, generates a baseline digital fingerprint of the original short video and dynamically generates a digital watermark and embeds it into the video bitstream. The original feature library update module, connected to the fingerprint-watermark collaborative generation module, is used to store the generated baseline digital fingerprint, digital watermark embedding key, original short video hash value, and video-associated supply chain commodity identifier in the original feature library. The indexing strategy and content update frequency of the original feature library are dynamically adjusted according to the cross-domain infringement hotspot map. The suspected infringement video screening module is connected to the original feature library update module. It is used to segment the real-time video stream data of the target monitoring platform using a sliding window and screen the first similar video segment based on fingerprint features. For the platform's existing video data, it screens the second similar video segment through watermark blind detection. The fingerprint feature screening and the watermark blind detection screening share the original feature library. The causal graph model construction module is connected to the suspected infringement video screening module. It is used to construct a causal graph model based on the time axis for similar video segments that are screened as suspected infringement videos, and generate the cross-domain infringement hotspot map based on the output of the causal graph model. The infringement confidence calculation module is connected to the causal graph model construction module. It is used to analyze the frame-level correspondence and content modification type between the suspected infringing video and the original video based on the constructed causal graph model, and to calculate the comprehensive infringement confidence. The watermark embedding algorithm adjustment module, connected to the infringement confidence calculation module and the fingerprint-watermark collaborative generation module, is used to extract the product traceability code from infringing videos with a comprehensive infringement confidence score higher than the threshold, synchronize it to the anti-counterfeiting traceability system of the supply chain in real time, and adjust the watermark embedding algorithm preference according to the offline high-counterfeit product feature information fed back by the system, and then notify the fingerprint-watermark collaborative generation module.