Audiovisual work material traceability method and system based on blockchain
By segmenting audiovisual works into sub-blocks and generating NFT identifiers using blockchain and AI technologies, and dynamically constructing logical sub-chains, the problem of long copyright confirmation cycles and difficulty in infringement identification in existing copyright management systems is solved, achieving efficient and accurate copyright protection.
Patent Information
- Application Number
- CN202511758235.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-11-27
AI Technical Summary
Existing digital rights management systems rely on centralized institutions, resulting in long copyright confirmation cycles, high costs, and difficulty in covering scenarios where creation occurs immediately. Furthermore, they cannot effectively identify covert infringement behaviors such as fragment-level reuse, editing, or semantic substitution.
A blockchain-based method for tracing audiovisual works is adopted. An AI segmentation engine is used to segment audiovisual works into sub-blocks based on semantic relevance, generate hash values and metadata, dynamically construct logical sub-chains and generate NFT identifiers, and use smart contracts to realize copyright response and permission management.
It enables more granular content tracking capabilities, improves the accuracy and efficiency of copyright protection, and solves pain points such as easy copying, difficulty in tracing, and ambiguity in rights confirmation within multimedia content, providing a highly reliable and efficient copyright protection solution for the digital content industry.
Smart Images

Figure CN121561880B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital content security management technology, and in particular to a blockchain-based method and system for tracing the source of audiovisual works. Background Technology
[0002] With the rapid development of the digital content industry, the creation, dissemination, and consumption patterns of audiovisual works worldwide are undergoing profound changes. The widespread adoption of short video platforms, streaming services, social networks, and metaverse applications has led to a highly decentralized, high-frequency, and fragmented nature of audiovisual content generation and distribution. Creators are no longer limited to professional institutions; a large number of individual users can produce high-quality content through smart devices, and the maturity of AI-generated content (AIGC) technology has further accelerated this trend, resulting in exponential growth in content output.
[0003] In the current technological environment, the confirmation of copyright and the determination of infringement for audiovisual works heavily rely on centralized institutions for registration and manual review, resulting in long confirmation cycles, high costs, and difficulty in covering scenarios where creation occurs immediately. Currently, most existing digital copyright management systems are based on centralized architectures, relying on third-party institutions for content registration, infringement monitoring, and rights enforcement. This not only poses a single point of failure risk but also leads to a lack of trust due to opaque information and inconsistent rules. Although blockchain technology has been introduced into the copyright field in recent years, and some systems have attempted to use hash-based on-chain methods for content notarization, most only generate a single fingerprint for complete files, failing to effectively identify covert infringement behaviors such as fragment-level reuse, editing, or semantic substitution, making it difficult to quickly identify and curb infringement. Summary of the Invention
[0004] To address the aforementioned technical issues, this application provides a blockchain-based method and system for tracing the source of audiovisual works.
[0005] Firstly, this application provides a blockchain-based method for tracing the source of audiovisual works, employing the following technical solution: A blockchain-based method for tracing the source of audiovisual works, the method comprising: The original audiovisual material is input into the AI segmentation engine, divided into several sub-blocks according to semantic relevance, and the hash value and metadata of each sub-block are extracted. Based on the semantic correlation between sub-blocks, at least one logical sub-chain is dynamically constructed, and the root hash value of each logical sub-chain is generated; The root hash value is bound to metadata to generate an NFT identifier, and the NFT identifier and associated data are anchored to the blockchain network; In response to a user query request, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and compare the hash deviation results between the user-submitted material and the hash recorded on the blockchain through the AI verification module. The hash deviation result triggers a smart contract to execute a predefined copyright response action. Based on the execution result of the smart contract, the permission status in the decentralized identity system is updated and synchronized to the blockchain network.
[0006] By adopting the above technical solutions, a full-lifecycle audiovisual work traceability system has been constructed, integrating content segmentation, structural modeling, rights confirmation and evidence storage, authenticity verification, and access control. Compared with traditional single hash-based on-chain or centralized copyright registration models, the technical solution of this application achieves more granular content tracking capabilities at the semantic level, adapts to complex creative structures with dynamic subchains, improves verification accuracy through the collaboration of AI and cryptography, and realizes rule automation and identity autonomy through smart contracts and DID. This fundamentally solves the pain points of easy copying, difficulty in traceability, and ambiguous rights confirmation in multimedia, providing a system-level solution with both technological innovation and engineering feasibility for copyright protection and value circulation in the digital content industry.
[0007] Secondly, this application provides a blockchain-based audiovisual work material traceability system, which adopts the following technical solution: A blockchain-based audiovisual work material traceability system, characterized in that the traceability system comprises: The material segmentation module is used to input the original audiovisual work material into the AI segmentation engine, segment it into several sub-blocks according to semantic relevance, and extract the hash value and metadata of each sub-block; The logical sub-chain construction module is used to dynamically construct at least one logical sub-chain based on the semantic correlation between sub-blocks, and generate the root hash value of each logical sub-chain; The NFT identifier generation module is used to bind the root hash value with metadata to generate an NFT identifier, and to anchor the NFT identifier and associated data to the blockchain network; The hash deviation comparison module is used to respond to user query requests, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and compare the user-submitted material with the hash deviation results recorded on the blockchain through the AI verification module. The copyright response execution module is used to trigger the execution of predefined copyright response actions by the smart contract based on the hash deviation result. Based on the execution result of the smart contract, the permission status in the decentralized identity system is updated and synchronized to the blockchain network.
[0008] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.
[0009] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect.
[0010] In summary, this application achieves at least one of the following beneficial technical effects: By constructing a full-link audiovisual material traceability system encompassing "semantic segmentation—dynamic sub-chain modeling—NFT ownership confirmation—AI collaborative verification—intelligent response," it realizes a deep integration from content structure analysis to automated copyright governance, improving the granularity, accuracy, and execution efficiency of digital copyright protection. In practical applications, it not only solves long-standing pain points such as easy copying, difficulty in tracing, ambiguous ownership confirmation, and delayed response within multimedia, but also provides highly reliable, efficient, and adaptive infrastructure support for the circulation of NFT digital assets and AIGC content, as well as cross-platform copyright collaboration, promoting the evolution of the digital content ecosystem towards automation, decentralization, and value programmability. Attached Figure Description
[0011] Figure 1 This is a schematic diagram of the first process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application.
[0012] Figure 2 This is a schematic diagram of the second process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application.
[0013] Figure 3 This is a schematic diagram of the third process of a blockchain-based method for tracing the source of audiovisual works, which is one embodiment of this application.
[0014] Figure 4 This is a schematic diagram of the fourth process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application.
[0015] Figure 5 This is a schematic diagram of the fifth process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application.
[0016] Figure 6 This is a schematic diagram of the sixth process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application.
[0017] Figure 7 This is a schematic diagram of the seventh process of a blockchain-based audiovisual work material tracing method according to one embodiment of this application. Detailed Implementation
[0018] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0019] This application discloses a blockchain-based method for tracing the source of audiovisual works.
[0020] Reference Figure 1 A blockchain-based method for tracing the source of audiovisual works, the method including: Step S101: Input the original audiovisual work material into the AI segmentation engine, segment it into several sub-blocks according to semantic relevance, and extract the hash value and metadata of each sub-block; Traditional audiovisual materials typically exist in the form of continuous streams, lacking inherent data boundary markers, making it difficult to support refined copyright tracking and content verification.
[0021] This application's embodiments introduce an AI segmentation engine and utilize deep learning models (such as convolutional neural networks CNN) to perform visual semantic analysis on video frame sequences, identifying key visual events such as scene switching, camera movement, and the appearance and disappearance of objects, thereby accurately dividing sub-blocks with independent semantic units.
[0022] For example, in a movie clip, AI can identify different scenes such as "city street scenes," "character dialogues," and "action fights," and accordingly divide the original video into several logically independent sub-segments. Simultaneously, the system performs multimodal feature extraction on each sub-segment, fusing visual features (such as color histograms and optical flow information), audio features (such as MFCC coefficients and volume changes), and potential textual information (such as subtitle content) to form a high-dimensional feature vector. This vector is then used to generate a unique and irreversible digital fingerprint, i.e., the sub-segment hash value, through cryptographic hash functions such as SHA-256, ensuring that any minor content tampering will result in a significant change in the hash value. Metadata further expands the information dimensions of this fingerprint system, including not only technical parameters such as timestamps (start and end times accurate to milliseconds) and spatial coordinates (GPS positioning or camera coordinates), but also the creator's decentralized identity (DID). This DID, built on the W3C standard, enables autonomous control and verifiable assertion of identity without relying on a centralized institution.
[0023] Understandably, this semantic-based rather than time-based segmentation method makes the segmentation results closer to the actual creation logic of the content and provides a semantic basis for subsequent correlation analysis.
[0024] Step S102: Based on the semantic correlation between sub-blocks, dynamically construct at least one logical sub-chain and generate the root hash value of each logical sub-chain; The system first calculates the semantic similarity matrix between all sub-blocks, typically using cosine similarity or Euclidean distance to measure the proximity of their multimodal feature vectors. Then, it employs clustering algorithms (such as DBSCAN or hierarchical clustering) to group highly similar sub-blocks into the same logical sub-chain. This step overcomes the limitations of linear or fixed topology in traditional hash chains or Merkle trees, introducing a dynamic clustering mechanism to adapt to the organizational patterns of different types of audiovisual content.
[0025] For example, in a documentary, multiple scattered shots about "polar ecology," though not consecutive on the timeline, are aggregated into an independent sub-chain due to their high thematic relevance; similarly, the repetition of the chorus in a music video can form a closed-loop semantic chain. Each logical sub-chain employs a Merkle tree structure for hash aggregation: the hashes of each sub-block serve as leaf nodes, and are merged pairwise layer by layer to generate the parent node hash, ultimately yielding a unique root hash value.
[0026] In this embodiment, this design not only achieves high efficiency in data integrity verification (verifying only the path rather than all data), but more importantly, it endows the content structure with programmability. Different material types (such as news, advertisements, and films / TV dramas) can be configured with different clustering thresholds and chain construction rules, enabling the system to be applicable across domains. Furthermore, the dynamism is reflected in the fact that the sub-chain structure can be adjusted in real time with the addition of new materials or re-editing, avoiding the rigidity of static structures in complex creative processes, thereby ensuring a high degree of consistency between the traceability path and the actual creative intent.
[0027] Step S103: Bind the root hash value with metadata to generate an NFT identifier, and anchor the NFT identifier and associated data to the blockchain network; In this context, NFTs (Non-Fungible Tokens) are not only a form of asset representation but also a structured container carrying complex copyright information. The system embeds the root hash of the logical subchain as the core anchor into the NFT's metadata field, while also attaching the IPFS (InterPlanetary File System) storage address of the original material. This achieves separate storage of content and metadata, ensuring permanent accessibility. IPFS provides content addressing to prevent link failures, while the blockchain ensures the immutability of the metadata. Copyright terms (such as scope of use, license period, and revenue sharing ratio) are also encoded in machine-readable JSON-LD format and written into the NFT, forming the basis for smart contract execution.
[0028] Furthermore, to enhance cross-platform interoperability, the system deploys NFTs to mainstream public chains (such as Ethereum and Polygon) via cross-chain bridges, generating cross-chain anchoring proofs between the source and target chains, and ensuring state consistency using light client verification or oracle mechanisms. This process essentially completes the mapping from "physical content" to "on-chain digital credentials," enabling each logical sub-chain to obtain a globally unique, verifiable, and transferable identity, greatly improving the automation level and legal validity of copyright registration.
[0029] Step S104: In response to the user's query request, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and compare the user-submitted material with the hash deviation result recorded on the blockchain through the AI verification module. Specifically, the query process begins with the user uploading the material to be verified or a specified content fragment. The system quickly locates the logical subchain to which it belongs and the corresponding NFT through local feature matching. Subsequently, the complete Merkle path of the subchain (including the hash sequence of sibling nodes) is extracted from the blockchain network, and the verification path is reconstructed locally.
[0030] In this process, the AI verification module plays a crucial role: unlike the traditional method of comparing hash values one by one, the system uses time-series models such as LSTM (Long Short-Term Memory Network) to analyze the hash change patterns of sub-block sequences and capture the differences between normal editing (such as clipping and color correction) and malicious tampering.
[0031] For example, legitimate color correction typically manifests as small, continuous fluctuations in hash values, while keyframe replacement leads to local hash mutations. LSTM can learn the dynamic characteristics of historical hash sequences and output a deviation index to quantify the degree of deviation between the current content and the original record. Once this index exceeds a preset deviation threshold, the system determines that there is potential tampering and automatically generates a visual traceability report containing the location of the tampering, the scope of impact, and a confidence score to assist human decision-making. This dual verification mechanism, which integrates cryptographic verification and AI semantic understanding, significantly reduces the false alarm rate and is particularly suitable for handling copyright disputes involving reasonable reprocessing scenarios.
[0032] Step S105: Trigger the smart contract to execute a predefined copyright response action based on the hash deviation result. Update the permission status in the decentralized identity system and synchronize it to the blockchain network based on the smart contract execution result.
[0033] Among them, smart contracts, as automated programs deployed on the blockchain network, respond instantly to the verification results according to pre-coded rules.
[0034] For example, if the deviation stems from unauthorized use or content tampering, the contract will initiate an infringement handling process: on the one hand, freezing the transfer, licensing, and other operational permissions of the relevant NFTs to prevent the continued circulation of infringing assets; on the other hand, packaging on-chain evidence (including the original hash path, verification logs, and fingerprints of user-submitted materials) to form an undeniable electronic evidence package, which can be retrieved by the judiciary or accessed through the arbitration system of the digital copyright trading platform.
[0035] Conversely, if a user request falls within the scope of compliant authorization (such as on-demand streaming or derivative works licenses), the smart contract automatically executes the billing logic, calculates fees based on the copyright terms in the NFT metadata (such as pay-per-use, subscription, or revenue sharing), completes on-chain payment settlement, and issues a digitally signed authorization certificate. This response mechanism achieves decentralization and automation of copyright management, eliminates the delays and costs of traditional intermediaries, and ensures the transparency and non-interference of rule enforcement.
[0036] Furthermore, as a user's self-identified carrier in the Web3 environment, the DID document stores information such as public keys, verification methods, and server endpoints. Whenever a copyright status change occurs (such as granting a license, revoking the right to use, or marking an infringement), the system will generate a corresponding status update statement, sign it, and write it into the creator's or user's DID document. The system will then synchronize the key status to multiple blockchain networks through a cross-chain oracle mechanism to ensure the global consistency of permission information.
[0037] For example, if a creator is penalized for copyright infringement, a "Restricted Account" label will be added to their DID document, restricting their automatic rights confirmation permissions for future uploaded content. Meanwhile, authorized users will have a "Permitted" credential added to their DID, allowing third-party platforms to quickly verify the legitimacy of their use. This dynamic permission management mechanism based on DID not only enhances the system's security and controllability but also provides a scalable framework for digital identity governance in a human-machine collaborative environment.
[0038] The above embodiments construct a full-lifecycle audiovisual work tracing system integrating content segmentation, structural modeling, rights confirmation and storage, authenticity verification, and access control. Compared to traditional single hash-based on-chain or centralized copyright registration models, the technical solution of this application achieves finer-grained content tracking capabilities at the semantic level, adapts to complex creative structures through dynamic subchains, improves verification accuracy through the collaboration of AI and cryptography, and automates rules and achieves identity autonomy through smart contracts and DID. This fundamentally solves the pain points of easy copying, difficulty in tracing, and ambiguous rights confirmation in multimedia content, providing a system-level solution with both technological innovation and engineering feasibility for copyright protection and value circulation in the digital content industry.
[0039] Reference Figure 2As one implementation of step S101, the steps of inputting the original audiovisual work material into the AI segmentation engine, segmenting it into several sub-blocks according to semantic relevance, and extracting the hash value and metadata of each sub-block include: Step S201: Receive the original audiovisual work materials and associated creator identity data; The original audiovisual materials usually exist in the form of digital files, such as container formats like MP4, MOV, or AVI, containing video streams, audio streams, and possibly subtitles and metadata tracks. These data are continuously distributed in the time dimension, lack inherent structured boundaries, and are difficult to use directly for refined copyright management or blockchain evidence storage.
[0040] Simultaneously, the system receives creator identity data. This data is not a traditional username or email address, but an encrypted identifier generated based on the Decentralized Identifier (DID) standard. It conforms to the W3C DID specification and possesses verifiability and self-control characteristics supported by Public Key Infrastructure (PKI). This design ensures that identity information does not rely on a centralized registration authority and that its ownership can be verified through a digital signature mechanism.
[0041] Understandably, the simultaneous input of source material and identity data provides a premise for the coupling of "content-subject" dual elements in all subsequent processing steps. This ensures that every piece of metadata generated naturally carries a credible identity anchor for the creator, fundamentally avoiding disputes over content ownership. Furthermore, this step implicitly includes a data integrity verification mechanism, such as using TLS encryption to ensure the confidentiality and integrity of data during transmission, preventing the original source material from being tampered with or the identity information from being misused before entering the system. This establishes a secure and reliable starting point for the entire preprocessing chain.
[0042] Step S202: The semantic relationships of the material are analyzed by the AI segmentation engine, and the material is divided into several sub-blocks; Traditional video segmentation methods often rely on fixed time intervals or simple frame difference detection, failing to accurately reflect the actual narrative logic or creative intent of the content. This application's embodiments employ a multimodal deep learning model working collaboratively to achieve accurate identification of complex semantic boundaries in audiovisual signals.
[0043] Specifically, the system utilizes convolutional neural networks (CNNs), particularly visual analysis models based on deep residual architectures such as ResNet-50, to extract features frame by frame from video frame sequences, identifying keyframes for scene cuts or shot transitions. These models capture low-level visual features such as edges, textures, and color distribution in the spatial dimension through convolutional kernels, and abstract high-level semantic concepts (such as "indoor dialogue" or "outdoor chase") through deep networks, thereby determining whether the visual content has undergone fundamental changes.
[0044] Meanwhile, the audio stream is temporally modeled using Long Short-Term Memory (LSTM) networks, especially Bidirectional LSTM (Bi-LSTM), to analyze continuity breaks in the audio signal, such as pauses in speech, changes in background music, and variations in ambient sound. For example, in an interview video, when the host finishes asking a question and the guest begins to answer, there is a noticeable silence gap and tone change in the audio signal. Bi-LSTM can capture this temporal pattern and mark it as a semantic segmentation point. The results from the visual boundary detection unit and the audio continuity analysis unit are not used independently, but are combined through a fusion strategy (such as weighted voting or attention mechanisms) to determine the final segmentation position, ensuring that the sub-blocks maintain semantic coherence in both the audiovisual and visual channels.
[0045] Understandably, this AI-based multimodal semantic segmentation method makes the generated sub-blocks no longer mechanical time slices, but content units with independent narrative functions, such as a complete dialogue round, an independent background music segment, or a complete action scene, providing semantically reasonable processing granularity for subsequent feature extraction and copyright tracking.
[0046] Step S203: Extract the multimodal feature data of each sub-block and generate the sub-block hash value; The goal of this step is to transform each semantic sub-block into a set of high-dimensional, computable, and interference-resistant digital feature vectors, and further compress them into unique and irreversible hash identifiers.
[0047] For video sub-blocks, the system extracts SIFT (Scale-Invariant Feature Transform) feature vectors from keyframes (such as I-frames or representative frames). This algorithm detects keypoints by performing Gaussian difference (DoG) on the image and describes its gradient direction and intensity distribution in scale space. It has good rotation, scaling and illumination invariance and can maintain feature stability under different encoding or slight image deformation conditions.
[0048] For audio sub-blocks, Mel-Frequency Cepstral Coefficients (MFCCs) are extracted. This feature simulates the characteristics of human auditory perception. The audio signal is mapped onto the Mel frequency standard after Fourier transform, and then the cepstral coefficients are extracted through discrete cosine transform, which effectively represents the timbre, rhythm and intonation features of speech or music.
[0049] The aforementioned visual and auditory feature vectors are then fed into a feature fusion engine, which generates a unified multidimensional feature matrix through weighted concatenation or deep network embedding. This matrix not only preserves the independent information of each modality, but also captures cross-modal correlations (such as the synchronization of a person's speech in the picture with their voice) through a fusion mechanism.
[0050] Building upon this foundation, the system uses the multi-dimensional feature matrix as input. First, it generates an initial hash value using the SHA-256 cryptographic hash algorithm. This algorithm boasts strong collision resistance and avalanche effect, ensuring that even minor changes in features result in drastic shifts in the hash value. To further enhance security and uniqueness, the system introduces a secondary hashing mechanism: the initial hash value is concatenated with creator identity data (such as DID hash) and a precise timestamp (UTC standard), and then subjected to another SHA-256 operation to generate the final unique hash value. This dual-binding design not only increases the hash value's sensitivity to the content itself but also forms an inseparable whole with the creator's identity and creation time, significantly enhancing the difficulty of forgery and the ability to trace the source.
[0051] Step S204: Encapsulate the sub-block hash value, multimodal feature data, and creator identity data into structured metadata.
[0052] This step is not simply data packaging, but rather constructing a metadata structure with interoperability and resolvability based on the data access specifications of the blockchain system. The system uses hash values as primary keys, multimodal feature data is organized in key-value pairs (such as "visual_features":[sift_vector],"audio_features":[mfcc_vector]), and creator identity data is embedded in DID format and attached with a digital signature to verify its authenticity.
[0053] In addition, the metadata also includes location identifiers (such as IPFS content address CID) and time codes (ISO 8601 format timestamps) required for blockchain storage. The former is used to point to the actual storage location of the original material or feature data, realizing the combination of off-chain large file storage and on-chain metadata anchoring, while the latter provides time sequence basis for all operations and supports subsequent copyright priority determination.
[0054] The entire metadata structure is typically encoded in JSON-LD (JSON for Linked Data) format, maintaining readability while supporting automatic machine parsing under the Semantic Web standard, facilitating cross-platform system calls and smart contract reading. The metadata wrapper, as the module performing this operation, also possesses adaptive capabilities, dynamically adjusting field weights or encryption strategies based on the content type (such as movies, advertisements, and short videos). For example, it can enable zero-knowledge proof fields in high-security scenarios and add licensing agreement terms in open and shared scenarios. This structured encapsulation method makes each sub-block a self-contained "digital asset package," which can be used for subsequent on-chain notarization or independently participate in authorization, transaction, or verification processes.
[0055] The above implementation achieves a systematic transformation of audiovisual works from raw data into verifiable digital assets. Its innovation lies in several aspects: First, it employs a multimodal semantic segmentation mechanism combining CNN and LSTM, overcoming the limitations of traditional single-modal segmentation and improving the semantic accuracy of sub-block partitioning; second, it constructs a robust digital fingerprint that combines content sensitivity and identity binding through SIFT and MFCC feature fusion and a dual hashing mechanism; finally, the encapsulation design based on DID and structured metadata meets the stringent requirements of the blockchain environment for data integrity, traceability, and interoperability. The overall solution not only provides a high-quality data foundation for subsequent copyright confirmation, infringement detection, and smart contract execution, but also achieves a deep integration of AI analysis capabilities and blockchain evidence storage requirements in its technical architecture, providing a scalable, verifiable, and tamper-resistant engineering solution for the trusted preprocessing stage in the digital content ecosystem.
[0056] Reference Figure 3 As one implementation of step S102, the step of dynamically constructing at least one logical sub-chain based on the semantic correlation between sub-blocks and generating the root hash value of each logical sub-chain includes: Step S301: Obtain multiple sub-blocks and their corresponding multimodal feature data and sub-block hash values; These sub-blocks are not arbitrary fragments from the original audiovisual stream, but semantically coherent units formed after processing by the preceding AI segmentation engine. Each sub-block represents a content fragment with independent narrative or functional meaning, such as a dialogue, a scene change, or a piece of background music.
[0057] The associated multimodal feature data contains a digital abstract representation of the content of the sub-block: visual feature vectors are usually composed of keyframe SIFT or CNN embedding vectors extracted by convolutional neural networks, reflecting the objects, composition and motion information in the picture; audio feature vectors mostly use Mel frequency cepstral coefficients (MFCC) or deep audio embedding (such as VGGish) to characterize speech content, timbre features or background music style.
[0058] In addition, each sub-block has a unique hash value generated by cryptographic hashing algorithms such as SHA-256. This hash value is not only calculated based on its multimodal feature data, but may also incorporate the creator identity identifier (DID) and timestamp to form a collision-resistant and tamper-proof content fingerprint.
[0059] Step S302: Calculate the semantic similarity between any two sub-blocks based on multimodal feature data, and generate a similarity matrix; The core of this process lies in measuring the proximity of different sub-blocks in a multimodal semantic space, rather than simply comparing visual or audio waveforms. The system first normalizes the visual and audio feature vectors of each sub-block to eliminate dimensional differences. Then, it generates a unified multimodal joint feature vector through a weighted fusion strategy (such as linear weighting or attention mechanisms). In this high-dimensional semantic space, the semantic distance between any two sub-blocks is calculated using a cosine similarity algorithm. This method reflects the directional consistency through the cosine of the angle between the vectors, with a value range of [-1, 1], where a value closer to 1 indicates greater semantic similarity.
[0060] For example, two sub-blocks displaying "city night scenes" accompanied by the same background music will have highly aligned multimodal feature vectors in space, resulting in higher similarity scores. The similarity calculation results between all sub-block pairs are organized into a symmetric N×N matrix (N being the total number of sub-blocks), i.e., the similarity matrix, where each row or column corresponds to a sub-block, and the matrix element (i,j) represents the semantic association strength between the i-th and j-th sub-blocks. This matrix is not only the input for subsequent clustering operations but can also be viewed as a "semantic topology map" of the entire audiovisual work, revealing implicit structural relationships between content fragments, such as thematic repetition, plot echoes, or stylistic continuity.
[0061] Understandably, this similarity calculation mechanism based on multimodal joint features is significantly better than single-modal analysis, and can identify content with large visual differences but consistent semantics (such as the monologues of the same character in different scenes), thereby improving the accuracy of overall association judgment.
[0062] Step S303: Aggregate sub-blocks with similarity higher than the preset clustering threshold into at least one logical sub-chain. Traditional clustering methods, such as K-means, rely on a preset number of clusters and are sensitive to initial centers, making them difficult to adapt to complex and ever-changing audiovisual content structures. This application's embodiments can employ density clustering algorithms (such as DBSCAN), which have the advantage of not requiring a pre-set number of clusters. Instead, they automatically identify high-density regions based on the "density reachability" principle. That is, in the similarity space, if a sub-block has a sufficient number of neighboring highly similar sub-blocks, then that region is determined to be a core cluster.
[0063] Specifically, the system first sets a clustering threshold (such as 0.75) on the similarity matrix, and regards sub-blocks with similarity higher than this value as "neighboring points". Then, through the density propagation mechanism, it identifies all connected sub-block sets to form a preliminary sub-block cluster.
[0064] However, relying solely on static similarity can lead to semantic breaks; for example, two visually similar but temporally disjointed advertisement clips might be incorrectly grouped into the same narrative unit. To address this, the system introduces an audio continuity constraint mechanism, utilizing the temporal series characteristics of audio features to correct the clustering results: if the sub-blocks within a certain sub-cluster are too discretely distributed on the original timeline, or if their audio signals exhibit significant discontinuities in spectral transitions (such as sudden silences or music changes), then the cluster is considered to fail to meet the semantic coherence condition and requires further splitting or boundary adjustment.
[0065] For example, in a documentary, multiple segments about "climate change" may be scattered across different chapters. However, if the background narration has a continuous tone and rhythm, it can be identified as belonging to the same logical subchain. Conversely, if an irrelevant advertisement is inserted, causing an audio interruption, then even if the visual themes are similar, they should belong to different subchains. This dual mechanism, combining density clustering and temporal continuity constraints, ensures that the generated logical subchains are not only highly related semantically but also maintain reasonable coherence in temporal evolution, thus more closely aligning with actual creative logic.
[0066] Step S304: Construct a hash tree structure from the hash values of the sub-blocks within each logical sub-chain; This step overcomes the limitations of traditional linear hash chains in terms of verifiability and scalability by introducing a tree structure based on time order. The system first uses a time-series sorting unit to arrange the hash values of each sub-block into leaf nodes according to their original playback order based on the timestamps (accurate to milliseconds), ensuring that the tree structure reflects the actual evolution path of the content.
[0067] For example, in a logical subchain containing "opening—interview—conclusion", even if some sub-blocks are highly similar in semantics, their hash values still need to be arranged in chronological order to avoid verification failure due to disordered order. Subsequently, the recursive aggregation unit pairs adjacent leaf nodes together, uses the SHA-256 algorithm to calculate their concatenated hash value as the parent node, and aggregates upwards layer by layer until a unique root node is generated.
[0068] This implementation process essentially constructs a Merkle tree, whose core advantage lies in supporting local verification: during subsequent copyright checks, it is not necessary to download all sub-blocks of the entire sub-chain; only the hash value of the target sub-block and its authentication path in the tree (i.e., the "Merkle path") are required to verify whether it belongs to the logical sub-chain through a small number of hash operations. Furthermore, the structural sensitivity of the Merkle tree means that any modification to a leaf node (such as replacing or deleting a sub-block) will cause unpredictable changes to the root hash value, thus providing strong cryptographic guarantees for content integrity.
[0069] Step S305: Calculate the root node value of the hash tree structure as the root hash value of the logical subchain.
[0070] The root hash value serves as the unique digital fingerprint of the entire logical sub-chain, centrally reflecting the content state, semantic relationships, and temporal order of all sub-blocks within that sub-chain. The system calculates the parent node hash layer by layer using the SHA-256 algorithm, ensuring that the output of each layer strictly depends on the input of the layer below, forming a bottom-up data dependency chain. If the content of any sub-block is tampered with, regardless of its depth, the root hash value will ultimately change through the tree's hierarchical propagation, thus being captured by the on-chain verification mechanism. This root hash is not only used for subsequent binding and anchoring to the blockchain with NFTs but also serves as the core basis for copyright determination in smart contract execution.
[0071] For example, when a user submits a piece of material requesting authorization verification, the system can quickly reconstruct its corresponding logical subchain hash tree and compare the calculated root hash with the original root hash stored on the chain. If they do not match, further AI verification or copyright response processes are triggered. This Merkle tree-based root hash generation mechanism not only improves the efficiency and security of data verification but also provides scalable technical support for fine-grained copyright management of complex audiovisual works.
[0072] In the above implementation, density clustering is combined with audio temporal continuity constraints, solving the problem of semantic breaks easily generated in audiovisual content by traditional clustering methods. Simultaneously, by constructing a Merkle tree sorted by timestamps, the hash structure is ensured to reflect both semantic aggregation and temporal continuity, significantly improving the accuracy and tamper resistance of the tracing path. This technical solution provides a high-fidelity structured evidence preservation mechanism for audiovisual works in a blockchain environment, realizing a technological leap from "content fragments" to "verifiable semantic units," demonstrating outstanding creativity and practicality in the field of digital copyright protection.
[0073] Reference Figure 4 As one implementation of step S104, in response to a user query request, the steps of locating the NFT identifier corresponding to the target sub-block, retrieving the sub-chain hash path on the blockchain network, and comparing the user-submitted material with the hash deviation result recorded on the blockchain through the AI verification module include: Step S401: Receive the user-submitted material fragment to be verified and the user's query request; The materials submitted by users may not be complete original works, but rather clips that have been edited, compressed, or partially modified, such as a montage of film and television clips uploaded to a short video platform, partial reuse of advertising materials, or film and television clips referenced in a live broadcast.
[0074] To address this, the system can receive the fragment and its accompanying query metadata (such as request time, user ID, usage scenario, etc.) through a standardized interface and initiate an automated verification process. This design ensures that regardless of the complexity of the source material, as long as it is submitted to the system, it will enter a unified verification channel, avoiding processing deviations due to differences in format or transmission paths.
[0075] More importantly, this step implicitly includes data integrity protection mechanisms, such as preventing intermediate tampering through TLS encrypted transmission and performing hash pre-calculation on the submitted content to record the original state, providing a benchmark for subsequent comparisons.
[0076] Step S402: Analyze the content features of the material segment to be verified and generate the target sub-block identifier; The core objective of this step is to establish a matching "identity tag" for fragments from unknown sources, enabling efficient retrieval from a vast on-chain database.
[0077] Specifically, the system first performs multimodal feature extraction on the materials to be verified: for the video part, representative image sequences are obtained through keyframe extraction algorithms (such as I-frame based or scene change detection), and input into a pre-trained convolutional neural network (such as ResNet or EfficientNet) to generate high-dimensional visual feature vectors. These vectors can capture the object categories, compositional structures and stylistic semantics in the images; for the audio part, its Mel-spectrogram is extracted and further converted into Mel-frequency cepstral coefficients (MFCC) or acoustic embedding vectors are generated using a deep audio model (such as VGGish).
[0078] Subsequently, the system integrates visual and auditory feature vectors into a multi-dimensional feature matrix through weighted concatenation or attention fusion mechanisms, fully preserving the collaborative expression of cross-modal information. Based on this, the SHA-256 hash value of this feature matrix is calculated as the unique target sub-block identifier for the segment. This target sub-block identifier is not a simple file hash, but a deep fingerprint based on content semantics, possessing a certain degree of resistance to compression and resolution changes, maintaining recognizability even after reasonable transcoding or minor editing of the material. For example, a film or television segment re-encoded into a different format but with unchanged content will still have highly consistent keyframes and audio features, and the generated identifier will match the original record, thus supporting effective tracking of "variant usage."
[0079] Step S403: Locate the corresponding NFT identifier and associated logical sub-chain information based on the target sub-block identifier; In this embodiment of the application, this step relies on a decentralized index library (such as a distributed index service built on IPFS or TheGraph) that stores the mapping relationship between "sub-block identifiers and NFT contract addresses" for all registered audiovisual materials.
[0080] Specifically, the system uses the target sub-block identifier as the key to perform an efficient search in the index. If a matching record is found, it returns the corresponding NFT contract address and key information from its metadata, particularly the root hash value of the logical sub-chain. This root hash is the Merkle tree root node value generated during the construction of the preceding dynamic sub-chain, representing the overall integrity status of a set of semantically coherent sub-blocks. Through this root hash, the system can further locate its specific block position and transaction record on the blockchain network (such as Ethereum, Polygon, or Filecoin), completing the full path tracing from content features to on-chain evidence storage.
[0081] Understandably, this mechanism avoids the inefficiency of full-chain scanning, improves query response speed, and ensures censorship resistance and persistence of data access through decentralized indexes. Furthermore, the localization process supports fuzzy matching extensions; for example, when an identifier cannot be precisely matched due to slight distortion, the system can enable approximate hashing (such as pHash) or feature distance thresholding mechanisms to retrieve candidate sets, enhancing the system's robustness.
[0082] Step S404: Retrieve the subchain hash path of the logical subchain from the blockchain network; The system connects to the distributed ledger through a blockchain access interface (such as Web3.js or Ethers.js), calls a smart contract to read the metadata bound to the NFT, and obtains the Merkle tree structure information corresponding to its logical subchain.
[0083] Specifically, the system starts from the leaf node (i.e., the original sub-block hash) and extracts the hash values of all intermediate nodes layer by layer upwards until the root node, forming a complete hash path sequence. This path not only contains the hash values of each level but also their position information in the tree (such as left / right nodes) to reconstruct the calculation process during subsequent verification. To ensure temporal consistency, the system also sorts the path sequence according to the timestamps of each sub-block, restoring its temporal structure in the original creation flow. This complete hash path constitutes the "digital skeleton" of the original content. Any modification to the content, order, or structure of the sub-blocks will cause a change in the hash value of a node in the path, thus affecting the root hash and being detected by the system.
[0084] Step S405: The AI verification module compares the real-time hash value of the material fragment to be verified with the hash path stored on the blockchain. Traditional verification methods typically employ point-by-point hash comparison, directly comparing the hash values of each sub-block to ensure complete consistency. However, this approach is overly sensitive to reasonable editing (such as color correction and subtitle addition), easily leading to false alarms. This application's embodiments introduce an AI verification core, particularly a temporal deviation analysis model based on Long Short-Term Memory (LSTM) networks, significantly improving the intelligence and fault tolerance of the verification.
[0085] Specifically, the system first divides the material to be verified into several verification sub-blocks according to the time structure of the original logical sub-chain to ensure consistent comparison granularity. Then, it generates the real-time hash value of each sub-block and performs a preliminary match with the corresponding node in the chain path. The time alignment unit of the AI verification module precisely aligns the two sets of hash sequences along the time axis, eliminating offsets caused by playback speed or editing. The deviation detection unit uses an LSTM model to analyze the dynamic deviation patterns between the real-time hash sequence and the chain path: LSTM captures the long-term dependencies of hash value changes through memory units, identifying the differences between normal editing (such as progressive filter application) and malicious tampering (such as keyframe replacement).
[0086] For example, legitimate color adjustments typically manifest as continuous small fluctuations in hash values, while frame replacement leads to local hash mutations. The system calculates a cumulative deviation index using a sliding window. When the deviation of multiple consecutive windows exceeds a preset tolerance, the corresponding node is marked as "inconsistent," and its position and confidence level are recorded. This verification mechanism based on temporal pattern recognition retains the high sensitivity of cryptographic hashing while introducing semantic-level fault tolerance, significantly improving the accuracy and practicality of the verification results.
[0087] Step S406: Output the hash deviation results and a visual traceability report.
[0088] Specifically, the system first generates a topological diagram of the logical sub-chains, graphically displaying the semantic relationships and temporal order between each sub-block. Nodes identified as "inconsistent" by AI are marked on the diagram, visually revealing potential tampering locations. Simultaneously, a heatmap rendering unit overlays color markers onto the timeline of the original material, with red areas indicating high-deviation segments, facilitating quick identification of problematic segments by users. Furthermore, the visual source tracing report can link creator information and copyright terms from the NFT metadata.
[0089] Understandably, the copyright status binding unit automatically associates creator information, license terms, and scope of authorization from the NFT metadata, forming a structured data package with a dual-path binding of "technical evidence + legal basis." For example, if a fragment is determined to be used without authorization, the report will simultaneously display the specific copyright terms violated (such as "prohibited for commercial use") and the rights holder's contact information, supporting subsequent automated copyright responses or judicial evidence collection.
[0090] The above implementation achieves high-precision and automated traceability of the authenticity and copyright compliance of audiovisual materials. Its core innovation lies in the deep integration of traditional cryptographic hash verification with AI time-series modeling, using LSTM to analyze the dynamic deviation patterns of hash sequences, thus overcoming the limitations of static comparison; at the same time, through decentralized indexing and NFT metadata binding, it achieves efficient association between content fingerprints and copyright information.
[0091] Reference Figure 5 As one implementation of step S105, the steps of triggering a predefined copyright response action in a smart contract based on the hash deviation result, updating the permission status in the decentralized identity system based on the smart contract execution result, and synchronizing it to the blockchain network include: Step S501: Receive the hash deviation result and associated tampering location data output by the AI verification module; The hash bias result is not a simple boolean value (such as "consistent / inconsistent"), but a composite data packet containing multidimensional information. It is usually encapsulated in JSON or Protobuf format, and includes the bias index, the duration of continuous bias, the timestamp interval of the affected sub-blocks, the spatial location (such as the range of video frames), and the tamper confidence score.
[0092] For example, when the AI verification module identifies a sudden change in the hash sequence from second 45 to second 52 of a video using the LSTM model, and the cumulative deviation of the sliding window exceeds a preset threshold of more than 3 seconds, the event is judged as a high-confidence tampering, and a corresponding deviation signal is generated.
[0093] Simultaneously, the system acquires metadata about the tampered location, such as the Merkle path of the sub-block in the original logical sub-chain, its associated NFT identifier, and the comparison result between the original hash value and the real-time hash value. This data together constitutes the context of the response decision, ensuring that subsequent rule matching is based not only on "whether tampering has occurred," but also on "the nature, scope, and persistence of the tampering."
[0094] The technical implementation of the above steps relies on a standardized event bus or message queue (such as Kafka or RabbitMQ) between modules to ensure the real-time transmission and processing order consistency of signals. At the same time, a digital signature mechanism is used to verify the authenticity of the signal source and prevent malicious forgery or man-in-the-middle attacks.
[0095] Step S502: Match the hash deviation result with a predefined copyright response rule base to obtain the response rule matching result; The copyright response rule base itself is a structured collection of knowledge, typically deployed in the form of a decision tree or rule engine (such as Drools), supporting the mapping of response actions by dimensions such as copyright type, usage scenario, and creator strategy. Therefore, this process is not a simple rule lookup, but rather dynamic reasoning and level determination based on multi-dimensional context.
[0096] Specifically, the system first analyzes key parameters in the hash deviation result, such as whether the tampering location involves core content (e.g., the opening logo, key plot points), whether the deviation value continuously exceeds a threshold for a certain duration (e.g., more than 5 seconds), and the copyright term type defined in the NFT metadata of the material (e.g., "No modification allowed," "Non-commercial editing permitted"). For example, if an NFT metadata states "Personal viewing only, distribution prohibited," and the system detects that the material has been completely reused on a public platform, it triggers the highest level of infringement response; conversely, if it is only a reasonable 15-second reference on a short video platform, and the deviation value is minor, it may trigger authorization guidance rather than direct freezing.
[0097] In particular, the system can also introduce a response level escalation mechanism: when the deviation signal persists for more than a set time (such as 10 consecutive verification failures), it will automatically escalate from "warning" to "asset freeze," reflecting a progressive governance approach to persistent infringement. This context-aware rule matching mechanism makes copyright response no longer a rigid "one-size-fits-all" approach, but a refined governance strategy that can adapt to different creative intentions and usage scenarios.
[0098] Step S503: Trigger the execution function corresponding to the matching result of the smart contract call and response rules; As an immutable program deployed on the blockchain, the function calls of smart contracts are entirely driven by the matching results of preceding rules, requiring no manual intervention. Based on the matched response rules, the system dynamically selects and calls predefined execution functions within the contract. If the response rule is infringement handling, the contract execution controller calls the "asset freeze function," which suspends the transfer, authorization, or trading rights of the asset by modifying state variables in the NFT contract (e.g., isFrozen = true), preventing further spread of infringing content. Simultaneously, it triggers the "evidence collection function," initiating an automated process for encapsulating infringement evidence.
[0099] Conversely, if the response rule is an authorization license (e.g., a user requests legal use but has not yet paid), the system calls the "certificate generation function" and the "billing function" to automatically calculate the fee and generate a legally valid digital license certificate.
[0100] Understandably, the essence of this dual-track response mechanism lies in the unified management of infringement prevention and compliant authorization through conditional branching within the same smart contract framework. This improves code reusability and ensures the consistency and transparency of the response logic. All function calls are broadcast via blockchain transactions, possessing traceability and non-repudiation characteristics, ensuring the execution process is open and trustworthy.
[0101] Step S504: Generate execution logs and state change data of the smart contract execution results; This step not only records "what was done", but also encapsulates a complete chain of evidence for "why it was done" and "how it was done".
[0102] In infringement scenarios, the evidence chain construction module initiates a dual-stream comparison unit to simultaneously acquire the original on-chain data (from IPFS storage) of the tampered sub-block and the user-submitted data to be verified. A spatiotemporal stamp generator then marks each comparison sample with precise time coordinates (UTC time) and spatial location (block height, transaction hash), forming spatiotemporally aligned evidence pairs. This data is encapsulated into a structured evidence package, containing the original Merkle path, real-time hash value, deviation analysis report, and IPFS Content Identifiers (CIDs) for both versions of the media fragment. Finally, the entire package is re-hashed and stored on the blockchain, forming an immutable electronic evidence chain.
[0103] For authorization scenarios, the system generates a digital license certificate containing the scope of authorization, usage period, fee details, and payment QR code, and writes its hash value to the execution log as a basis for subsequent verification. The entire execution log is encoded in the W3C Verifiable Credentials standard format, supporting cross-platform verification and long-term archiving, providing an authoritative data source for judicial evidence collection, platform arbitration, or automated auditing.
[0104] Step S505: Synchronize the state change data to the decentralized identity system and blockchain network.
[0105] This step ensures that changes in copyright status not only remain at the asset level but are also reflected in the credible expression of the subject's identity. The system will write the execution result (such as "asset frozen" or "authorized") into the relevant creator's or user's DID document in the form of a claim, for example, adding a "content_frozen" field to the creator's DID and a "license_issued" credential to the user's DID.
[0106] More importantly, through the cross-chain oracle mechanism, the system synchronizes key state changes to multiple blockchain networks: the oracle listens for updates to DID documents and automatically triggers updates to the state flags of NFT contracts on the target chain, ensuring consistency of permission states even in a cross-chain environment. For example, if an NFT is frozen on Ethereum, its mirror asset on the Polygon chain is also locked synchronously through the oracle, preventing cross-chain arbitrage. This atomic design of cross-chain state solves the problem of permission fragmentation in multi-chain ecosystems, achieving true global governance.
[0107] In the above implementation, AI-driven deviation detection is deeply integrated with the dual-track response mechanism of smart contracts. This supports both the rapid containment of infringement and the convenient authorization of compliant use, forming a dynamic governance system that emphasizes both punishment and incentives. Through dynamic evidence chain encapsulation and cross-chain state synchronization, the system not only ensures the auditability and consistency of the execution process but also provides a programmable, verifiable, and scalable copyright governance infrastructure for the digital content ecosystem, enhancing the self-control capabilities and legal enforcement power of audiovisual works in an open network environment.
[0108] Reference Figure 6 As a further implementation of the method for tracing the source of audiovisual works, the steps after comparing the user-submitted materials with the hash deviation results recorded on the blockchain through the AI verification module include: Step S601: Extract sub-blocks and associated logical sub-chains whose hash deviation results exceed a preset deviation threshold; Specifically, when the system identifies a significant deviation between the real-time hash value of a sub-block and its hash path stored on the blockchain, and the deviation continues to exceed a preset threshold (e.g., more than 3 consecutive seconds, or a cumulative deviation index greater than 0.8), the system not only records the abnormal node but also actively traces its logical sub-chain structure.
[0109] The logical subchain identifier, serving as the attribution label for the sub-block during semantic clustering, contains its path information in the Merkle tree, timestamp sequence, and cluster ID, and is a key index for locating the entire semantic unit. By extracting this identifier, the system can elevate the abnormal behavior of a single sub-block to a questioning of the integrity of the entire logical subchain, avoiding the risk of the entire copyright certificate becoming invalid due to partial tampering without being detected.
[0110] For example, in an NFT version of a movie, if a key plot sub-block is maliciously replaced with advertising content, the system not only marks the sub-block as tampered, but also includes its "main plot sub-chain" in the verification scope to ensure that the overall content security is not misjudged because the root hash on the chain is still valid (if it has not been recalculated).
[0111] Step S602: Start the subchain integrity verification protocol and recalculate the root hash value according to the associated logical subchain; The system first retrieves all original sub-block data contained in the logical sub-chain associated with the given block using the IPFS Content Identifier (CID). This data is stored in a decentralized distributed file system, possessing content addressing and tamper-proof characteristics. Subsequently, the system reconstructs the arrangement structure of each sub-block in the original creation stream based on its timestamp order, and calculates the parent node hash value layer by layer from the leaf node according to the Merkle tree construction rules, ultimately generating a new root hash value.
[0112] Understandably, this process is essentially a "complete replay verification" of the original content. Its technical value lies in the fact that even if the root hash stored on-chain remains unchanged, as long as the underlying original data changes (such as IPFS nodes being maliciously replaced or storage failing), the newly calculated root hash will differ from the on-chain record, thus exposing potential data decoupling risks. This step also introduces fault tolerance mechanisms, such as enabling redundant replica recovery strategies when data retrieval fails, or using interpolation compensation algorithms for partially missing sub-blocks, to ensure the robustness and feasibility of the verification process.
[0113] Step S603: When the recalculated root hash value is inconsistent with the record on the blockchain, mark the subchain as an abnormal subchain. The system converts the old and new root hash values into binary sequences, calculates their Hamming distance, and compares it with a preset threshold. The Hamming distance measures the number of different bits at the same position in two equal-length binary strings. If the distance exceeds the threshold (e.g., a difference of more than 5 bits), it is considered that there is a substantial difference between the two, sufficient to affect the judgment of content integrity.
[0114] Compared to traditional exact match mechanisms, the Hamming distance method offers stronger noise resistance, distinguishing between "minor transmission errors" and "substantial content tampering," thus avoiding misjudgments caused by the hash algorithm's avalanche effect. For example, if the audio gain of only one sub-block is slightly adjusted, its hash value may be completely different. However, if the overall semantic structure remains unchanged, the Hamming distance may still be within an acceptable range, allowing the system to issue a warning instead of immediately freezing. This judgment logic based on quantization differences improves the system's stability and accuracy in complex network environments.
[0115] Step S604: Freeze the NFT transfer permissions associated with the abnormal sub-chain through a smart contract and trigger an alarm event.
[0116] Specifically, once a logical subchain is marked as abnormal, the system immediately updates the status flag of the NFTs bound to that subchain via a smart contract call, setting it to "frozen" or "pending_audit." This prohibits their trading, transfer, or licensing on the secondary market, preventing the continued circulation of contaminated content and avoiding wider copyright confusion. It should be noted that this freeze is reversible; after subsequent manual review or automatic repair, the freeze can be lifted through re-verification and governance voting.
[0117] Simultaneously, the system triggers multi-channel alarm events, including sending encrypted notifications to the creator's DID, pushing audit logs to regulatory nodes, and broadcasting anomaly records in a decentralized message bus (such as OrbitDB), ensuring that relevant parties are promptly informed of the risks. Alarm information includes the abnormal subchain ID, deviation location, comparison results of the old and new root hashes, and suggested handling measures, supporting automated response processes or manual intervention. This mechanism not only enhances the system's proactive defense capabilities but also provides technical assurance for the value stability of NFT assets, preventing the "hollowing out" of digital credentials due to underlying data distortion.
[0118] The above implementation constructs a proactive data self-inspection closed loop of "deviation extraction—integrity verification—root hash reconstruction—state freezing," achieving deep security hardening of the blockchain copyright system. Its core innovation lies in breaking through the passive mode of traditional traceability systems that rely solely on on-chain hash storage, introducing periodic or event-driven reverse verification of the original stored data to ensure that on-chain credentials and off-chain content remain consistent. Through Hamming distance quantification and a dynamic NFT permission freezing mechanism, the system ensures security while also considering fault tolerance for misjudgments and governance flexibility, effectively addressing real-world challenges in decentralized environments such as heterogeneous data storage, untrustworthy nodes, and susceptibility to tampering. This technical solution not only enhances the robustness and credibility of the copyright traceability system but also provides a verifiable and executable technical anchor for the binding relationship between "digital credentials and real content" in the NFT ecosystem, propelling digital assets from "formal ownership confirmation" to a higher-level governance stage of "substantive credibility."
[0119] Reference Figure 7 As a further implementation of the method for tracing the source of audiovisual works, the method for tracing the source of materials also includes: Step S701: Capture the spatiotemporal feature data of the user query request in real time; Among them, the spatiotemporal feature data includes geographic coordinates, query timestamps, and device fingerprints; Specifically, when the system receives a user's source tracing query request, it simultaneously extracts the geographic coordinate information carried by the user (usually obtained through GPS, Wi-Fi positioning, or IP geographic mapping technology), accurate to the city or region level; records the timestamp of the query, accurate to the millisecond level, to analyze the time distribution pattern of access; and collects device fingerprints, which are unique identifier combinations generated by the user terminal, including but not limited to device model, operating system version, browser characteristics, screen resolution, and hardware hash value, to distinguish the access patterns of different user groups.
[0120] Understandably, this spatiotemporal data not only reflects "who, when, and where" queries, but also implies potential content consumption preferences and online behavior trends. For example, during an international film festival, mobile devices from North America frequently searched for specific clips of an award-winning film during the evening hours. The system can use this spatiotemporal feature to identify that the clip may become a regional hot topic, thus providing a basis for decision-making regarding subsequent resource preloading.
[0121] Step S702: Calculate the access frequency of each logical sub-chain in the spatiotemporal feature data within a preset time window, and generate a sub-chain access matrix containing popularity weight values. The generation of the sub-chain access matrix includes: generating a first heat density map by geographic region, generating a second heat waveform map by time series, and generating a third terminal distribution map by device type; Specifically, the system uses logical sub-chains as the basic statistical unit, divides time into fixed windows (such as every 15 minutes, every hour, or every day), and aggregates and analyzes the query counts of each sub-chain within each window. Unlike simple access counting, this embodiment employs a multi-dimensional aggregation strategy, constructing a visual feature map from three dimensions: space, time, and device. The first heat map statistically analyzes the distribution of sub-chain accesses by geographical region (such as country, city, or operator region), forming a spatial heat map to reveal the intensity of attention to specific content in different regions; the second heat waveform map unfolds the access frequency by time series, showing the periodic fluctuations, sudden increases, or decreases in sub-chain popularity, for example, a documentary clip shows a regular peak during weekday mornings at educational institutions; the third terminal distribution map statistically analyzes the access percentage by device type (such as smartphone, tablet, smart TV, or VR headset), reflecting the differences in content format preferences among different terminal users.
[0122] Understandably, these feature maps do not exist independently, but rather serve as input tensors to a Graph Convolutional Neural Network (GCN). Through graph structure modeling, geographical regions, time segments, and device types are treated as nodes, and access frequency is used as edge weights to construct a multimodal heterogeneous graph. The GCN extracts high-order features through a local neighborhood aggregation mechanism, ultimately outputting a comprehensive popularity weight value for each logical sub-chain. This value not only reflects the frequency of access but also incorporates deeper semantics such as spatial concentration, temporal persistence, and terminal coverage breadth, forming a quantitative assessment of content popularity.
[0123] Step S703: Construct a heat propagation model based on the sub-chain access matrix to predict the sub-chain access distribution in future periods and obtain the prediction results; This popularity propagation model is not a simple extrapolation prediction, but rather a fusion architecture based on graph neural networks and time series modeling, capable of capturing the diffusion patterns of content popularity in the spatiotemporal dimension. The system uses historical access matrices as training data, combined with external factors (such as social media spread indices, holiday effects, or platform recommendation strategies) as auxiliary inputs, to train a spatiotemporal graph convolutional network (ST-GCN). This model can simulate the propagation path of popularity spreading from the core area to the periphery and from high-activity periods to adjacent periods.
[0124] For example, when a film clip first becomes popular in Asia, the model can predict that it will attract increased attention in Europe and America within 24 hours, and deploy resources accordingly. The prediction results are output in the form of a probability distribution, indicating the expected access intensity of each logical subchain in different geographical regions within a specific time window (such as the next hour or the next day), providing accurate guidance for the replication of edge nodes. This prediction mechanism has self-learning capabilities, continuously optimizing weight parameters as data accumulates, improving prediction accuracy, and its measured error rate is significantly lower than that of traditional ARIMA or LSTM models.
[0125] Step S704: Based on the prediction results, subchains whose heat weight values exceed the preset weight threshold are marked as high-frequency subchains; The preset weight threshold is not a fixed constant, but is dynamically adjusted based on the overall system load, edge node capacity, and content lifecycle. For example, during peak platform traffic periods, the system can appropriately increase the threshold to avoid excessive migration leading to edge resource overload; while for new content in the promotion phase, the threshold can be lowered to support early popularity cultivation.
[0126] Understandably, content marked as high-frequency subchains typically possesses characteristics such as consistently high access volume, cross-regional dissemination potential, or critical copyright verification requirements. Its data integrity and access response speed directly impact user experience and copyright governance efficiency. These subchains often involve popular film clips, classic music samples, or high-value NFT-related materials; access delays or verification failures could lead to copyright disputes or commercial losses. Therefore, including them in the priority acceleration service reflects the system's intelligent priority management of resource allocation.
[0127] Step S705: Migrate the complete data copy of the high-frequency subchain to the edge computing node cluster that is geographically closest to the user.
[0128] This step fully leverages the "proximity to users and low latency" advantages of edge computing, replicating logical sub-chain data, originally stored centrally on centralized servers or the IPFS mainnet, to edge nodes (such as CDN edge servers, 5G MEC nodes, or regional data centers) closer to end users as needed. The migration process is driven by a scheduling engine, comprehensively considering the user's geographical coordinates and the topology of the edge nodes, selecting the target cluster with the fewest network hops and the lowest RTT for deployment.
[0129] For example, for a documentary subchain frequently queried by European users, the system can push a copy to edge nodes in Frankfurt or Amsterdam, allowing local users to initiate traceability verification without having to go back to the Asian data center, significantly reducing transmission latency. The migration operation includes not only the synchronization of the original sub-block data and hash path, but also the complete replication of associated metadata, Merkle tree structure, and NFT mapping index, ensuring that edge nodes have the ability to independently respond to queries.
[0130] In addition, the system updates the routing table of the blockchain light nodes after migration, enabling them to intelligently select the nearest data source based on the user's location, avoiding the bandwidth overhead caused by full node synchronization, while maintaining consistency with the main chain state.
[0131] The above implementation achieves dynamic performance optimization of the audiovisual material traceability system, transforming traditional passive content distribution into proactive intelligent preloading. It uses graph convolutional neural networks to fuse geographic, temporal, and device features to generate a high-precision popularity weight assessment, and combines genetic algorithms and erasure coding mechanisms to achieve optimal resource allocation. This technical solution not only significantly reduces user query latency and improves system throughput, but also effectively alleviates the pressure on centralized storage and bandwidth costs while ensuring data integrity and copyright verification accuracy. It provides scalable and adaptive technical support for large-scale, high-concurrency, low-latency digital copyright services, driving the evolution of blockchain traceability systems from "usable" to "highly efficient and usable."
[0132] This application also discloses a blockchain-based audiovisual work material traceability system.
[0133] A blockchain-based audiovisual work material traceability system, the traceability system comprising: The material segmentation module is used to input the original audiovisual work material into the AI segmentation engine, segment it into several sub-blocks according to semantic relevance, and extract the hash value and metadata of each sub-block; The logical sub-chain construction module is used to dynamically construct at least one logical sub-chain based on the semantic correlation between sub-blocks, and generate the root hash value of each logical sub-chain; The NFT identifier generation module is used to bind the root hash value with metadata to generate an NFT identifier, and to anchor the NFT identifier and associated data to the blockchain network; The hash deviation comparison module is used to respond to user query requests, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and compare the user-submitted material with the hash deviation results recorded on the blockchain through the AI verification module. The copyright response execution module is used to trigger the execution of predefined copyright response actions by the smart contract based on the hash deviation result. Based on the execution result of the smart contract, the permission status in the decentralized identity system is updated and synchronized to the blockchain network.
[0134] The audiovisual work material tracing system based on blockchain according to the present application embodiment can implement any of the above-mentioned audiovisual work material tracing methods, and the specific working process of each module in the audiovisual work material tracing system can refer to the corresponding process in the above-mentioned method embodiment.
[0135] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0136] This application also discloses a computer device.
[0137] Computer equipment, including memory, processor, and computer program stored on memory and executable on processor, wherein the processor executes the computer program to implement a blockchain-based method for tracing the source of audiovisual works as described above.
[0138] This application also discloses a computer-readable storage medium.
[0139] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the blockchain-based methods for tracing the source of audiovisual works.
[0140] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0141] It should be noted that the computer device and storage medium in this application embodiment are respectively electronic devices and storage media that apply the above-described method for tracing the source of audiovisual works based on blockchain. Therefore, all embodiments of the above-described method for tracing the source of audiovisual works are applicable to the computer device and storage medium, and can achieve the same or similar beneficial effects. For the computer device / storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple; relevant details can be found in the descriptions of the method embodiments.
[0142] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0143] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0144] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A method for tracing the source of audiovisual works based on blockchain, characterized in that, The material tracing method includes: The original audiovisual material is input into the AI segmentation engine, divided into several sub-blocks according to semantic relevance, and the hash value and metadata of each sub-block are extracted. Based on the semantic correlation between sub-blocks, at least one logical sub-chain is dynamically constructed, and the root hash value of each logical sub-chain is generated; The root hash value is bound to metadata to generate an NFT identifier, and the NFT identifier and associated data are anchored to the blockchain network; In response to user query requests, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and use the AI verification module to analyze the hash change pattern of the sub-block sequence using a time series model to capture the difference between normal editing and malicious tampering, and compare the hash deviation results between the user-submitted material and the hash deviation recorded on the blockchain. Based on the hash deviation result, the smart contract is triggered to execute a predefined copyright response action. Based on the execution result of the smart contract, the permission status in the decentralized identity system is updated. The execution result is written into the creator's or user's decentralized identity document in the form of a declaration and synchronized to the blockchain network. The steps after comparing the user-submitted materials with the hash deviation results recorded on the blockchain include: extracting sub-blocks and associated logical sub-chains whose hash deviation results exceed a preset deviation threshold; initiating a sub-chain integrity verification protocol and recalculating the root hash value based on the associated logical sub-chains; marking the sub-chain as an abnormal sub-chain when the recalculated root hash value is inconsistent with the record on the blockchain; freezing the NFT circulation permissions associated with the abnormal sub-chain and triggering an alarm event through a smart contract. The steps for dynamically constructing at least one logical sub-chain based on the semantic correlation between sub-blocks and generating the root hash value of each logical sub-chain include: Obtain multiple sub-blocks and their corresponding multimodal feature data and sub-block hash values; Based on the multimodal feature data, the semantic similarity between any two sub-blocks is calculated, and a similarity matrix is generated; Based on a preset clustering threshold, sub-blocks with similarity higher than the threshold are aggregated into at least one logical sub-chain; wherein, the time series characteristics of audio features are used to correct the clustering results, and sub-block clusters that do not meet the semantic coherence condition are split or their boundaries are adjusted. Construct a hash tree structure from the hash values of the sub-blocks within each logical sub-chain; Calculate the root node value of the hash tree structure as the root hash value of the logical subchain; The AI verification module uses a time-series model to analyze the hash change patterns of sub-block sequences, capturing the differences between normal editing and malicious tampering. The steps for comparing the user-submitted materials with the hash deviation results recorded on the blockchain include: The material to be verified is divided into real-time hash values according to the time structure of the logical sub-chain. The time alignment unit of the AI verification module accurately aligns the real-time hash sequence with the on-chain path according to the time axis. The deviation detection unit uses the long short-term memory network to analyze the dynamic deviation pattern between the real-time hash sequence and the on-chain path to capture the long-term dependency relationship of hash value changes. The cumulative deviation index is calculated by sliding window. When the deviation of multiple consecutive windows exceeds the preset tolerance, the corresponding node is marked as inconsistent.
2. The method for tracing the source of audiovisual works based on blockchain according to claim 1, characterized in that, The steps of inputting the original audiovisual material into the AI segmentation engine, dividing it into several sub-blocks according to semantic relevance, and extracting the hash value and metadata of each sub-block include: Receive original audiovisual works and associated creator identity data; The AI segmentation engine analyzes the semantic relationships of the material and divides it into several sub-blocks. Extract the multimodal feature data of each sub-block and generate the sub-block hash value; The sub-block hash value, multimodal feature data, and creator identity data are encapsulated as structured metadata.
3. The method for tracing the source of audiovisual works based on blockchain according to claim 1, characterized in that, In response to a user query request, the steps of locating the NFT identifier corresponding to the target sub-block, retrieving the sub-chain hash path on the blockchain network, and comparing the user-submitted material with the hash deviation result recorded on the blockchain through the AI verification module include: Receive user-submitted material fragments to be verified and user query requests; Analyze the content features of the material segment to be verified and generate a target sub-block identifier; Locate the corresponding NFT identifier and associated logical sub-chain information based on the target sub-block identifier; Retrieve the subchain hash path of the logical subchain from the blockchain network; The AI verification module compares the real-time hash value of the material fragment to be verified with the hash path stored on the blockchain. Output hash bias results and a visual traceability report.
4. The method for tracing the source of audiovisual works based on blockchain according to claim 3, characterized in that, The steps involved in triggering a predefined copyright response action in a smart contract based on the hash bias result, updating the permission status in the decentralized identity system based on the smart contract execution result, and synchronizing it to the blockchain network include: Receive the hash deviation result and associated tampering location data output by the AI verification module; The hash deviation result is matched against a predefined copyright response rule base to obtain the response rule matching result. Trigger the smart contract to call the execution function corresponding to the matching result of the response rule; Generate execution logs and state change data for smart contract execution results; The status change data will be synchronized to the decentralized identity system and blockchain network.
5. A method for tracing the source of audiovisual works based on blockchain according to any one of claims 1 to 4, characterized in that, The material tracing method also includes: Real-time capture of spatiotemporal feature data of user query requests; wherein, the spatiotemporal feature data includes geographic coordinates, query timestamp, and device fingerprint; The access frequency of each logical sub-chain in the spatiotemporal feature data within a preset time window is statistically analyzed to generate a sub-chain access matrix containing a heat weight value. Based on the sub-chain access matrix, a heat propagation model is constructed to predict the sub-chain access distribution in future periods, and the prediction results are obtained. Based on the prediction results, subchains whose heat weight values exceed a preset weight threshold are marked as high-frequency subchains; The complete data copy of the high-frequency subchain is migrated to the edge computing node cluster that is geographically closest to the user.
6. A blockchain-based audiovisual work material traceability system, characterized in that, A blockchain-based method for tracing the source of audiovisual works as described in any one of claims 1 to 5, the tracing system comprising: The material segmentation module is used to input the original audiovisual work material into the AI segmentation engine, segment it into several sub-blocks according to semantic relevance, and extract the hash value and metadata of each sub-block; The logical sub-chain construction module is used to dynamically construct at least one logical sub-chain based on the semantic correlation between sub-blocks, and generate the root hash value of each logical sub-chain; The NFT identifier generation module is used to bind the root hash value with metadata to generate an NFT identifier, and to anchor the NFT identifier and associated data to the blockchain network; The hash deviation comparison module is used to respond to user query requests, locate the NFT identifier corresponding to the target sub-block, retrieve the sub-chain hash path on the blockchain network, and compare the user-submitted material with the hash deviation results recorded on the blockchain through the AI verification module. The copyright response execution module is used to trigger the execution of predefined copyright response actions by the smart contract based on the hash deviation result. Based on the execution result of the smart contract, the permission status in the decentralized identity system is updated and synchronized to the blockchain network.
7. A computer device, characterized in that: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Electronic data storage verification method and system based on hierarchical hash and smart contract
CN120415825A
Video stream dynamic fragment encryption and block chain evidence storage method
CN120416543A