An image end-edge-cloud collaborative authentication method based on adaptive feature compression
Patent Information
- Application Number
- CN202610871763.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明的目的在于克服现有技术的不足,提供一种基于自适应特征压缩的图片端-边-云协同确权方法,解决传统图片确权技术终端算力开销大、特征压缩固化、特征重建精度低、批量存证效率低、版权数据易篡改、设备适配性差的技术问题,实现轻量化边缘终端、高性能边缘服务器、不可篡改区块链网络的高效协同,完成图片自适应特征压缩、精准确权、安全存证与快速验证
[0029]1、本发明采用模型分割与端边算力分层协同机制,将基础卷积模型拆分部署于边缘设备与边缘服务器,轻量化边缘终端运算压力,适配手机、嵌入式设备等低算力边缘终端,大幅降低终端算力消耗与本地存储开销。
Smart Images

Figure CN122818318A_ABST
Abstract
Description
Technical Field
[0001] This invention provides an edge-cloud collaborative image copyright confirmation method based on adaptive feature compression, belonging to the fields of image copyright confirmation, edge computing, feature compression and blockchain evidence storage technology. Background Technology
[0002] With the rapid development of new media technologies, digital photography, and online communication technologies, massive amounts of image resources are rapidly disseminated, reused, and circulated on internet platforms. Problems such as image piracy, unauthorized alteration, and infringing dissemination occur frequently, making it increasingly difficult to confirm, trace, and verify image copyrights. This seriously damages the legitimate rights and interests of creators and hinders the healthy development of the digital cultural and creative industry.
[0003] Traditional methods for confirming image copyright mainly include watermark embedding, hash-based global feature confirmation, and centralized cloud-based evidence storage. Among these, watermarking is susceptible to damage from image processing operations such as cropping, compression, filtering, and scaling, resulting in poor robustness. Traditional global hash-based confirmation methods require complete feature extraction and hash calculation on the terminal, consuming significant computing power and data transfer volume, making them unsuitable for lightweight edge devices such as mobile phones, cameras, and embedded terminals. Centralized cloud-based confirmation uploads all image data to the cloud for processing, leading to high transmission latency, high bandwidth consumption, high risk of user privacy leaks, and concentrated server computing power pressure.
[0004] Existing edge-cloud collaborative rights confirmation solutions mostly adopt a fixed-ratio feature compression method, which cannot adaptively adjust the compression ratio according to the device's computing power and the redundancy of image features. This easily leads to problems such as loss of key features, low feature reconstruction accuracy, and insufficient uniqueness of hash fingerprints. At the same time, the lack of a systematic model joint optimization mechanism results in low rights confirmation accuracy and poor robustness, making it difficult to adapt to batch image rights confirmation scenarios in complex network environments and diverse edge devices.
[0005] In view of the shortcomings of existing technologies, such as high terminal computing power consumption, poor feature compression flexibility, low rights confirmation accuracy, large data transmission volume, easy tampering of evidence data, and weak system collaboration, there is an urgent need for an image copyright confirmation method that features adaptive feature compression, efficient collaboration between edge and cloud, and accurate and secure rights confirmation. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an edge-cloud collaborative image rights confirmation method based on adaptive feature compression. This method solves the technical problems of traditional image rights confirmation technologies, such as high terminal computing power consumption, fixed feature compression, low feature reconstruction accuracy, low efficiency of batch evidence storage, easy tampering of copyright data, and poor device adaptability. It achieves efficient collaboration between lightweight edge terminals, high-performance edge servers, and tamper-proof blockchain networks, and completes adaptive feature compression, accurate rights confirmation, secure evidence storage, and rapid verification of images.
[0007] To achieve the above objectives, this invention provides an image copyright confirmation method based on adaptive feature compression, which relies on a three-tier collaborative architecture of edge-cloud, including an edge device, an edge server, and a blockchain network. The method includes a joint model pre-training step, an image copyright notarization step, and an image copyright verification step. The specific technical solution is as follows:
[0008] An image edge-cloud collaborative rights determination method based on adaptive feature compression includes the following steps:
[0009] S1. On the edge device, the segmented front-end lightweight convolutional neural network model is used to perform front-end inference on the local image to be stored, and the intermediate feature tensor is obtained.
[0010] S2. On the edge device, based on the sparse mask output by the sparse gated network, channel pruning is performed on the intermediate feature tensor, and entropy encoding is performed on the pruned feature data to generate compressed feature data packets.
[0011] S3. Send the compressed feature data packet to the edge server;
[0012] S4. On the edge server, the compressed feature data packet is entropy decoded, and the decoded features are reconstructed using a feature recovery neural network model to obtain the recovered features;
[0013] S5. On the edge server, perform deep hashing on the recovery features to generate the binary feature hash code of the image;
[0014] S6. Construct a Merkle tree from the feature hash codes of multiple images belonging to the same batch, and write the root hash of the Merkle tree into the blockchain to complete the copyright certificate.
[0015] S7. After receiving the image to be verified, verify its copyright information on the blockchain.
[0016] Preferably, the convolutional neural network model segmentation step includes: selecting a basic convolutional neural network model, determining the segmentation point based on the computing power ratio between edge devices and edge servers, and segmenting the basic convolutional neural network model into a front-end lightweight convolutional neural network deployed on edge devices and a back-end convolutional neural network deployed on edge servers, thereby realizing the hierarchical allocation of model computing power and adapting to the differences in computing power between edge devices.
[0017] Preferably, the sparse gating network adopts a differentiable Top-K sparsity mechanism. The specific process includes: first, performing global average pooling on the input intermediate feature tensor to generate channel statistics; inputting the channel statistics into a lightweight gating network to output a gating weight vector; introducing Gumbel noise and generating a sparse mask using the Gumbel-Softmax method; generating the sparse mask using a continuous approximation method during the training phase to ensure that the model can backpropagate; and generating the sparse mask using a hard Top-K selection method during the inference phase, retaining only the top K feature channels with the highest weights, where K is calculated from a preset global sparsity rate and the total number of channels, thereby achieving adaptive pruning of feature channels and eliminating redundant features.
[0018] Preferably, the feature recovery neural network model is a lightweight convolutional structure, and its processing includes: inputting the pruned sparse feature map, performing channel upsampling through 1×1 convolution to supplement channel dimension information; extracting deep effective features through 3×3 convolution combined with the ReLU activation function; and then restoring the original number of channels through 1×1 convolution to output the reconstructed complete feature map, maximizing the restoration of the original image feature information and ensuring the accuracy of subsequent hash generation.
[0019] Preferably, the deep hashing process includes: inputting the recovered features into a fully connected layer and mapping them to a low-dimensional intermediate vector; during the training phase, using the tanh function to approximate the sign function to solve the problem of the sign function being non-differentiable and to achieve iterative optimization of the model through backpropagation; during the inference phase, using the sign function to generate a standard binary hash code, and then converting it into a 0-1 format hash code, which is used as a unique digital fingerprint for the image to ensure the uniqueness and recognizability of the hash features of each image.
[0020] Preferably, the Merkle tree construction and blockchain notarization steps include: an edge server collecting feature hash codes, timestamps, and author information for the same batch of images; using each image's feature hash code as a leaf node, iteratively calculating hash values layer by layer to construct a complete Merkle tree, and calculating a unique Merkle tree root hash; packaging the Merkle tree root hash, timestamp, and author information into blockchain transaction data, uploading and writing it into the blockchain network, generating a unique transaction ID, realizing the aggregated notarization of batch image copyright information, and reducing blockchain storage overhead.
[0021] Preferably, the copyright verification step includes: repeating the complete process S1-S5 for the image to be verified to generate a target feature hash code for the image to be verified; querying the Merkle root hash and the corresponding Merkle verification path of the image to be verified on the blockchain based on the blockchain transaction ID; iteratively calculating the root hash value layer by layer using the target feature hash code and the Merkle path; comparing the calculated root hash with the root hash natively stored on the blockchain; if they match, the image copyright verification passes and the image is determined to be genuine original content; if they do not match, the verification fails and the image is determined to have the risk of being tampered with or infringing.
[0022] Preferably, the method further includes a joint model pre-training step, specifically: sequentially connecting the front-end lightweight convolutional neural network, sparse gating network, feature recovery neural network, back-end convolutional neural network, and deep hashing module to construct a complete model feedforward link; using a composite loss function for iterative training, wherein the composite loss function integrates four types of losses: feature reconstruction loss, deep hashing loss, sparse constraint loss, and robustness loss, to comprehensively constrain the model training accuracy; and employing a three-stage hierarchical training strategy to improve the overall performance and adaptability of the model: in the first stage, the parameters of the front-end and back-end convolutional neural networks are fixed, and only the sparse gating network and feature recovery neural network are trained to ensure the accuracy of feature compression and reconstruction; in the second stage, all model module parameters are unfrozen, and end-to-end joint training is performed to optimize the overall link synergy; in the third stage, adversarial examples and image data augmentation operations are introduced to enhance the model's anti-interference ability and robustness.
[0023] Preferably, the present invention constructs an edge-cloud collaborative system adapted to the above-mentioned rights confirmation method, the system comprising an edge device, an edge server, and a blockchain network:
[0024] The edge device is equipped with a front-end lightweight convolutional neural network, a sparse gating network, and an entropy coding unit, which are used to extract intermediate feature tensors from local images to be stored, adaptively prune redundant feature channels, compress feature data through entropy coding, and generate lightweight compressed data packets to reduce terminal transmission and computing power overhead.
[0025] The edge server is equipped with an entropy decoding unit, a feature recovery neural network, a post-end convolutional neural network, a deep hashing unit, and a Merkle tree construction and management module. These are used to receive compressed data packets uploaded by edge devices and complete decoding, feature reconstruction, deep hashing encoding, batch Merkle tree construction, and evidence data packaging.
[0026] The blockchain network adopts a public or consortium blockchain architecture to permanently store core copyright data such as Merkle tree root hash, timestamp, and author information. Relying on the immutable and traceable characteristics of blockchain, it provides secure copyright notarization and public verification services.
[0027] The edge device, edge server, and blockchain network work together to execute the aforementioned adaptive feature compression image edge-cloud collaborative rights confirmation method, completing the entire process of copyright registration and verification.
[0028] Due to the adoption of the above technical solution, the beneficial effects of the image edge-cloud collaborative rights confirmation method based on adaptive feature compression of the present invention are as follows:
[0029] 1. This invention adopts a model segmentation and edge computing power hierarchical collaboration mechanism to split and deploy the basic convolution model on edge devices and edge servers, thereby reducing the computing pressure on edge terminals, adapting to low computing power edge terminals such as mobile phones and embedded devices, and significantly reducing terminal computing power consumption and local storage overhead.
[0030] 2. This invention introduces a sparse gating network based on differentiable Top-K, which can adaptively complete feature channel pruning according to a preset sparsity rate, replacing the traditional fixed-ratio compression method. While eliminating redundant features and significantly reducing the amount of data transmission, it retains the core rights features of the image, balancing compression efficiency and feature integrity.
[0031] 3. This invention designs a lightweight feature recovery network that can accurately restore the sparse features after pruning, ensuring the uniqueness and accuracy of the digital fingerprint generated by deep hashing. At the same time, through three-stage model joint training and composite loss function constraints, the feature reconstruction accuracy, hash discrimination and anti-interference robustness of the model are greatly improved.
[0032] 4. This invention adopts a Merkle tree batch aggregation evidence storage mechanism, which aggregates the hash codes of batch images into a root hash for on-chain storage, avoiding the high storage cost and high latency caused by storing a single image on the chain, greatly improving the efficiency of batch image ownership confirmation and evidence storage, and adapting to the scenario of batch evidence storage of massive images.
[0033] 5. This invention relies on the immutable and traceable characteristics of blockchain to store core copyright data, and combines it with the Merkle path verification mechanism to achieve rapid and accurate verification of image copyright, effectively preventing copyright data tampering and forgery, and ensuring the authority and reliability of image copyright confirmation.
[0034] 6. This invention constructs a complete end-edge-cloud three-level collaborative architecture, with clear division of labor and efficient linkage among modules, realizing full-process automation of image feature extraction, compression, transmission, reconstruction, hashing, evidence storage and verification. It is adaptable to complex network environments and diverse terminal devices, and has strong versatility and practicality. Attached Figure Description
[0035] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0036] Figure 1 This is a flowchart of an image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to the present invention;
[0037] Figure 2 This is a system architecture deployment diagram of an edge-cloud collaborative image rights confirmation method based on adaptive feature compression according to the present invention.
[0038] Figure 3 This is a flowchart illustrating the edge device to edge server process of an image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to the present invention.
[0039] Figure 4 This is an adaptive feature compression calculation diagram for an edge-cloud collaborative image rights confirmation method based on adaptive feature compression according to the present invention.
[0040] Figure 5 This invention relates to a deep hashing and Merkle tree construction method for image edge-cloud collaborative rights determination based on adaptive feature compression. Figure 1 ;
[0041] Figure 6 This invention relates to a deep hashing and Merkle tree construction method for image edge-cloud collaborative rights determination based on adaptive feature compression. Figure 2 . Detailed Implementation
[0042] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] This invention provides an edge-cloud collaborative image rights confirmation method based on adaptive feature compression: the entire process covers edge device feature processing, edge server feature reconstruction and hash generation, blockchain batch notarization and copyright verification, and the specific implementation steps are detailed as follows:
[0044] S1. Edge Device Front-End Feature Inference Implementation. The basic convolutional neural network (CNN) model is pre-segmented and deployed. The edge device is equipped with a lightweight front-end CNN model, receiving locally collected or stored image data to be authenticated. The images are pre-processed, including size normalization, pixel value standardization, and dimension padding, ensuring consistent image data dimensions for the input model. The pre-processed images are then input into the lightweight front-end CNN model to perform forward inference operations, extracting shallow and mid-level visual features layer by layer, and outputting a fixed-dimensional intermediate feature tensor. This intermediate feature tensor contains core weighting features such as image texture, contour, and color distribution, while retaining some redundant channel features to provide a data foundation for subsequent adaptive pruning.
[0045] S2. Implementation of Adaptive Feature Compression and Encoding in Edge Devices. Edge devices incorporate a sparse gating network module. Taking the intermediate feature tensor output from S1 as input, the sparse gating network adaptively filters and prunes feature channels, eliminating invalid and redundant channels and retaining only core feature channels with clear identity. After pruning, sparse feature data with reduced dimensions is obtained. This data is then losslessly compressed using an entropy coding algorithm to eliminate redundant information, resulting in smaller, lower-cost compressed feature data packets. This effectively reduces bandwidth consumption and transmission latency in subsequent network transmissions.
[0046] S3. Cross-end data transmission implementation. Edge devices transmit the generated compressed feature data packets unidirectionally to the corresponding edge server via wired or wireless network links. The transmission process uses an encrypted transmission protocol to prevent feature data from being stolen or tampered with, ensuring the security of rights confirmation data transmission. After the transmission is completed, the edge device can selectively clear the temporary feature data locally, saving local storage resources.
[0047] S4. Edge Server Feature Decoding and Reconstruction Implementation. After receiving the compressed feature data packet uploaded by the edge device, the edge server calls its built-in entropy decoding unit to perform lossless decoding of the compressed data packet according to the decoding rules corresponding to the entropy encoding of the edge device, restoring the pruned sparse feature map. The decoded sparse feature map is input into the trained feature recovery neural network model. Through the model's convolution operation and feature mapping mechanism, the pruned non-redundant feature information is supplemented, restoring a complete restored feature map with the same dimension and feature distribution as the original front-end inference output, ensuring the integrity and accuracy of subsequent hash feature generation.
[0048] S5. Deep Hash Feature Generation Implementation. The edge server inputs the reconstructed recovered features into the deep hash processing module. A fully connected layer maps the high-dimensional features to a low-dimensional intermediate vector, removing redundant information and retaining the image's unique feature representation. Based on the mapped intermediate vector, hash encoding is performed to generate a fixed-length binary feature hash code uniquely corresponding to each image. This hash code serves as the image's unique digital fingerprint and is the core basis for subsequent batch notarization and copyright verification.
[0049] S7. Blockchain Copyright Verification Implementation. Upon receiving images to be verified from scenarios such as platform review, infringement complaints, and originality verification, the entire feature processing steps from S1 to S5 are repeated. Front-end feature inference, adaptive compression, transmission decoding, feature reconstruction, and deep hash encoding are performed on the images to be verified to generate the target feature hash code. Based on the unique blockchain transaction ID of the batch to which the image belongs, the corresponding Merkle root hash data and the corresponding Merkle verification path stored in the blockchain network are retrieved. The root hash value is calculated iteratively layer by layer using the target feature hash code combined with the Merkle path node information. The calculation result is compared with the original root hash stored on the chain to complete the verification of the image's copyright authenticity and whether it has been tampered with.
[0050] S6. Batch Merkle Tree Construction and Blockchain Notarization Implementation. Edge servers continuously aggregate the binary feature hash codes of all images to be notified within the same batch, batch collecting the creation timestamp, author identity information, and work identification information corresponding to each image. Using the feature hash code of each image as a leaf node of the Merkle tree, the hash value of the parent node is calculated by merging pairs of each image layer by layer according to the Merkle tree iterative hash calculation rules, iterating upwards level by level, ultimately generating a unique Merkle tree root hash. The generated Merkle tree root hash, batch timestamp, author rights confirmation information, and work batch information are uniformly packaged and encapsulated into standard blockchain transaction data, submitted to the blockchain network to complete the on-chain writing. After the blockchain network confirms the transaction, a unique transaction ID is generated, completing the batch copyright notarization of the entire batch of images. All on-chain data possesses the characteristics of being tamper-proof and traceable.
[0051] The core of this invention lies in constructing a collaborative computing and evidence storage closed loop of "lightweight compression at the edge, precise recovery at the edge, and trusted evidence storage at the cloud." Through a three-level system division of labor, the computing load and data flow are intelligently distributed, systematically solving the problems of limited edge resources, high transmission costs, and lack of trust in rights confirmation.
[0052] On the edge device side, i.e., resource-constrained terminals, a lightweight front-end CNN is deployed to extract image features after segmentation, and an adaptive feature compression module based on sparse gating is introduced. This module dynamically selects key feature channels according to the input image content to achieve "on-demand compression," and then uploads the image after lossless entropy encoding, minimizing the amount of data transmission at the source.
[0053] At the edge (edge server), after receiving compressed data, the server first performs entropy decoding, and then reconstructs high-fidelity complete features using a lightweight feature recovery neural network. Subsequently, the complete features are used for subsequent CNN inference. The inference result is flattened into a one-dimensional feature vector, and then deep hashing is used to map the feature vector into a fixed-length binary hash code, which serves as the unique digital fingerprint of the image, ensuring robustness against content-preserving modifications.
[0054] On the cloud side (blockchain network), edge servers efficiently organize the hash codes of a batch of images into a Merkle tree data structure, submitting only the root hash of the tree and the evidence metadata to the blockchain. This reduces the evidence storage cost of N images to a single transaction, enabling low-cost, verifiable, and tamper-proof batch registration of massive amounts of images.
[0055] During verification, the image to be checked undergoes the same compression, recovery, and hashing process described above to obtain the target fingerprint. This fingerprint is then quickly calculated and compared with the corresponding Merkle path stored on the blockchain, thus completing the copyright confirmation and traceability of a single image without needing to query the original batch data.
[0056] The overall process is as follows:
[0057] CNN model segmentation: Determine the segmentation points of the CNN model. The original large CNN model is divided into a lighter front-end CNN deployed on the edge device and a back-end CNN deployed on the edge server, according to the computing power limit that the edge device can bear.
[0058] Joint pre-training: The front-end feature extraction CNN, adaptive feature compression module (including sparse gating), feature recovery network, back-end CNN, and deep hashing module are sequentially connected to form a differentiable complete feedforward link. Using a training dataset containing a large number of images, a corresponding composite loss function is designed, and joint training is implemented.
[0059] System initialization: A lightweight image feature extraction model (including a segmented front-end CNN model and an adaptive feature compression module) deployed on edge devices, a feature recovery model and a deep hashing module deployed on edge servers, and a blockchain network evidence storage interface.
[0060] Copyright registration process:
[0061] Step 1: The edge device receives the image to be stored and extracts the intermediate feature tensor through a lightweight CNN in the front end;
[0062] Step 2: The extracted intermediate feature tensors are pruned through the adaptive feature compression module to remove secondary channels, and the retained features are entropy encoded to generate compressed data packets and transmit them to the edge server.
[0063] Step 3: After decoding by the edge server, the complete features are reconstructed through the feature recovery model, flattened into a one-dimensional feature vector after being calculated by the subsequent CNN convolution, and then a binary hash code is generated using deep hashing.
[0064] Step 4: Collect hash codes from the same batch, construct a Merkle tree, and write the root hash, timestamp, and author information into the blockchain.
[0065] 5) Copyright verification process:
[0066] Step 1: Obtain the image to be verified, and generate the target hash code through feature extraction, compression, restoration, and hashing;
[0067] Step 2: Retrieve the Merkle path and on-chain root hash of the batch to which the image belongs;
[0068] Step 3: Calculate the root hash using the target hash code and the Merkle path, compare it with the on-chain record, and determine the verification result.
[0069] System architecture deployment such as Figure 2 As shown, the system comprises three core components that work together to complete the entire process of evidence storage and verification:
[0070] Edge devices (endpoints): Deploy lightweight feature extraction and compression modules, including CNN front-end models, adaptive feature compression modules, and entropy coding units, supporting local low-power computation.
[0071] Edge server (edge): Deploys the feature recovery and hash calculation module, including entropy decoding unit, lightweight feature recovery network, CNN back-end model, deep hashing unit, and Merkle tree construction and management module.
[0072] Blockchain network (cloud): Adopts public chain or consortium chain architecture, provides immutable transaction storage function, and only receives Merkle root hash and metadata sent by edge servers.
[0073] The process from edge devices to edge servers, such as Figure 3 As shown, model segmentation and joint pre-training
[0074] Step 1: Select a CNN model (such as ResNet-50), and determine the segmentation point based on the computing power ratio of edge devices and edge servers to divide the model into front and back segments.
[0075] Step 2: Insert an adaptive feature compression module after the predetermined segmentation point (after the front-end model). This module is followed by the feature recovery module, the back-end model, and the deep hashing module. The overall model is jointly pre-trained using the training dataset (including images of various scenes).
[0076] Joint pre-training aims to achieve a balance between compression and reconstruction, and between discriminability and sparsity, through collaborative optimization. Specific objectives include: ensuring high-fidelity reconstruction of compressed features to minimize information loss; improving the discriminability of deep hash codes so that semantically similar images have similar hash representations; controlling the amount of data transmitted through sparsity constraints to achieve efficient compression; and enhancing the model's robustness to content-preserving modifications to ensure stable performance under real-world perturbations.
[0077] Loss function design: Joint training uses a multi-task loss function, with a total loss of:
[0078]
[0079] in, To balance the hyperparameters.
[0080] For feature reconstruction loss, mean squared error (MSE) is used, as shown in the following formula:
[0081]
[0082] in For batch size, These represent the number of channels, height, and width of the feature map, respectively.
[0083] For depth hashing loss, triplet loss is used to encourage similar images to have similar hash codes and dissimilar images to have distant hash codes. Anchor images are provided. Positive samples and Semantic similarity, negative samples (and Semantically dissimilar), their hash codes are respectively ,but:
[0084]
[0085] in Cosine distance This is the interval hyperparameter.
[0086] For sparse constraint loss, such as Regularization, applied to the output vector of a sparse gated network. This is used to control the sparsity of gated outputs, as detailed below:
[0087]
[0088] in For batch size, This represents the number of channels.
[0089] To mitigate the robustness loss, two enhanced versions are obtained by applying two different slight perturbations (such as scaling, rotation, and noise) to the same image. and To encourage consistency in their hash codes:
[0090]
[0091] in This represents the hash code output by the model (a continuous approximation during training). The distance is Euclidean.
[0092] The model training process consists of three steps:
[0093] Step 1: In the initialization phase, the front-end CNN, back-end CNN, and deep hashing module are initialized using pre-trained weights; the sparse gating network and feature recovery network are randomly initialized.
[0094] Step 2: Phased training. Phase 1: Fix the front-end and back-end CNNs, and train only the sparse gating network and feature recovery network, focusing on feature compression and reconstruction capabilities. Phase 2: Unfreeze all modules and perform end-to-end joint training to optimize the total loss function. Phase 3: Introduce adversarial examples and data augmentation to further enhance the robustness and generalization ability of the model.
[0095] Step 3: Use large-scale image datasets (such as ImageNet, COCO, etc.) and construct specialized datasets that include copyright-related scenarios (such as artistic creations, news photos, commercial photography, etc.) to enhance task adaptability.
[0096] Adaptive feature compression computation, such as Figure 4 As shown, this is the key process of this patent. Let the input feature tensor be... ( Batch size Number of channels (Feature map size), the module flow is as follows:
[0097] Step 1: First, global average pooling is performed for feature aggregation, for each batch of samples. and channels Summarize spatial dimension features:
[0098]
[0099] The generated batch summary characteristics are
[0100] Step 2: The channel importance is quantified using a Lightweight Gated Network (MLP), which is essentially two layers of fully connected neural networks.
[0101]
[0102] in (Dimensional reduction weights, To reduce the dimensionality ratio, the default is 16). (Restore dimensional weights) Let be the gating weight vector. Its overall expression is: .
[0103] Step 3: Differentiable Top-K sparsification, generating importance sparse masks m, and finally only... The channel characteristics are used for subsequent processing and transmission.
[0104] The process of generating the importance sparse mask m uses the Gumbel-Softmax method to ensure gradient propagation. First, Gumbel noise is added: Generate noise gating values:
[0105]
[0106] in For the sigmoid function, This is a temperature parameter (controlling the sharpness of the distribution).
[0107] Use continuous approximation during the training phase:
[0108]
[0109] Hard choice during the reasoning phase:
[0110]
[0111] in To reserve the number of channels for the target, in one instance, By presetting the global sparsity (e.g., S=0.75) and total number of channels The calculation shows that, Under this fixed K-value constraint, the system adaptively and dynamically selects the K most important channels based on the content of the input image through a differentiable Top-K mechanism, thereby achieving a balance between feature compression rate and accuracy.
[0112] Step 4: Feature channel filtering, i.e., feature compression, applies the sparse mask generated in the previous step to the original features:
[0113]
[0114] Vectorized representation as ( (For channel-wise multiplication), sparse features are obtained. This sparse feature This is the K-channel feature map after channel pruning.
[0115] Feature compression and recovery:
[0116] Step 1: Edge devices: Based on the attention map, retain the Top-K channels, perform entropy encoding (such as Huffman encoding) on the pruned feature data, and generate compressed data packets;
[0117] Step 2: Edge server: Entropy decoding is performed on the received data packets, and the decoded features are input into the pre-trained feature recovery CNN. This network learns to predict approximate values of the pruned channels and outputs the reconstructed high-dimensional features.
[0118] The specific structure of this pre-trained feature recovery CNN is as follows:
[0119]
[0120]
[0121]
[0122]
[0123] Depth hashing and Merkle tree construction, such as Figure 5 , Figure 6 As shown.
[0124] Deep hashing
[0125] Nonlinear transformation: The reconstructed features are input into the fully connected layer and mapped to an intermediate vector of length K. ;
[0126] Binarization: In the inference stage, through symbolic functions Generate binary hash code Then convert ;
[0127] Training optimization: During training, the following methods are used: approximate function( (This is a scaling factor) to facilitate backpropagation.
[0128] Batch evidence storage process
[0129] Step 1: The edge server collects the hash codes, timestamps, and author information of the same batch (e.g., 1000 images);
[0130] Step 2: Construct a Merkle tree using hash codes as leaf nodes and calculate the root hash. ;
[0131] Step 3: Root hash The batch metadata is packaged into a blockchain transaction, written to the blockchain network, and a unique transaction ID is generated.
[0132] Copyright verification process:
[0133] Step 1: Repeat the "feature extraction-compression-decoding-restoration-hashing" process for the image to be verified to generate the target hash code. ;
[0134] Step 2: Query the on-chain root hash based on the transaction ID The Merkle path (e.g., the set of adjacent hash nodes) corresponding to the image;
[0135] Step 3: Utilize Calculate the root hash layer by layer with the Merkle path ,like If the verification is successful (copyright valid), then the copyright is valid; if there is no discrepancy, then the copyright is deemed to have been infringed or the content has been altered.
[0136] This invention utilizes mainstream image feature extraction networks as the basic convolutional neural network model, including but not limited to ResNet, MobileNet, and VGG, which are suitable for visual feature extraction. Before model deployment, model segmentation is performed based on the hardware computing power ratio between edge devices and edge servers. The computing power ratio indicators include core parameters such as device CPU computing power, GPU computing power, memory size, power consumption, and parallel processing capability. The optimal model segmentation point is determined through quantitative calculations of the computing power ratio, splitting the complete basic convolutional neural network model into two segments: the front-end structure simplifies the number of convolutional layers, reduces the number of parameters and computational load, constructing a lightweight convolutional neural network model suitable for deployment and operation on edge devices with low computing power, primarily completing shallow and mid-level image feature extraction; the back-end structure retains the complete deep convolution, pooling, and feature mapping structures, deployed on edge servers with sufficient computing power to cooperate with the front-end features to complete complete feature representation. Model segmentation achieves hierarchical adaptation of computing power, avoiding inference latency and device lag issues caused by insufficient computing power on edge devices, while fully leveraging the high-performance computing capabilities of edge servers to ensure the efficient operation of the overall rights confirmation process.
[0137] The present invention and its embodiments have been described above. This description is not restrictive. In short, if a person skilled in the art is inspired by this description and designs a similar structure and embodiment without departing from the spirit of the present invention, such design should fall within the protection scope of the present invention.
Claims
1. A collaborative image rights confirmation method based on adaptive feature compression, characterized in that, Includes the following steps: S1. On the edge device, the segmented front-end lightweight convolutional neural network model is used to perform front-end inference on the local image to be stored, and the intermediate feature tensor is obtained. S2. On the edge device, based on the sparse mask output by the sparse gated network, channel pruning is performed on the intermediate feature tensor, and entropy encoding is performed on the pruned feature data to generate compressed feature data packets. S3. Send the compressed feature data packet to the edge server; S4. On the edge server, the compressed feature data packet is entropy decoded, and the decoded features are reconstructed using a feature recovery neural network model to obtain the recovered features; S5. On the edge server, perform deep hashing on the recovery features to generate the binary feature hash code of the image; S6. Construct a Merkle tree from the feature hash codes of multiple images belonging to the same batch, and write the root hash of the Merkle tree into the blockchain to complete the copyright certificate. S7. After receiving the image to be verified, verify its copyright information on the blockchain.
2. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The convolutional neural network model segmentation step includes: selecting a basic convolutional neural network model, determining the segmentation point based on the computing power ratio between edge devices and edge servers, and segmenting the basic convolutional neural network model into a front-end lightweight convolutional neural network deployed on edge devices and a back-end convolutional neural network deployed on edge servers.
3. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The sparse gated network employs a differentiable Top-K sparsity mechanism, which includes first performing global average pooling on the input intermediate feature tensor to generate channel statistics; inputting the channel statistics into a lightweight gated network to output a gate weight vector; and introducing Gumbel noise to generate a sparse mask using the Gumbel-Softmax method. During the training phase, a continuous approximation method is used to generate sparse masks. During the inference phase, a hard Top-K selection method is used to generate sparse masks, retaining only the top K feature channels with the highest weights. K is calculated from the preset global sparsity rate and the total number of channels.
4. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The feature recovery neural network model is a lightweight convolutional structure. Its processing includes inputting a pruned sparse feature map, performing channel upsampling through a 1×1 convolution, and extracting features through a 3×3 convolution combined with the ReLU activation function. Then, a 1×1 convolution is used to restore the original number of channels, and the reconstructed complete feature map is output.
5. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The deep hashing process includes inputting the recovered features into a fully connected layer and mapping them to an intermediate vector; during the training phase, the tanh function is used to approximate the sign function to achieve backpropagation; during the inference phase, the sign function is used to generate a binary hash code, which is then converted into a 0-1 format hash code as the unique digital fingerprint of the image.
6. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The Merkle tree construction and blockchain evidence storage steps include edge servers collecting feature hash codes, timestamps, and author information for the same batch of images; A Merkle tree is constructed using the characteristic hash code as the leaf node, and the Merkle tree root hash is calculated. The Merkle tree root hash, timestamp, and author information are packaged into a blockchain transaction, written into the blockchain network, and a unique transaction ID is generated.
7. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The copyright verification step includes repeating the S1-S5 process on the image to be verified to generate a target feature hash code. Based on the transaction ID, query the Merkle root hash and the Merkle path corresponding to the image to be verified stored on the blockchain; The root hash is calculated layer by layer using the target feature hash code and Merkle path, and compared with the root hash stored in the blockchain. If they match, the verification is successful; otherwise, the verification fails.
8. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that: The method also includes a joint model pre-training step, specifically: sequentially connecting the front-end lightweight convolutional neural network, sparse gating network, feature recovery neural network, back-end convolutional neural network, and deep hashing module to form a complete feedforward link; training is performed using a composite loss function, which includes feature reconstruction loss, deep hashing loss, sparse constraint loss, and robustness loss; a three-stage training strategy is adopted: in the first stage, the front-end and back-end convolutional neural networks are fixed, and only the sparse gating network and feature recovery neural network are trained; The second stage unfreezes all modules and performs end-to-end joint training; the third stage introduces adversarial examples and data augmentation to enhance the model's robustness.
9. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 1, characterized in that, The method includes the following system: at the edge device end, a front-end lightweight convolutional neural network, a sparse gated network, and an entropy coding unit are deployed to extract intermediate feature tensors from images, adaptively prune redundant feature channels, and generate compressed data packets using entropy coding; On the edge server side, an entropy decoding unit, a feature recovery neural network, a post-convolutional neural network, a deep hashing unit, and a Merkle tree construction and management module are deployed to receive compressed data packets, decode and reconstruct features, generate binary feature hash codes, and build Merkle trees in batches. On the blockchain network side, a public chain or consortium chain architecture is adopted to store Merkle tree root hashes, timestamps, and author information, providing tamper-proof copyright notarization and verification services.
10. The image edge-cloud collaborative rights confirmation method based on adaptive feature compression according to claim 9, characterized in that: The image edge-cloud collaborative rights confirmation method, which involves the collaborative execution of adaptive feature compression by the edge device, edge server, and blockchain network, is described.