Image decryption method and device, computer equipment and storage medium

By standardizing encrypted images, extracting spatial and temporal features, performing cross-attention fusion, and optimizing tensor neural networks, the problem of low decryption efficiency caused by relying on prior knowledge of encryption algorithms in existing technologies is solved, achieving efficient image decryption that is applicable to the financial and medical fields.

CN120856833APending Publication Date: 2025-10-28PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510683139.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing image decryption methods rely on prior knowledge of encryption algorithms, resulting in low decryption efficiency and an inability to adapt to unknown or dynamically changing encryption environments. This is especially true in cross-border payment credential encryption scenarios in the financial sector, where low decryption efficiency can lead to business process interruptions.

Method used

A preprocessing module is used to standardize the encrypted image, extract features through spatial and temporal paths, perform feature fusion using a cross-attention gating module, and optimize the image using a tensor neural network. Finally, the target decoder reconstructs the decrypted image, achieving efficient decryption without requiring encryption algorithm parameters or key information.

Benefits of technology

It enables efficient and accurate decryption of encrypted images without relying on prior knowledge of the encryption algorithm, improving the efficiency and adaptability of image decryption, and is suitable for image decryption scenarios in the financial and medical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856833A_ABST
    Figure CN120856833A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an image decryption method and device, computer equipment and a storage medium. Standardizing the encrypted image based on a preprocessing module to obtain a target encrypted image; performing spatial feature extraction on the target encrypted image based on a spatial path processing strategy to obtain spatial features; performing time sequence feature extraction on the target encrypted image based on a time sequence path processing strategy to obtain time sequence features; performing dynamic feature fusion on the spatial features and the time sequence features based on a cross-attention gating module to obtain fusion features; performing optimization processing on the fusion features based on a tensor neural network to obtain optimized feature information; and performing image reconstruction processing on the optimized feature information based on the target decoder to obtain a target decrypted image. In addition, the target decryption image can be stored in the block chain. The image decryption method and device can be applied to image decryption scenes in the financial field and the medical field, and the decryption efficiency of image decryption is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to image decryption methods, devices, computer equipment, and storage media. Background Technology

[0002] In traditional image encryption technology systems, encryption schemes based on algorithms such as AES and RSA, while possessing high security, rely on strict key management systems and predefined algorithm logic, resulting in drawbacks such as complex key distribution and high management costs. Existing decryption models (such as deep learning methods based on ViT) typically require pre-obtaining the parameters or structural information of the encryption algorithm (such as key generation rules and permutation patterns), which limits their generalization ability, enabling them to decrypt only specific encryption schemes and failing to adapt to unknown or dynamically changing encryption environments.

[0003] For example, in the context of cross-border payment credential encryption in the financial sector, insurance companies or banks may use self-developed hybrid encryption algorithms (such as a customized solution combining AES and chaotic mapping) to encrypt electronic remittance slips. However, existing decryption models cannot resolve the proprietary design of these algorithms, requiring manual reverse engineering or authorization from the algorithm provider, resulting in low decryption efficiency. Such technical bottlenecks not only prolong the decryption cycle of sensitive information (such as the review time for cross-border transaction credentials) but may also lead to business process interruptions due to decryption failures, impacting the operational efficiency of financial institutions and customer trust.

[0004] Therefore, there is an urgent need to propose a general image decryption method that does not rely on prior knowledge of encryption algorithms, so as to achieve efficient decryption of unknown encrypted images and improve the timeliness of data circulation. Summary of the Invention

[0005] The purpose of this application is to provide an image decryption method, apparatus, computer device, and storage medium to solve the technical problem that existing image decryption methods rely on prior knowledge of encryption algorithms, resulting in low decryption efficiency.

[0006] Firstly, an image decryption method is provided, including:

[0007] Obtain the encrypted image to be processed;

[0008] The encrypted image is standardized based on a preset preprocessing module to obtain the corresponding target encrypted image;

[0009] Based on a preset spatial path processing strategy, spatial features are extracted from the target encrypted image to obtain the corresponding spatial features.

[0010] Based on a preset temporal path processing strategy, temporal features are extracted from the target encrypted image to obtain the corresponding temporal features;

[0011] Based on a preset cross-attention gating module, the spatial features and the temporal features are dynamically fused to obtain the corresponding fused features;

[0012] The fused features are optimized based on a preset tensor neural network to obtain corresponding optimized feature information;

[0013] Based on a preset target decoder, the optimized feature information is used to perform image reconstruction processing to obtain a target decrypted image corresponding to the encrypted image.

[0014] Secondly, an image decryption device is provided, comprising:

[0015] The acquisition module is used to acquire the encrypted image to be processed;

[0016] The processing module is used to perform standardization processing on the encrypted image based on a preset preprocessing module to obtain the corresponding target encrypted image;

[0017] The first extraction module is used to extract spatial features from the target encrypted image based on a preset spatial path processing strategy to obtain the corresponding spatial features.

[0018] The second extraction module is used to extract temporal features from the target encrypted image based on a preset temporal path processing strategy to obtain the corresponding temporal features;

[0019] The fusion module is used to dynamically fuse the spatial features and the temporal features based on a preset cross-attention gating module to obtain the corresponding fused features;

[0020] An optimization module is used to optimize the fused features based on a preset tensor neural network to obtain corresponding optimized feature information;

[0021] The reconstruction module is used to perform image reconstruction processing on the optimized feature information based on a preset target decoder to obtain a target decrypted image corresponding to the encrypted image.

[0022] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described image decryption method.

[0023] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described image decryption method.

[0024] In the above-mentioned image decryption method, apparatus, computer equipment, and storage medium, the following steps are taken: First, an encrypted image to be processed is acquired; then, the encrypted image is standardized based on a preset preprocessing module to obtain a corresponding target encrypted image; next, spatial features are extracted from the target encrypted image based on a preset spatial path processing strategy to obtain corresponding spatial features; and temporal features are extracted from the target encrypted image based on a preset temporal path processing strategy to obtain corresponding temporal features; then, the spatial features and temporal features are dynamically fused based on a preset cross-attention gating module to obtain corresponding fused features; subsequently, the fused features are optimized based on a preset tensor neural network to obtain corresponding optimized feature information; finally, the optimized feature information is reconstructed based on a preset target decoder to obtain a target decrypted image corresponding to the encrypted image. This application, after obtaining the encrypted image to be processed, standardizes the encrypted image using a preprocessing module to obtain the target encrypted image. Then, it extracts spatial features from the target encrypted image using a spatial path processing strategy, and extracts temporal features using a temporal path processing strategy. Next, it dynamically fuses the spatial and temporal features using a cross-attention gating module to obtain fused features. Then, it optimizes the fused features using a tensor neural network to obtain optimized feature information. Finally, it uses a target decoder to perform image reconstruction processing on the optimized feature information, thereby efficiently and accurately obtaining the target decrypted image corresponding to the encrypted image. This application, through the combined use of a preprocessing module, spatial path processing strategy, temporal path processing strategy, cross-attention gating module, tensor neural network, and target decoder, requires only the encrypted image as input, without any encryption algorithm parameters or key information. It can achieve high-quality and efficient image decryption through spatiotemporal joint modeling and dynamic optimization without relying on prior knowledge of the encryption algorithm, effectively improving the decryption efficiency of images. Attached Figure Description

[0025] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0027] Figure 2 This is a flowchart of an embodiment of the image decryption method according to this application;

[0028] Figure 3 This is a schematic diagram of the structure of an embodiment of the image decryption apparatus according to this application;

[0029] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0031] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0033] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0034] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0035] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0036] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0037] It should be noted that the image decryption method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the image decryption device is generally located in the server / terminal device.

[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0039] Continue to refer Figure 2 A flowchart illustrating an embodiment of the image decryption method according to this application is shown. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The image decryption method provided by this application embodiment can be applied to any scenario requiring image decryption, and thus can be applied to products in these scenarios, such as image decryption scenarios in the financial and medical fields. The image decryption method includes the following steps:

[0040] Step S201: Obtain the encrypted image to be processed.

[0041] In this embodiment, the image decryption method operates on an electronic device (e.g., Figure 1The server / terminal device shown can acquire the encrypted image to be processed via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The executing entity of this application is specifically an image decryption system, which can be simply referred to as the system. The image decryption system includes a preprocessing module, a cross-attention gating module, a tensor neural network (or ZNN dynamic optimization layer), and a target decoder (i.e., a U-Net decoder). The system's workflow is as follows: The system takes an encrypted image as input and first performs standardization processing through the preprocessing module, zero-padding and block segmentation of the input image. Subsequently, the system employs a dual-path architecture to extract features of the encrypted image in parallel: the spatial path extracts multi-scale spatial features based on an improved ResNet-50 architecture; the temporal path uses a Transformer encoder to model the relationship between image block sequences. Next, the innovative cross-attention gating module (CAGM) receives the features from both paths and achieves dynamic fusion. Subsequently, the ZNN dynamic optimization layer adjusts the convergence direction of the decryption process in real time by constructing a system of time-varying differential equations. Finally, the U-Net decoder integrates and fuses features, reconstructing the decrypted image through multi-level upsampling.

[0042] This application can be applied to image decryption processing scenarios in the financial and medical fields. For example, in the financial field, the encrypted images may include encrypted scans of paper insurance policies, accident scene photos, invoices, and check anti-counterfeiting patterns. In the medical field, the encrypted images may include encrypted medical images (such as CT slices, MRI slices), full-view digital slices, and dermoscopy photographs.

[0043] Step S202: The encrypted image is standardized based on a preset preprocessing module to obtain the corresponding target encrypted image.

[0044] In this embodiment, the specific implementation process of standardizing the encrypted image based on the preset preprocessing module to obtain the corresponding target encrypted image will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0045] Step S203: Based on a preset spatial path processing strategy, spatial features are extracted from the target encrypted image to obtain the corresponding spatial features.

[0046] In this embodiment, the core of the aforementioned spatial path processing strategy is to extract multi-scale spatial features from the encrypted image. The specific implementation process of extracting spatial features from the target encrypted image based on the preset spatial path processing strategy to obtain the corresponding spatial features will be described in further detail in subsequent embodiments of this application, and will not be elaborated upon here.

[0047] Step S204: Based on a preset temporal path processing strategy, temporal features are extracted from the target encrypted image to obtain the corresponding temporal features.

[0048] In this embodiment, the aforementioned temporal path processing strategy aims to model the temporal dependencies between encrypted image blocks, which is crucial for understanding the sequence transformations during image encryption. The specific implementation process of extracting temporal features from the target encrypted image based on the preset temporal path processing strategy to obtain the corresponding temporal features will be further described in detail in subsequent embodiments of this application, and will not be elaborated upon here.

[0049] Step S205: Based on the preset cross-attention gating module, the spatial features and the temporal features are dynamically fused to obtain the corresponding fused features.

[0050] In this embodiment, the Cross-Attention Gating Module (CAGM) is the core innovation of this application, solving the problem of heterogeneous fusion of spatial and temporal features. The CAGM employs a three-stage processing flow: first, it enhances the feature representations of each path; then, it performs adaptive weight fusion. The specific implementation process of dynamically fusing the spatial and temporal features based on the preset CAGM to obtain the corresponding fused features will be further described in detail in subsequent embodiments of this application, and will not be elaborated upon here.

[0051] Step S206: Optimize the fused features based on a preset tensor neural network to obtain corresponding optimized feature information.

[0052] In this embodiment, the tensor neural network is specifically a spatiotemporal attention tensor neural network (ST-ZNN), or a ZNN dynamic optimization layer. The ZNN dynamic optimization layer represents a significant breakthrough in traditional neural network decryption methods, transforming the static optimization paradigm into a time-varying dynamic system. The specific implementation process of optimizing the fused features based on a preset tensor neural network to obtain the corresponding optimized feature information will be further described in detail in subsequent embodiments and will not be elaborated upon here.

[0053] Step S207: Based on a preset target decoder, perform image reconstruction processing on the optimized feature information to obtain a target decrypted image corresponding to the encrypted image.

[0054] In this embodiment, the specific implementation process of performing image reconstruction processing on the optimized feature information based on the preset target decoder to obtain the target decrypted image corresponding to the encrypted image will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0055] This application first acquires the encrypted image to be processed; then, it performs standardization processing on the encrypted image based on a preset preprocessing module to obtain the corresponding target encrypted image; subsequently, it extracts spatial features from the target encrypted image based on a preset spatial path processing strategy to obtain the corresponding spatial features; and extracts temporal features from the target encrypted image based on a preset temporal path processing strategy to obtain the corresponding temporal features; then, it performs dynamic feature fusion of the spatial features and the temporal features based on a preset cross-attention gating module to obtain the corresponding fused features; subsequently, it optimizes the fused features based on a preset tensor neural network to obtain the corresponding optimized feature information; finally, it performs image reconstruction processing on the optimized feature information based on a preset target decoder to obtain the target decrypted image corresponding to the encrypted image. This application, after obtaining the encrypted image to be processed, standardizes the encrypted image using a preprocessing module to obtain the target encrypted image. Then, it extracts spatial features from the target encrypted image using a spatial path processing strategy, and extracts temporal features using a temporal path processing strategy. Next, it dynamically fuses the spatial and temporal features using a cross-attention gating module to obtain fused features. Then, it optimizes the fused features using a tensor neural network to obtain optimized feature information. Finally, it uses a target decoder to perform image reconstruction processing on the optimized feature information, thereby efficiently and accurately obtaining the target decrypted image corresponding to the encrypted image. This application, through the combined use of a preprocessing module, spatial path processing strategy, temporal path processing strategy, cross-attention gating module, tensor neural network, and target decoder, requires only the encrypted image as input, without any encryption algorithm parameters or key information. It can achieve high-quality and efficient image decryption through spatiotemporal joint modeling and dynamic optimization without relying on prior knowledge of the encryption algorithm, effectively improving the decryption efficiency of images.

[0056] In some alternative implementations, step S203 includes the following steps:

[0057] Invoke the preset deep residual network.

[0058] In this embodiment, the aforementioned deep residual network can specifically employ the ResNet-50 network. ResNet-50 is a variant of the ResNet series, featuring 50 convolutional layers. ResNet addresses the vanishing gradient problem in deep network training by introducing residual connections, enabling deeper training and improved performance. ResNet-50 consists of multiple residual blocks, each containing a shortcut connection that directly connects the input to the output, forming a residual unit. This design allows the network to learn the residuals between the input and output during training, helping to alleviate the vanishing gradient problem and enabling the network to learn more complex feature representations.

[0059] The deep residual network is improved based on a preset module introduction strategy to obtain the corresponding target deep residual network.

[0060] In this embodiment, the strategy introduced by the above module includes: introducing an SE-ResBlock structure, embedding the compression-excitation module into the standard residual block, which enables the network to adaptively adjust the importance of feature channels. Specifically, for each channel c, we first obtain its statistical information z through global average pooling. c Then, a two-layer fully connected network is applied to model the interdependencies between channels, expressed by the formula:

[0061] s c =σ(W2δ(W1z) c ))

[0062] Among them, s c W1 and W2 are the recalibration coefficients for channel c, respectively. The compression ratio r is set to 16 to achieve a balance between computational efficiency and performance. The δ function uses ReLU activation (the output equals the input when the input is positive, otherwise the output is zero), while σ uses the Sigmoid function (mapping the input to between 0 and 1) to ensure a reasonable range for the recalibration coefficients. Furthermore, the architecture of the deep residual network can be improved based on the strategy introduced in the above modules, thereby constructing the desired target deep residual network.

[0063] Obtain a pre-defined multi-scale feature extraction strategy.

[0064] In this embodiment, the aforementioned multi-scale feature extraction strategy is a 5-level layer-by-layer downsampling feature extraction strategy constructed to capture multi-scale spatial information. The first level uses a 7×7 convolution kernel with a stride of 2 to downsample the input image to half its original resolution. Subsequently, the fourth level uses a 3×3 convolution kernel with a stride of 2 to progressively reduce the feature map size to 1 / 32 of its original resolution. Specifically, in the deeper layers of the network, we introduce a dilated spatial pyramid pooling (ASPP) module. By using convolution operations with dilation rates of 6, 12, and 18 in parallel, we expand the receptive field while maintaining computational efficiency. The dilation rate refers to the number of pixels between adjacent elements in the convolution kernel. A larger dilation rate can obtain broader contextual information without increasing the number of parameters, which is crucial for parsing the global structure in encrypted images.

[0065] Based on the target depth residual network, the multi-scale feature extraction strategy is used to extract spatial features from the target encrypted image to obtain the corresponding multi-scale spatial features.

[0066] In this embodiment, spatial features can be extracted from the target encrypted image based on the target depth residual network and the strategy content of the multi-scale feature extraction strategy, thereby obtaining the corresponding multi-scale spatial features.

[0067] The multi-scale spatial features are used as the spatial features.

[0068] This application improves a target deep residual network by invoking a pre-defined deep residual network and then enhancing it using a pre-defined module introduction strategy. A pre-defined multi-scale feature extraction strategy is then obtained. Subsequently, based on the target deep residual network, the multi-scale feature extraction strategy is used to extract spatial features from the target encrypted image, yielding corresponding multi-scale spatial features. Finally, these multi-scale spatial features are used as the spatial features. This application improves the deep residual network using a module introduction strategy to obtain the target deep residual network. The combined use of the improved target deep residual network and the multi-scale feature extraction strategy enables efficient and accurate spatial feature extraction of the target encrypted image, improving the extraction efficiency and ensuring the accuracy of the obtained spatial features.

[0069] In some optional implementations of this embodiment, step S204 includes the following steps:

[0070] The target encrypted image is divided into blocks and projected to obtain the corresponding first image block.

[0071] In this embodiment, the input target encrypted image i can be divided into a series of 16×16 pixel non-overlapping blocks using a Transformer-based hierarchical architecture, each block containing 3 channels of RGB information. For each image block i p We use a learnable linear projection matrix W e Mapping it to a 256-dimensional latent space representation, we can express it as:

[0072] z p =W e ·vec(I p )+b e

[0073] Here, the vec(·) operation flattens a two-dimensional image patch into a one-dimensional vector. · represents matrix multiplication. Linear projection can be regarded as a preliminary abstraction of the original pixel information.

[0074] The first image block is positionally encoded based on a preset two-dimensional sinusoidal position encoding strategy to obtain the corresponding second image block.

[0075] In this embodiment, a two-dimensional sinusoidal position encoding mechanism is employed to preserve the positional information of image patches. For an image patch located at spatial coordinates (x, y), its position encoding uses sine and cosine functions in dimensions 2i and 2i+1, respectively:

[0076] PE (x,y,2i) =sin(x / 10000) (2i / D) )

[0077] PE (x,y,2i+1) =cos(y / 10000) (2i / D) )

[0078] This encoding method can reflect spatial distance relationships through frequency changes and has good extrapolation performance.

[0079] Specifically, the first image block can be positionally encoded based on the encoding method corresponding to the two-dimensional sinusoidal position encoding strategy to obtain the corresponding second image block.

[0080] Invoke the preset target Transformer encoder.

[0081] In this embodiment, the target Transformer encoder is an improved version of the initial Transformer encoder. Specifically, the core of the Transformer encoder is the self-attention mechanism, which this application improves for image decryption tasks. Standard multi-head self-attention allows direct interaction between any two positions, which can lead to redundant computation when processing long sequences. Therefore, this application introduces a local window attention mechanism, which restricts each position to focus only on image patches within a 7×7 neighborhood using a mask matrix M.

[0082]

[0083] Among them, K T This represents the transpose of matrix K. This design balances local detail awareness and computational efficiency, making it particularly suitable for processing regional transformation patterns in encrypted images. Meanwhile, the encoder retains a 12-layer stacked structure to ensure sufficient model capacity, with each layer containing 8 parallel attention heads to enhance the diversity of feature representations.

[0084] Temporal features are extracted from the second image block based on the target Transformer encoder to obtain the corresponding specified features.

[0085] In this embodiment, the temporal features of the second image block can be extracted by using an improved target Transformer encoder, and the obtained specified features can be used as the final temporal features.

[0086] The specified feature is used as the temporal feature.

[0087] This application obtains a first image block by segmenting and projecting the target encrypted image; then, it performs positional encoding on the first image block based on a preset two-dimensional sinusoidal positional encoding strategy to obtain a second image block; subsequently, it calls a preset target Transformer encoder; and extracts temporal features from the second image block based on the target Transformer encoder to obtain corresponding specified features; these specified features are then used as the temporal features. This application obtains a first image block by segmenting and projecting the target encrypted image, then performs positional encoding on the first image block based on a two-dimensional sinusoidal positional encoding strategy to obtain a second image block, and then extracts temporal features from the second image block based on an improved target Transformer encoder to obtain temporal features. This allows for efficient and accurate temporal feature extraction of the target encrypted image, improving the extraction efficiency and ensuring the accuracy and diversity of the obtained temporal features.

[0088] In some alternative implementations, step S205 includes the following steps:

[0089] The spatial features are recalibrated based on the cross-attention gating module to obtain the corresponding specified spatial features.

[0090] In this embodiment, the aforementioned cross-attention gating module adopts a three-stage processing flow. First, it enhances the feature representations of the spatial path and the temporal path, and then performs adaptive weight fusion. The enhancement of the spatial path's feature representation corresponds to the spatial feature recalibration stage. In the spatial feature recalibration stage, this application comprehensively considers the average response and maximum response of the feature channels. These two statistical information reflect the overall activity and saliency peak of the channel, respectively. Specifically, average pooling and max pooling are applied in parallel, and then concatenated along the channel dimension: [AvgPool(F s MaxPool(F) s Then, 1×1 convolutions are used to extract inter-channel relationships, and normalized weights are generated using the Sigmoid function. Finally, element-wise multiplication is used to apply the weights to the original spatial features.

[0091] F s ′ =Sigoid(f (1×1) ([AvgPool(F s MaxPool(F) s )]))⊙F s

[0092] Here, ⊙ represents the element-wise multiplication of two matrices of the same dimension, and [A; B] represents the concatenation of two tensors along a specific dimension. This design allows the model to dynamically focus on the discriminative information carried by different channels, enhancing the feature representation related to decryption. Furthermore, the spatial features described above can be recalibrated according to the processing steps corresponding to the spatial feature recalibration stage, thereby obtaining the corresponding specified spatial features.

[0093] The time-series features are enhanced to obtain the corresponding specified time-series features.

[0094] In this embodiment, the enhancement processing of temporal features leverages the advantages of the Transformer architecture, further optimizing the temporal features through residual connections and layer normalization: F′ t =LayerNorm(F t +FFN(MSA(F t Here, Multi-head Self-Attention (MSA) first captures the dependencies within the sequence, the Feedforward Network (FFN) further enhances the expressive power of the features, and residual connections and layer normalization help stabilize the training process and alleviate gradient problems.

[0095] Invoke the preset fusion network.

[0096] In this embodiment, the final dynamic feature fusion stage is the key to CAGM. This application designs a three-layer MLP-structured fusion network f fusion It receives the preprocessed spatial features F_s' and the shape-adjusted temporal features Reshape(F′). t This outputs a pixel-level fused weight map α. This weight map achieves an adaptive combination of spatial and temporal features:

[0097] F out =αF′ s +(1-α)Reshape(F′ t )

[0098] This approach allows the model to autonomously decide which features to emphasize at each location in the space, thereby making an optimal response to the encryption characteristics of different regions. For example, for regions with obvious spatial structure, the weights may be biased towards spatial features; while for regions that require temporal context understanding, the weights may be biased towards temporal features.

[0099] Based on the fusion network, the specified spatial features and the specified temporal features are dynamically fused to obtain the corresponding fusion weight map.

[0100] In this embodiment, the specified spatial features and specified temporal features can be fused by using the fusion network described above, and a pixel-level fusion weight map can be output as the final fusion feature.

[0101] The fusion weight map is used as the fusion feature.

[0102] In this embodiment, the innovative cross-attention gating module achieves adaptive fusion of spatial and temporal features, dynamically adjusts the weight of each path according to content characteristics, and improves decryption robustness in complex encryption scenarios.

[0103] This application obtains corresponding specified spatial features by recalibrating the spatial features based on the cross-attention gating module; then, it enhances the temporal features to obtain corresponding specified temporal features; subsequently, it calls a preset fusion network; and dynamically fuses the specified spatial features and the specified temporal features based on the fusion network to obtain a corresponding fusion weight map; finally, it uses the fusion weight map as the fused feature. This application obtains specified spatial features by recalibrating the spatial features based on the cross-attention gating module, and obtains specified temporal features by enhancing the temporal features. Then, it dynamically fuses the specified spatial features and specified temporal features based on the fusion network, thereby automatically and accurately achieving adaptive combination of spatial and temporal features, improving the intelligence and adaptability of the generated fused features, and enhancing decryption robustness in complex encryption scenarios.

[0104] In some alternative implementations, step S206 includes the following steps:

[0105] Obtain the pre-constructed error function corresponding to the tensor neural network.

[0106] In this embodiment, the tensor neural network is specifically a spatiotemporal attention tensor neural network (ST-ZNN), or a ZNN dynamic optimization layer. The ZNN dynamic optimization layer represents a significant breakthrough in traditional neural network decryption methods, transforming the static optimization paradigm into a time-varying dynamic system. The tensor neural network first constructs an error function E(t) to measure the difference between the prediction after feature fusion decoding and the current decryption result:

[0107] E(t)=F(F out (t))-I dec (t),

[0108] Here, F represents the feature decoding function, which projects the fused features onto the image space.

[0109] The corresponding dynamic error constraints are generated based on the error function.

[0110] In this embodiment, traditional methods typically minimize the sum of squared errors directly, while this application designs a dynamic error constraint based on the time derivative:

[0111] dI dec (t) / dt=γsign(E(t))min(|E(t)|,τ)

[0112] Here, d(·) / dt represents the derivative of (·) with respect to time t, |x| represents the absolute value of x, the sign function extracts the direction information of the error, min returns the minimum value, min(|E(t)|,τ) implements adaptive pruning of the error amplitude to prevent instability in the decryption process caused by extreme values, and γ is an adjustable parameter controlling the convergence speed. This design enables the decryption process to dynamically adjust the optimization direction and step size according to the current error situation, accelerating convergence while maintaining stability.

[0113] Obtain the solution strategy corresponding to the preset continuous-time differential equation.

[0114] In this embodiment, in order to solve the continuous-time differential equation in a digital computing environment, this application adopts the second-order Runge-Kutta method as the solution strategy for the continuous-time differential equation, which has higher numerical accuracy than the simple Euler method.

[0115] Based on the tensor neural network, the dynamic error constraints and the solution strategy are used to optimize the fused features in accordance with the direction and step size to obtain the corresponding optimization results.

[0116] In this embodiment, for each time step, the update amount of the decrypted image is determined through a two-stage calculation: first, the rate of change k1 in the current state is calculated; then, a more accurate rate of change k2 is estimated based on intermediate states; and finally, the decrypted image is updated.

[0117] i dec (n+1) =i dec n +k2

[0118] Where n represents the index of the discrete time step, I dec n Let represent the decrypted image at the nth time step. A fixed step size h of 0.5 is used to strike a balance between accuracy and computational efficiency. Experiments show that for a typical 256×256 image, excellent decryption results can be achieved in 50 iterations (approximately 0.2 seconds of computation time). Furthermore, by solving differential equations, a precise update direction can be provided for each iteration of the decrypted image, ensuring the stability and accuracy of the reconstruction process.

[0119] The optimization result is used as the optimization feature information.

[0120] This application obtains a pre-constructed error function corresponding to the tensor neural network; generates corresponding dynamic error constraints based on the error function; then obtains a solution strategy corresponding to a preset continuous-time differential equation; subsequently, based on the tensor neural network, it uses the dynamic error constraints and the solution strategy to optimize the fused features in accordance with the direction and step size, obtaining the corresponding optimization result; subsequently, the optimization result is used as the optimized feature information. This application obtains a pre-constructed error function corresponding to the tensor neural network, generates corresponding dynamic error constraints based on the error function, obtains a solution strategy corresponding to a preset continuous-time differential equation, and then uses the dynamic error constraints and solution strategy to optimize the fused features in accordance with the direction and step size based on the tensor neural network to intelligently and efficiently obtain the final optimized feature information. This application constructs a time-varying differential equation system based on the use of tensor neural networks, and achieves autonomous convergence of the decryption process through real-time adjustment of the error function, breaking through the limitations of traditional static optimization methods. This is beneficial for subsequent decryption processing of encrypted images based on optimized feature information, achieving high-quality decryption results.

[0121] In some optional implementations of this embodiment, step S207 includes the following steps:

[0122] The optimized feature information is input into the target decoder.

[0123] In this embodiment, the target decoder can specifically be a U-Net decoder.

[0124] The target decoder integrates the optimized feature information to obtain the corresponding hierarchical features.

[0125] In this embodiment, the optimized feature information obtained after optimization by the tensor neural network can be integrated with multi-scale feature information by using the target decoder described above, so as to transform the optimized feature information into a hierarchical feature representation suitable for reconstruction.

[0126] Obtain the preset multi-level upsampling strategy.

[0127] In this embodiment, the multi-level upsampling strategy includes: restoring the low-resolution feature map to a high-resolution decrypted image through progressive upsampling (such as transposed convolution or interpolation). The strategy of this multi-level upsampling strategy is to transform the optimization result of the tensor neural network into a visually understandable image while preserving detailed information (such as edges and textures) from the optimization process.

[0128] The hierarchical features are progressively upsampled based on the multi-level upsampling strategy to reconstruct the corresponding high-resolution image.

[0129] In this embodiment, the hierarchical features can be progressively upsampled based on the strategy content of the multi-level upsampling strategy to reconstruct the corresponding high-resolution image.

[0130] The high-resolution image is used as the target decryption image corresponding to the encrypted image.

[0131] In this embodiment, the goal of the tensor neural network's dynamic optimization is to reduce the error function value, while the goal of decryption image reconstruction is to generate a high-quality decrypted image. The two are closely related through the error function and fusion features. Furthermore, the tensor neural network achieves adaptive adjustment of the optimization step size through dynamic error constraints, while the U-Net decoder achieves stable restoration of spatial resolution through multi-level upsampling. Together, they ensure a balance between dynamic optimization and stable reconstruction in the decryption process. In addition, the optimization results of the tensor neural network guide the reconstruction of the decrypted image, and the reconstruction results are fed back to the tensor neural network for the next round of optimization. This closed-loop iterative mechanism is key to the efficient operation of the decryption system. In summary, the tensor neural network's dynamic optimization provides the optimization direction and step size for decryption image reconstruction through error function and differential equation solving, while the decryption image reconstruction transforms the optimization results into a high-quality decrypted image through the U-Net decoder and multi-level upsampling. The two form a closed loop through iterative updates, jointly achieving efficient and accurate image decryption.

[0132] This application inputs the optimized feature information into the target decoder; then, the target decoder integrates the optimized feature information to obtain corresponding hierarchical features; subsequently, a preset multi-level upsampling strategy is obtained; and based on the multi-level upsampling strategy, the hierarchical features are progressively upsampled to reconstruct the corresponding high-resolution image; subsequently, the high-resolution image is used as the target decryption image corresponding to the encrypted image. This application obtains corresponding hierarchical features by integrating optimized feature information using the target decoder, and then progressively upsamples the hierarchical features based on the multi-level upsampling strategy to reconstruct the target decryption image corresponding to the encrypted image. This image decryption ensures that the decrypted image has both global consistency and local authenticity, guaranteeing the accuracy of the obtained target decryption image.

[0133] In some optional implementations of this embodiment, step S202 includes the following steps:

[0134] Call the preset preprocessing module.

[0135] In this embodiment, the preprocessing module is a pre-built functional module with image standardization processing capabilities.

[0136] The encrypted image is zero-padding processed by the preprocessing module to obtain the corresponding first encrypted image.

[0137] In this embodiment, the zero-padding process includes: zero-padding the input encrypted image to ensure that the image size meets the actual requirements of subsequent network processing.

[0138] The first encrypted image is divided into blocks to obtain the corresponding second encrypted image.

[0139] In this embodiment, the above-mentioned block processing includes: dividing the first encrypted image after zero padding into a series of non-overlapping blocks of fixed size (e.g., 16×16 pixel blocks) for subsequent processing.

[0140] The second encrypted image is used as the target encrypted image.

[0141] This application calls a preset preprocessing module; then, based on the preprocessing module, zero-padding is applied to the encrypted image to obtain a corresponding first encrypted image; subsequently, the first encrypted image is divided into blocks to obtain a corresponding second encrypted image; finally, the second encrypted image is used as the target encrypted image. This application, by calling a preset preprocessing module and then using the preprocessing module to perform zero-padding and block processing on the encrypted image, achieves efficient and accurate standardization of the encrypted image, ensuring the accuracy and standardization of the obtained target encrypted image.

[0142] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0143] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0144] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0145] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned target decryption image, the target decryption image can also be stored in a node of a blockchain.

[0146] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0147] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0148] Foundational technologies in artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0149] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0150] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0151] Further reference Figure 3As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of an image decryption apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0152] like Figure 3 As shown, the image decryption device 300 described in this embodiment includes: an acquisition module 301, a processing module 302, a first extraction module 303, a second extraction module 304, a fusion module 305, an optimization module 306, and a reconstruction module 307. Wherein:

[0153] Acquisition module 301 is used to acquire the encrypted image to be processed;

[0154] The processing module 302 is used to perform standardization processing on the encrypted image based on a preset preprocessing module to obtain the corresponding target encrypted image;

[0155] The first extraction module 303 is used to extract spatial features from the target encrypted image based on a preset spatial path processing strategy to obtain the corresponding spatial features.

[0156] The second extraction module 304 is used to extract temporal features from the target encrypted image based on a preset temporal path processing strategy to obtain the corresponding temporal features.

[0157] The fusion module 305 is used to dynamically fuse the spatial features and the temporal features based on a preset cross-attention gating module to obtain the corresponding fused features;

[0158] The optimization module 306 is used to optimize the fused features based on a preset tensor neural network to obtain corresponding optimized feature information;

[0159] The reconstruction module 307 is used to perform image reconstruction processing on the optimized feature information based on a preset target decoder to obtain a target decrypted image corresponding to the encrypted image.

[0160] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0161] In some optional implementations of this embodiment, the first extraction module 303 includes:

[0162] The first calling submodule is used to call the preset deep residual network;

[0163] An improvement submodule is used to improve the deep residual network based on a preset module introduction strategy to obtain the corresponding target deep residual network.

[0164] The first acquisition submodule is used to acquire a preset multi-scale feature extraction strategy;

[0165] The first extraction submodule is used to extract spatial features from the target encrypted image based on the target deep residual network and using the multi-scale feature extraction strategy to obtain the corresponding multi-scale spatial features.

[0166] The first determining submodule is used to use the multi-scale spatial features as the spatial features.

[0167] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0168] In some optional implementations of this embodiment, the second extraction module 304 includes:

[0169] The first processing submodule is used to perform block division and projection processing on the target encrypted image to obtain the corresponding first image block;

[0170] The encoding submodule is used to perform position encoding on the first image block based on a preset two-dimensional sinusoidal position encoding strategy to obtain the corresponding second image block;

[0171] The second calling submodule is used to call the preset target Transformer encoder;

[0172] The second extraction submodule is used to extract temporal features from the second image block based on the target Transformer encoder to obtain the corresponding specified features;

[0173] The second determining submodule is used to use the specified feature as the temporal feature.

[0174] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0175] In some optional implementations of this embodiment, the fusion module 305 includes:

[0176] The second processing submodule is used to recalibrate the spatial features based on the cross-attention gating module to obtain the corresponding specified spatial features.

[0177] An enhancement submodule is used to enhance the temporal features to obtain the corresponding specified temporal features;

[0178] The third calling submodule is used to call the preset fusion network;

[0179] The fusion submodule is used to dynamically fuse the specified spatial features and the specified temporal features based on the fusion network to obtain the corresponding fusion weight map.

[0180] The third determining submodule is used to use the fusion weight map as the fusion feature.

[0181] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0182] In some optional implementations of this embodiment, the optimization module 306 includes:

[0183] The second acquisition submodule is used to acquire a pre-constructed error function corresponding to the tensor neural network;

[0184] A generation submodule is used to generate corresponding dynamic error constraints based on the error function;

[0185] The third acquisition submodule is used to acquire the solution strategy corresponding to the preset continuous-time differential equation;

[0186] The optimization submodule is used to perform optimization processing on the fused features based on the tensor neural network, using the dynamic error constraints and the solution strategy, in accordance with the direction and step size, to obtain the corresponding optimization results.

[0187] The fourth determining submodule is used to use the optimization result as the optimization feature information.

[0188] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0189] In some optional implementations of this embodiment, the reconstruction module 307 includes:

[0190] An input submodule is used to input the optimized feature information into the target decoder;

[0191] The integration submodule is used to integrate the optimized feature information through the target decoder to obtain the corresponding hierarchical features;

[0192] The fourth acquisition submodule is used to acquire the preset multi-level upsampling strategy;

[0193] The upsampling submodule is used to progressively upsample the hierarchical features based on the multi-level upsampling strategy in order to reconstruct the corresponding high-resolution image;

[0194] The fifth determining submodule is used to use the high-resolution image as the target decryption image corresponding to the encrypted image.

[0195] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0196] In some optional implementations of this embodiment, the processing module 302 includes:

[0197] The fourth submodule is used to call the preset preprocessing module;

[0198] The third processing submodule is used to perform zero-padding processing on the encrypted image based on the preprocessing module to obtain the corresponding first encrypted image;

[0199] The fourth processing submodule is used to perform block processing on the first encrypted image to obtain the corresponding second encrypted image;

[0200] The sixth determining submodule is used to use the second encrypted image as the target encrypted image.

[0201] In this embodiment, the operations performed by the above modules or units correspond one-to-one with the steps of the image decryption method in the aforementioned embodiments, and will not be repeated here.

[0202] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0203] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0204] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0205] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for image decryption methods. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0206] In some embodiments, the processor 42 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, such as executing computer-readable instructions for the image decryption method.

[0207] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0208] Compared with the prior art, the embodiments of this application have the following beneficial effects:

[0209] In this embodiment, after obtaining the encrypted image to be processed, the encrypted image is standardized using a preprocessing module to obtain the target encrypted image. Then, spatial features are extracted from the target encrypted image using a spatial path processing strategy, and temporal features are extracted using a temporal path processing strategy. Next, the spatial and temporal features are dynamically fused using a cross-attention gating module to obtain fused features. Then, the fused features are optimized using a tensor neural network to obtain optimized feature information. Finally, the optimized feature information is used for image reconstruction using a target decoder, thereby efficiently and accurately obtaining the target decrypted image corresponding to the encrypted image. This application, through the combined use of a preprocessing module, spatial path processing strategy, temporal path processing strategy, cross-attention gating module, tensor neural network, and target decoder, requires only the encrypted image as input, without any encryption algorithm parameters or key information. It can achieve high-quality and efficient image decryption through spatiotemporal joint modeling and dynamic optimization without relying on prior knowledge of the encryption algorithm, effectively improving the decryption efficiency of images.

[0210] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the image decryption method described above.

[0211] Compared with the prior art, the embodiments of this application have the following main advantages:

[0212] In this embodiment, after obtaining the encrypted image to be processed, the encrypted image is standardized using a preprocessing module to obtain the target encrypted image. Then, spatial features are extracted from the target encrypted image using a spatial path processing strategy, and temporal features are extracted using a temporal path processing strategy. Next, the spatial and temporal features are dynamically fused using a cross-attention gating module to obtain fused features. Then, the fused features are optimized using a tensor neural network to obtain optimized feature information. Finally, the optimized feature information is used for image reconstruction using a target decoder, thereby efficiently and accurately obtaining the target decrypted image corresponding to the encrypted image. This application, through the combined use of a preprocessing module, spatial path processing strategy, temporal path processing strategy, cross-attention gating module, tensor neural network, and target decoder, requires only the encrypted image as input, without any encryption algorithm parameters or key information. It can achieve high-quality and efficient image decryption through spatiotemporal joint modeling and dynamic optimization without relying on prior knowledge of the encryption algorithm, effectively improving the decryption efficiency of images.

[0213] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0214] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. An image decryption method, characterized in that, Includes the following steps: Obtain the encrypted image to be processed; The encrypted image is standardized based on a preset preprocessing module to obtain the corresponding target encrypted image; Based on a preset spatial path processing strategy, spatial features are extracted from the target encrypted image to obtain the corresponding spatial features. Based on a preset temporal path processing strategy, temporal features are extracted from the target encrypted image to obtain the corresponding temporal features; Based on a preset cross-attention gating module, the spatial features and the temporal features are dynamically fused to obtain the corresponding fused features; The fused features are optimized based on a preset tensor neural network to obtain corresponding optimized feature information; Based on a preset target decoder, the optimized feature information is used to perform image reconstruction processing to obtain a target decrypted image corresponding to the encrypted image.

2. The image decryption method according to claim 1, characterized in that, The step of extracting spatial features from the target encrypted image based on a preset spatial path processing strategy to obtain the corresponding spatial features specifically includes: Invoke the preset deep residual network; The deep residual network is improved based on a preset module introduction strategy to obtain the corresponding target deep residual network. Obtain a pre-defined multi-scale feature extraction strategy; Based on the target deep residual network, the multi-scale feature extraction strategy is used to extract spatial features from the target encrypted image to obtain the corresponding multi-scale spatial features. The multi-scale spatial features are used as the spatial features.

3. The image decryption method according to claim 1, characterized in that, The step of extracting temporal features from the target encrypted image based on a preset temporal path processing strategy to obtain the corresponding temporal features specifically includes: The target encrypted image is divided into blocks and projected to obtain the corresponding first image block; The first image block is position-encoded based on a preset two-dimensional sinusoidal position encoding strategy to obtain the corresponding second image block; Invoke the preset target Transformer encoder; Based on the target Transformer encoder, temporal features are extracted from the second image block to obtain the corresponding specified features; The specified feature is used as the temporal feature.

4. The image decryption method according to claim 1, characterized in that, The step of dynamically fusing the spatial features and the temporal features based on a preset cross-attention gating module to obtain the corresponding fused features specifically includes: The spatial features are recalibrated based on the cross-attention gating module to obtain the corresponding specified spatial features. The time-series features are enhanced to obtain the corresponding specified time-series features; Invoke the preset fusion network; Based on the fusion network, the specified spatial features and the specified temporal features are dynamically fused to obtain the corresponding fusion weight map; The fusion weight map is used as the fusion feature.

5. The image decryption method according to claim 1, characterized in that, The step of optimizing the fused features based on a preset tensor neural network to obtain corresponding optimized feature information specifically includes: Obtain the pre-constructed error function corresponding to the tensor neural network; Generate corresponding dynamic error constraints based on the error function; Obtain the solution strategy corresponding to the preset continuous-time differential equation; Based on the tensor neural network, the dynamic error constraint and the solution strategy are used to optimize the fused features in accordance with the direction and step size to obtain the corresponding optimization results. The optimization result is used as the optimization feature information.

6. The image decryption method according to claim 1, characterized in that, The step of performing image reconstruction processing on the optimized feature information based on a preset target decoder to obtain a target decrypted image corresponding to the encrypted image specifically includes: The optimized feature information is input into the target decoder; The target decoder integrates the optimized feature information to obtain the corresponding hierarchical features. Obtain the preset multi-level upsampling strategy; The hierarchical features are progressively upsampled based on the multi-level upsampling strategy to reconstruct the corresponding high-resolution image. The high-resolution image is used as the target decryption image corresponding to the encrypted image.

7. The image decryption method according to claim 1, characterized in that, The step of standardizing the encrypted image based on a preset preprocessing module to obtain the corresponding target encrypted image specifically includes: Call the preset preprocessing module; The encrypted image is zero-padding processed by the preprocessing module to obtain the corresponding first encrypted image. The first encrypted image is divided into blocks to obtain the corresponding second encrypted image; The second encrypted image is used as the target encrypted image.

8. An image decryption device, characterized in that, include: The acquisition module is used to acquire the encrypted image to be processed; The processing module is used to perform standardization processing on the encrypted image based on a preset preprocessing module to obtain the corresponding target encrypted image; The first extraction module is used to extract spatial features from the target encrypted image based on a preset spatial path processing strategy to obtain the corresponding spatial features. The second extraction module is used to extract temporal features from the target encrypted image based on a preset temporal path processing strategy to obtain the corresponding temporal features; The fusion module is used to dynamically fuse the spatial features and the temporal features based on a preset cross-attention gating module to obtain the corresponding fused features; An optimization module is used to optimize the fused features based on a preset tensor neural network to obtain corresponding optimized feature information; The reconstruction module is used to perform image reconstruction processing on the optimized feature information based on a preset target decoder to obtain a target decrypted image corresponding to the encrypted image.

9. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the image decryption method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image decryption method as described in any one of claims 1 to 7.