Data processing method and device based on cross-domain feature alignment, equipment and medium
Features are extracted through block processing, linear projection, position coding and timing processing units, and combined with domain adaptation parameters and hierarchical attention weight allocation mechanism, the problem of insufficient alignment of multi-scale features in encrypted images is solved, and the high accuracy and robustness of decrypted images are achieved.
Patent Information
- Application Number
- CN202510651049.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, the multi-scale feature alignment is insufficient, resulting in the loss of local detail information of the encrypted image and the mismatch of the global structural consistency, affecting the decryption effect.
By obtaining the image data to be processed for block processing, a local block sequence is generated, and a potential spatial representation is generated through linear projection and position encoding transformation. The pre-trained timing processing unit is used to extract short-range local details, medium-range regional features and long-range global structural features, and aligned processing is performed in combination with domain adaptation parameters, and finally feature fusion is performed through the hierarchical attention weight allocation mechanism.
Improve local detail restoration and global structural consistency of decrypted images, solve the problems of local detail loss and global structural mismatch caused by insufficient multi-scale feature alignment, and enhance the accuracy and robustness of decrypted models.
Smart Images

Figure CN120543992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a data processing method, apparatus, device and storage medium based on cross-domain feature alignment. Background Art
[0002] In the field of image security, image encryption technology has been widely used in multiple areas, including privacy protection and secure communications. It plays an important role in scenarios where sensitive image data requires secure isolation and access control. However, existing image decryption technologies generally suffer from problems such as a single method and poor adaptability. Traditional decryption methods rely on knowing the specific encryption algorithm and key information. Although accurate recovery can be achieved in theory, they are difficult to apply if the key encryption parameters cannot be obtained. While adaptive decryption methods based on deep learning can attempt decryption in the absence of encryption details, they still have significant shortcomings in handling complex encryption structures and capturing dynamic characteristics.
[0003] In the fintech sector, with the rise of applications such as remote transactions and financial data storage, encrypted images are widely used in scenarios such as identity verification, contract tracking, and transaction vouchers. Current demands for decrypting encrypted images are highly dynamic and diverse, but traditional neural networks have significant limitations in capturing the temporal characteristics of financial image data. In particular, during dynamic encryption processing, the lack of effective temporal dependency modeling often leads to localized distortion or loss of structural information in the decrypted image, impacting authenticity verification and compliance traceability for financial transactions.
[0004] In the healthcare sector, the need to protect the privacy of medical images has driven the widespread use of encryption technology for the storage and transmission of sensitive data such as medical records and diagnostic reports. Existing decryption methods face difficulties in recovering structural details when processing encrypted medical images. Traditional convolutional neural network models primarily focus on static spatial features and struggle to accurately recover the temporal dynamic features introduced by encryption perturbations. This results in insufficient recovery of subtle lesions and fine-grained tissue structures in decrypted images, hindering subsequent clinical analysis and auxiliary diagnosis.
[0005] Existing methods have made some progress in processing cross-domain data (e.g., recommendation systems and behavior recognition) in general domains. However, encrypted and non-encrypted images exhibit more complex structural differences. Existing cross-domain transfer methods often lack domain adaptation for encryption scenarios. This results in an inability to effectively bridge the gap between encrypted and decrypted data distributions during transfer learning, thereby reducing the generalization ability of decryption models across different data domains.
[0006] Furthermore, currently used attention mechanisms are mostly single-scale or static, making it difficult to simultaneously account for both local perturbations and global structural features. This is especially true when encrypted images experience significant dynamic changes, as traditional attention mechanisms cannot flexibly adjust the range of feature attention based on the time scale. This single-attention strategy exhibits significant deficiencies in fine-grained tissue reconstruction of medical images and tamper-proof verification of financial document images. Furthermore, existing tensor neural network methods often employ fixed convergence strategies and lack the ability to adjust attention weights in real time based on dynamic feature changes, further limiting the adaptability and robustness of decryption tasks. Summary of the Invention
[0007] The main purpose of the present invention is to provide a data processing method, device, equipment and storage medium based on cross-domain feature alignment, aiming to solve the technical problem in the prior art that insufficient multi-scale feature alignment leads to loss of local detail information and mismatch of global structural consistency, affecting the accurate reconstruction and decryption effect of encrypted images.
[0008] To achieve the above objectives, the present invention provides a data processing method based on cross-domain feature alignment, comprising:
[0009] Acquiring image data to be processed, and performing block processing on the image data to be processed to generate a local block sequence;
[0010] Transforming the local block sequence through linear projection and position encoding to generate a latent space representation;
[0011] Inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structure features;
[0012] Aligning the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain, respectively, through the domain adaptation parameters in the temporal processing unit;
[0013] Through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fusion features;
[0014] Target image data is generated based on the multi-scale fusion features.
[0015] Furthermore, to achieve the above-mentioned object, the present invention provides a data processing device based on cross-domain feature alignment, comprising:
[0016] A block preprocessing module is used to obtain image data to be processed, and perform block processing on the image data to be processed to generate a local block sequence;
[0017] A feature encoding module, configured to transform the local block sequence through linear projection and position encoding to generate a latent space representation;
[0018] a temporal feature extraction module, configured to input the latent space representation into a pre-trained temporal processing unit, extract target domain features from the latent space representation, and decompose the target domain features into short-range local detail features, medium-range regional features, and long-range global structural features;
[0019] a feature alignment module, configured to align the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain using the domain adaptation parameters in the temporal processing unit;
[0020] The feature fusion module is used to fuse the aligned short-range local detail features, medium-range regional features, and long-range global structural features through a hierarchical attention weight distribution mechanism to generate multi-scale fused features;
[0021] An image generation module is used to generate target image data based on the multi-scale fusion features.
[0022] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a data processing program based on cross-domain feature alignment stored in the memory and run on the processor. When the data processing program based on cross-domain feature alignment is executed by the processor, the steps of the data processing method based on cross-domain feature alignment as described above are implemented.
[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a data processing program based on cross-domain feature alignment is stored. When the data processing program based on cross-domain feature alignment is executed by a processor, the steps of the data processing method based on cross-domain feature alignment as described above are implemented.
[0024] Beneficial effects: The present invention relates to the field of image processing technology and can be applied to business scenarios such as financial technology and medical health. A data processing method, device, equipment and medium based on cross-domain feature alignment are disclosed, including: obtaining image data to be processed and performing block processing to generate a local block sequence; converting the local block sequence through linear projection and position encoding to generate a latent space representation; inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features and decomposing them into short-range local detail features, medium-range regional features and long-range global structure features; through the domain adaptation parameters in the temporal processing unit, aligning the short-range local detail features, medium-range regional features and long-range global structure features with the corresponding scale feature distribution of the pre-trained source domain; through the hierarchical attention weight allocation mechanism, fusing the aligned short-range local detail features, medium-range regional features and long-range global structure features to generate multi-scale fusion features; generating target image data based on the multi-scale fusion features. The present invention combines local block segmentation processing with linear projection modeling of latent space features, uses a temporal processing unit to extract and decompose features of different scales, and aligns the distribution of short-range, medium-range and long-range scale features based on domain adaptation parameters. It adaptively allocates feature weights and fuses multi-scale information through a hierarchical attention mechanism to finally generate target image data, thereby effectively improving the accuracy of decrypted images in local detail restoration and global structural consistency, and solving the problems of local detail loss and global structure mismatch caused by insufficient multi-scale feature alignment in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0026] Figure 1 A schematic diagram of an application environment of a data processing method based on cross-domain feature alignment in an embodiment of the present invention;
[0027] Figure 2 This is a flow chart of an embodiment of a data processing method based on cross-domain feature alignment according to the present invention;
[0028] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a data processing device based on cross-domain feature alignment of the present invention;
[0029] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0030] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0032] The data processing method based on cross-domain feature alignment provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can obtain the image data to be processed through the user end and perform block processing to generate a local block sequence; the local block sequence is transformed through linear projection and position encoding to generate a latent space representation; the latent space representation is input into a pre-trained temporal processing unit to extract the target domain features and decompose them into short-range local detail features, medium-range regional features and long-range global structure features; through the domain adaptation parameters in the temporal processing unit, the short-range local detail features, medium-range regional features and long-range global structure features are aligned with the corresponding scale feature distribution of the pre-trained source domain; through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fused features; and the target image data is generated based on the multi-scale fused features. The present invention combines local block segmentation processing with linear projection modeling of latent space features, utilizes a temporal processing unit to extract and decompose features of different scales, and aligns the distribution of short-range, medium-range, and long-range scale features based on domain adaptation parameters. It adaptively allocates feature weights and fuses multi-scale information through a hierarchical attention mechanism, and finally generates target image data, thereby effectively improving the accuracy of decrypted images in terms of local detail restoration and global structural consistency, and solving the problems of local detail loss and global structural mismatch caused by insufficient multi-scale feature alignment in the prior art. Among them, the user end can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server end can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0033] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a data processing method based on cross-domain feature alignment provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0034] like Figure 2 As shown, the data processing method based on cross-domain feature alignment proposed in the present invention includes the following steps:
[0035] S10, acquiring image data to be processed, and performing block processing on the image data to be processed to generate a local block sequence;
[0036] In this embodiment, the purpose of obtaining the image data to be processed is to provide raw data support for subsequent feature encoding and feature extraction. The image data to be processed refers to data entities from any source, including medical images, financial bill scans, surveillance images, satellite images, or other visual image files with data steganography or encryption. The acquisition method can be based on direct collection of real-time acquisition equipment, such as DICOM format images generated by medical imaging equipment, high-definition bill images generated by scanning financial terminals, or retrieving preset image data sets through server storage systems, or receiving data from remote encrypted storage through asynchronous data synchronization interfaces.
[0037] The purpose of performing block processing on the image data to be processed is to divide the overall large-scale image into multiple small-scale sub-regions, so as to improve the local feature expression ability and computational efficiency in subsequent processing. Block processing is usually completed based on a two-dimensional sliding window operation. The size of the sliding window is determined by the application scenario. In the field of medical imaging, a size that can cover fine-grained organizational structures can be selected, while in the field of financial imaging, a size that can completely cover the bill information block can be selected. The sliding window can be set with a fixed step size in space. By controlling the step size, the degree of overlap between blocks can be achieved, thereby improving the expression ability of local continuity features. Edge filling operations can be performed during the block process to avoid the problem of insufficient window size at the boundary. The filling method can be mirror filling, zero filling or reflection filling, which can be dynamically adjusted according to different application requirements.
[0038] Generating a local block sequence means organizing each sub-image block into a vector set in order after the block processing to form a local block sequence for subsequent feature encoding and modeling. When constructing the local block sequence, the position index information of each block can be retained to form a correspondence between the block and the spatial position information of the original image. This position mark helps to ensure the accuracy of the subsequent position encoding process. When organizing the local block sequence, a one-dimensional arrangement can be formed according to the order in which the blocks are generated (for example, scanning from left to right, from top to bottom), or a secondary grouping can be constructed according to the application scenario. For example, in the medical and health field, local blocks can be classified according to tissue and organ regions, and in the financial field, they can be classified according to the format structure of bills. The local block sequence not only contains the pixel data of the sub-image block, but can also be accompanied by auxiliary attributes such as boundary markers, timestamps, and source identifiers to adapt to complex data fusion requirements.
[0039] In practical applications, there are various ways to obtain image data for processing. For example, in healthcare scenarios, a direct connection to a PACS system can be used to extract DICOM-compliant medical image data from the hospital's internal server, and different sampling parameters can be selected based on the image type. In financial scenarios, high-resolution scanners deployed at counters or on mobile terminals can capture images of customer-provided receipts in real time and pre-process the captured images for noise removal and enhancement.
[0040] Block processing can use a fixed window scanning method, setting the window size to the pixel size that matches the target application, and dynamically adjusting the step size based on the desired overlap ratio between blocks. During boundary processing, if the target image size is not divisible by the sliding window, a mirror edge extension method can be used to symmetrically expand the image edge pixels to the required size of the full block, ensuring the integrity of the information within each block.
[0041] Local block sequences can be generated using a row-major ordering method, where block elements are arranged sequentially according to row scanning order. Spatial index encoding can also be inserted as needed to attach absolute position information or a local coordinate system identifier to each block. When processing heterogeneous data, local block sequence generation can be accompanied by intra-block normalization, such as performing mean or standard deviation normalization on each block, to reduce cross-block data distribution differences and improve the consistency of subsequent feature encoding.
[0042] When faced with different production environments, technology can be adapted based on the data scale and processing capabilities. In high-throughput environments, batch processing can be used to process multiple image data files at once. In edge computing environments, streaming block processing and local real-time processing can be used to adapt to data processing needs under resource-constrained conditions.
[0043] Example: In the healthcare business, pathology slice images can be processed in blocks, with each block covering a limited range of cell tissue. By constructing a local block sequence, the neural network can accurately capture the boundary features of tiny lesions and internal tissue heterogeneity, improving the subsequent image decryption and reconstruction accuracy.
[0044] In the financial technology business field, electronic bill images can be processed in blocks. Each local block corresponds to a specific field or security watermark area on the bill. Through local block sequence management and organization, it helps to reconstruct the original format and anti-counterfeiting features of the bill during decryption, thereby improving the reliability of bill verification and archiving management.
[0045] This embodiment performs standardized block processing on the image data to be processed to generate a structured local block sequence, which not only improves the data organization efficiency in the subsequent feature encoding process, but also enhances the expression ability of local fine-grained features in the model training stage, improves the comprehensive restoration level of local disturbances and global structural information during decryption modeling, and improves the problem of detail loss caused by overall feature abstraction in traditional processing methods.
[0046] S20, transforming the local block sequence through linear projection and position encoding to generate a latent space representation;
[0047] In this embodiment, the purpose of transforming the local block sequence through linear projection and position encoding is to map the original local block sequence from a high-dimensional pixel space to a unified latent space representation, thereby facilitating subsequent feature extraction and domain alignment. A local block sequence is a set of image blocks with a clear spatial order obtained through normalization. It is usually organized in the form of a vector or tensor, and each block retains its own pixel information and relative position information.
[0048] Linear projection is an operation that maps input features to a new feature space through matrix transformation. Its core is to set a learnable weight matrix and multiply it by the input vector to obtain a compressed or reconstructed feature representation. In this scenario, the linear projection operation transforms each block element in the local block sequence through a low-rank matrix, converting the original high-dimensional block vector into a fixed-dimensional feature vector, thereby reducing the complexity of the feature space while ensuring the basic amount of information. Linear projection can use fully connected layer parameters initialized with decimal weights from a standard normal distribution, or it can use a specific distribution to initialize the weight matrix based on the distribution of pre-training data to improve initial learning stability.
[0049] Positional encoding is the process of adding explicit positional identifiers to each block element in a local block sequence. Its primary purpose is to preserve the spatial position information of image blocks in the latent space, preventing the loss of spatial structure caused by pure feature transformations. Positional encoding can be generated in a variety of ways, such as periodic encoding of sine and cosine functions, or dynamically generated through a learnable position embedding matrix. Sine-cosine encoding, by superimposing different frequencies, can capture both local fine-grained and global positional information, offering the advantages of being parameter-free and generalizable. Learnable positional encoding, on the other hand, can automatically adjust spatial sensitivity based on specific tasks, improving the ability to adapt to local permutations during decryption.
[0050] The fusion of linear projection and position encoding is typically performed using element-wise addition. Each linearly transformed feature vector is summed with its corresponding position information vector along the same dimension to form the final latent space representation. This latent space representation has a uniform feature dimension and embeds position information, offering strong feature expression and spatial structure preservation capabilities, enabling subsequent feature extraction modules to perform fine-grained, multi-scale processing.
[0051] In practical applications, transforming a local block sequence through linear projection and positional encoding can be achieved through a series of sequential operations. First, the local block sequence is flattened into a set of one-dimensional vectors. For each block, the pixel values are arranged in row-major or column-major order to form a fixed-length block vector. Next, a set of linear transformation matrices with learnable parameters is designed. The number of columns in the matrix is equal to the feature dimension of the target latent space, and the number of rows is equal to the original length of the block vector. When initializing the weights, either the Xavier uniform distribution or the He normal distribution strategy can be used to ensure gradient stability during the initial stages of network training.
[0052] Perform matrix multiplication on each block vector by the linear transformation matrix to obtain the corresponding compressed feature vector. While generating the compressed feature vector, the spatial position of each block is determined. This position information can be generated using the block's upper left corner coordinates, center coordinates, or block index number in the original image. This position information can be generated using a position encoding module. If sine-cosine position encoding is selected, sine and cosine values of different frequencies are generated based on the block index and concatenated. If learnable position embedding is selected, a separate learnable vector is assigned to each position index.
[0053] After the positional encoding is generated, the linearly transformed block feature vector is added to the corresponding positional encoding vector in the element-wise dimension to form a fused feature. The fused feature can be normalized using a normalization module to ensure that the feature has a uniform scale and stable value range during subsequent processing, avoiding gradient explosion or vanishing issues.
[0054] The structures of the linear projection and position encoding modules can be flexibly adjusted in different application scenarios. For example, in resource-constrained environments, low-rank matrix decomposition can be used to approximate the linear projection operation, reducing computational complexity. In large-scale image processing tasks, a sparse position encoding strategy can be adopted to reduce memory usage. In scenarios where enhanced perception of spatial detail is required, multi-frequency position encoding or local adaptive position encoding can be introduced to improve sensitivity to subtle structural changes.
[0055] Example: In the healthcare business, for medical images at different levels, such as CT or MRI slice images, the high-dimensional pixel information of each local tissue block can be compressed into a unified latent feature vector through linear projection, while the spatial position information is superimposed. This enables subsequent models to distinguish fine structural changes in adjacent organs or lesion areas during the decryption and reconstruction process.
[0056] In the financial technology business field, for multi-format bill images, linear projection and position coding processing can be used to convert local fields of the bill, such as amount, date, and signature area, into a unified latent space representation. This ensures that even if encryption interference exists, the bill format details and field arrangement logic can still be accurately restored during the decryption stage, thereby improving the accuracy and stability of bill identification and verification.
[0057] This embodiment performs linear projection and position coding fusion processing on local block sequences, which can fully retain the local detail information and global position information in the image data while reducing the feature dimension and computational complexity, thereby achieving the unification of feature expression ability and spatial perception ability, providing a high-quality latent space foundation for subsequent cross-domain feature extraction and multi-scale dynamic alignment, and avoiding the information loss and spatial structure loss problems caused by traditional high-dimensional feature redundancy.
[0058] S30, inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structure features;
[0059] In this embodiment, the latent space representation is fed into a pre-trained temporal processing unit to further mine hidden dynamic temporal relationships and cross-block dependencies within a unified feature space. The latent space representation, derived from the previous stage through linear projection and positional encoding of the local block sequence, offers a unified dimensional feature representation and positional awareness, making it suitable as input for the temporal modeling module.
[0060] A pre-trained TPU is a neural network module that has been trained and optimized offline on a large number of related tasks or pre-trained data. Its structure typically includes a local attention window mechanism, a temporal convolutional unit, a layer normalization module, and a cross-domain alignment submodule. Pre-training enables the TPU to have strong cross-domain generalization and temporal feature modeling capabilities, improving feature extraction during transfer learning between encrypted and unencrypted images.
[0061] Extracting target domain features from the latent space representation involves extracting local-to-global information from input features within the temporal processing unit through local windowing or multi-scale convolution operations. The local window mechanism allows for dynamic modeling of neighborhood features around each local block, capturing short-range local detail variations; while multi-scale convolution or attention mechanisms help exploit structural connections between mid-range regions and long-range global regions across different scales.
[0062] Decomposing target domain features into short-range local detail features, medium-range regional features, and long-range global structural features is achieved by performing branch extraction operations on feature maps with different temporal and spatial receptive fields. Short-range local detail features focus on small local perturbations and high-frequency details in the image, such as small encryption perturbations or local noise changes; medium-range regional features focus on pattern and structural changes within the region, and are suitable for describing medium-scale phenomena such as continuous text regions or organ boundaries; long-range global structural features model the spatial layout and sequence evolution trends within the entire image, reflecting the macroscopic encryption logic or layout structure.
[0063] In practical applications, the latent space representation can be input into a pre-trained temporal processing unit (TSPU) through a tensor input interface. The latent space representation is first organized into batches of three- or four-dimensional tensors and input into the TPU. The TPU can employ a Transformer encoder architecture with local attention windows or a lightweight convolutional network with dilated convolutions and a dynamic receptive field adjustment mechanism. During feature extraction, local windows of varying sizes can be introduced for different levels of features. Short-range windows can be set to cover adjacent local blocks, medium-range windows cover larger groups of blocks, and long-range windows use global attention or global pooling strategies to extract holistic features. The feature extraction module outputs feature branches at different scales separately. Each branch undergoes independent normalization and dimensionality reduction, ensuring that short-range local detail features, medium-range regional features, and long-range global structural features maintain a unified format, facilitating subsequent fusion or alignment. For different scenarios, the decomposition method can be dynamically adjusted based on the characteristics of the image data. For example, in healthcare, the extraction of short-range local detail features can be enhanced to address blurred organ boundaries or subtle changes in lesion areas. In the financial instrument sector, the extraction of long-range global structural features can be appropriately strengthened to address the importance of global layout. To adapt to heterogeneous data environments, a configurable time window size or multi-receptive field selection mechanism can be introduced, allowing the time series processing unit to adaptively switch feature extraction modes based on the characteristics of the input data.
[0064] Example description: In the medical and health business field, the latent space representation can be input into the timing processing unit for the MRI sequence images obtained by continuous scanning. The changes in tiny lesions on the surface of organs can be extracted through short-range local detail features, the changes in relative displacement between organs can be extracted through medium-range regional features, and the overall position change trend of organs in the entire scanning sequence can be extracted through long-range global structural features, thereby effectively identifying the true pathological characteristics masked by encrypted perturbations.
[0065] In the field of financial technology business, the latent space representation of encrypted and stored electronic contract image data can be input into the time series processing unit, and the document detail watermark area can be extracted through short-range local detail features, the text block layout structure can be extracted through medium-range regional features, and the overall alignment relationship of the layout can be extracted through long-range global structure features. This helps to achieve accurate restoration and authenticity verification of the contract layout during the decryption and restoration process.
[0066] This embodiment, by inputting the latent space representation into a pre-trained temporal processing unit and extracting target domain features of different scales, can accurately model the local perturbations of the image, medium-scale structural changes, and the overall spatial layout evolution process in a unified latent feature space, thereby achieving coordinated perception of local details and global structures in the encrypted image, compensating for the detail loss and structural mismatch problems caused by the lack of temporal dynamic modeling in traditional static feature extraction methods, and providing multi-granularity, multi-level rich feature support for subsequent cross-domain alignment and feature fusion.
[0067] S40, aligning the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain using the domain adaptation parameters in the temporal processing unit;
[0068] In this embodiment, the domain adaptation parameters in the temporal processing unit are used to align short-range local detail features, medium-range regional features, and long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain. This aims to reduce the statistical distribution differences between encrypted and unencrypted data in different scale feature spaces and enhance the model's cross-domain migration capabilities during decryption inference. The temporal processing unit is an integrated module that performs feature extraction, domain adaptation, and parameter updates, and can execute adversarial training strategies based on the input target domain features.
[0069] Domain adaptation parameters refer to the set of weights learned during the training phase based on adversarial optimization between unencrypted source domain data and encrypted target domain data. These parameters include short-range, medium-range, and long-range domain adaptation parameters, each corresponding to a discriminator subnetwork at a specific scale. These domain adaptation parameters control how the discriminator assesses the consistency of input features with the source domain feature distribution. Specifically, they can take the form of convolution kernel weights, normalization parameters, or fully connected mapping matrices.
[0070] Short-range local detail features correspond to fine-grained features generated by small-scale local perception units. Their alignment emphasizes restoring the high-frequency structural distribution of source domain samples, such as edges and textures, within a localized area. Medium-range regional features focus on continuous pattern changes within a specific area. Alignment requires maintaining overall regional consistency to avoid distortion of the medium-scale structure introduced by densification. Long-range global structural features describe spatial relationships and temporal patterns across the entire image. During alignment, the overall layout logic must be consistent with the source domain.
[0071] Alignment involves calculating the discriminant between target and source domain features and then backpropagating gradient information to adjust the target domain feature distribution, making it impossible to effectively distinguish between source and target domain data within the discriminator output space. This process typically combines random interpolation sample generation, gradient penalty calculation, and weight update operations to form a complete adversarial domain adaptation optimization chain.
[0072] In practical implementation, independent discriminator subnetworks can be constructed during training for short-range local detail features, medium-range regional features, and long-range global structural features, with each discriminator using independent domain adaptation parameters. The target features at each scale are input to the corresponding discriminator, which outputs a discriminant score, representing the probability that the feature belongs to the source or target domain. To improve alignment, each pair of source and target domain features can be randomly interpolated within the feature space to generate cross-domain interpolated samples. By increasing the number of interpolated samples in the training process, feature distribution boundaries are smoothed and the gradient explosion problem is alleviated. Uniform random sampling can be used for the interpolation coefficients to cover more distribution transition regions. For gradient penalty, the gradient magnitude of the output with respect to the input can be calculated based on the discriminant output of the interpolated samples. By introducing a gradient penalty loss, the discriminator is forced to maintain Lipschitz continuity, thereby avoiding overfitting to local perturbations. During the parameter update process, independent optimization steps can be performed for the discriminators at the short, medium, and long scales, and their domain adaptation parameters are updated independently based on their respective adversarial losses and backpropagation. To adapt to different data types, the depth or receptive field size of the discriminator network can be adjusted based on the difference in feature scale between the source and target domains. For example, for super-resolution encrypted samples in medical imaging data, the receptive field of the short-range discriminator can be increased; for layout-encrypted samples in financial documents, the global modeling capabilities of the long-range discriminator can be strengthened.
[0073] Example description: In the medical and health business field, for encrypted continuous ultrasound image sequences, the continuity of lesion edges can be accurately restored by aligning short-range local detail features, the consistency of internal organ structure patterns can be maintained by aligning mid-range regional features, and the temporal consistency of the entire scanning path can be restored by aligning long-range global structural features, thereby assisting in the clinical usability decryption of encrypted ultrasound sequences.
[0074] In the financial technology business field, for encrypted and stored electronic bills or contract documents, short-range local detail feature alignment can be used to accurately restore tiny details such as watermarks and signatures, mid-range regional feature alignment can be used to maintain the consistency of the text typesetting area, and long-range global structural feature alignment can be used to repair the entire document layout structure, thereby achieving high-fidelity decryption and structured information extraction of encrypted bill or contract images.
[0075] This embodiment uses an independent domain adaptation parameter mechanism based on scale decomposition, combined with cross-domain random interpolation and gradient penalty strategies, to align short-range local detail features, medium-range regional features, and long-range global structural features with the source domain feature distribution, effectively making up for the shortcomings of traditional overall alignment methods in local detail preservation and overall structure restoration, improving the accuracy of local texture restoration and the rationality of global layout during decryption reasoning, significantly enhancing the model's cross-domain migration capability and inference robustness, and providing a unified and fine-grained aligned multi-scale feature input for subsequent feature fusion and image reconstruction.
[0076] S50, through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fused features;
[0077] In this embodiment, the aligned short-range local detail features, medium-range regional features and long-range global structural features are fused through a hierarchical attention weight allocation mechanism to generate multi-scale fusion features. The purpose is to dynamically adjust the contribution ratio of features of different scales in the final feature representation based on multi-granularity perception, thereby integrating local details and global layout information to form a fusion representation with unified spatial semantics and retaining the integrity of details.
[0078] The hierarchical attention weight allocation mechanism introduces independent attention weight generation modules for features at different scales. This mechanism weights short-range local detail features, mid-range regional features, and long-range global structural features, and then weights the features based on these weights before fusion. The hierarchical nature of attention weights is that attention weights are not uniformly allocated to all features at once, but are calculated and normalized separately at different scales and levels, preserving the independence and complementarity of features at different scales.
[0079] Short-range local detail features typically focus on high-frequency textures and small-scale structural changes. Therefore, their attention weights should emphasize the importance of local fine structure in the fusion process. Medium-range regional features focus on identifying coherent patterns at the regional level, requiring a moderate tilt in attention allocation to prevent regional information from being weakened. Long-range global structural features model the overall layout and temporal evolution logic. Attention allocation must ensure that overall structural cues are not overwhelmed by local noise.
[0080] During the fusion process, attention weight matrices are first applied to each of the short-range local detail features, medium-range regional features, and long-range global structural features, each undergoing a linear transformation to obtain attention weights. Then, through a normalization operation, such as the Softmax function, the weights of features at different scales are jointly normalized, ensuring that the fused features maintain a stable numerical range while maintaining the expressive power of each feature.
[0081] After weighting, the short-range local detail features, mid-range regional features, and long-range global structural features are multiplied point-by-point by their corresponding normalized weights, channel by channel, to generate a weighted feature map. Finally, channel concatenation is used to concatenate the weighted short-range, mid-range, and long-range features into a unified feature tensor. A linear compression module can be introduced as needed to reduce the dimensionality of the concatenated tensor, resulting in the final multi-scale fused feature.
[0082] In practical applications, independent small attention weight prediction modules can be designed for short-range local detail features, medium-range regional features, and long-range global structural features. Each module inputs features at the corresponding scale and outputs an attention weight matrix through one or more layers of small convolutional or fully connected neural networks. Layer normalization can be used to process the weight tensors within each scale to improve numerical stability. During the weight normalization stage, normalization can be applied dynamically to each sample to ensure that the sum of weights at different scales is equal to one, ensuring that the feature fusion process is not biased by amplifying the value of features at a single scale. Normalization methods can include softmax normalization, L2 normalization, or more complex normalization mechanisms based on sparsity regularization to adapt to different data distribution characteristics. During the weighting operation, the channel dimension remains consistent, and each scale feature is multiplied by its corresponding normalized weight through channel-by-channel element-by-element multiplication. After weighting, the weighted short-range, medium-range, and long-range features are concatenated along the channel dimension. To control the dimensionality of the final fused features and prevent feature redundancy, a one-dimensional convolution or fully connected mapping layer can be added after concatenation for linear dimensionality reduction. For different application scenarios, the structure of the attention weight module can be adjusted based on the complexity of the input image. For example, in medical and health images, to better capture fine-grained lesion signals, the resolution of attention weight prediction for short-range local detail features can be improved. In financial document image processing, to emphasize overall layout consistency, the receptive field size of the weight prediction module for long-range global structural features can be enhanced.
[0083] This embodiment uses a hierarchical attention weight distribution mechanism to achieve dynamic and fine-grained feature importance adjustment in the process of fusing short-range local detail features, medium-range regional features and long-range global structural features. It not only improves the fidelity of local details, but also ensures the coherence of medium-range patterns and the consistency of global layout, effectively avoiding the problem of detail loss or structural mismatch caused by the dominance of single-scale features, thereby constructing a unified, multi-scale, and highly semantically consistent fusion feature representation, providing a more solid and complete feature foundation for subsequent image reconstruction.
[0084] S60: Generate target image data based on the multi-scale fusion features.
[0085] In this embodiment, generating target image data based on multi-scale fusion features means that a unified feature representation that has completed the fusion of short-range, medium-range and long-range scale features is finally restored to visual high-quality image data through a series of operations such as upsampling, skip connection, feature splicing, convolutional decoding and normalization.
[0086] Multi-scale fusion features are a unified feature tensor formed by dynamically weighting the aligned short-range local detail features, mid-range regional features, and long-range global structural features. This unified feature tensor incorporates rich local detail information while maintaining mid-range structural coherence and global layout consistency. Image reconstruction based on multi-scale fusion features effectively compensates for the local information loss and global structural distortion that occurs in traditional single-scale feature reconstruction.
[0087] Layer-by-layer upsampling involves gradually restoring multi-scale fused features to the target image resolution through a succession of stacked upsampling units. Upsampling can be implemented using methods such as transposed convolution, nearest neighbor interpolation plus convolution, and sub-pixel convolution. Each upsampling step expands the feature map size and refines the feature space's expressiveness.
[0088] Skip connections involve passing the local block sequence generated in the encoding phase to the corresponding upsampling layer in the decoding phase, where it is fused with the current upsampled features to supplement the fine-grained spatial information lost in the decoder due to downsampling, linear transformation, and other operations. The local block sequence provides low-level detail constraints in the form of standardized block features, helping to maintain the coherence of image textures and boundaries during decoding.
[0089] The initial upsampled feature map is concatenated with the sequence of local blocks passed by the skip connection in the channel dimension to generate an enhanced feature map. The concatenation operation directly stacks two tensors in the channel direction, allowing subsequent convolution operations to simultaneously utilize information from different stages, achieving multi-source feature fusion.
[0090] Residual convolution involves applying convolutional units, activation units, and a residual connection mechanism to an enhanced feature map to extract higher-order, more stable feature representations. Convolutional units capture local patterns through local receptive fields, activation units introduce nonlinear representations, and residual connections facilitate gradient flow, preventing degradation in deep networks.
[0091] Finally, pixel-by-pixel normalization is performed on the reconstructed feature map, independently mapping the feature value of each pixel to a standard numerical range. This outputs a standardized target image data that meets visualization requirements and is numerically stable. Pixel-by-pixel normalization can be implemented using a sigmoid function or other smooth activation functions to ensure that the brightness and contrast of the generated image meet the expected standards.
[0092] In specific implementations, a multi-layer transposed convolutional module can be used to form an upsampling path. After each upsampling, the upsampled feature map is concatenated with the local block sequence features at the corresponding scale along the channel dimension. To ensure consistent feature dimensionality after concatenation, a one-dimensional convolution or linear transformation can be used before the skip connection input to adjust the number of channels of the local block sequence features to match that of the upsampled feature map. The upsampling unit can be flexibly configured based on the target output resolution. For example, if the target image is twice the original input resolution, a transposed convolution kernel can be designed with a doubling of the size for each upsampling. If the target image is four times the input, two layers of upsampling modules can be stacked, each doubling the size. During convolutional decoding, multiple small convolutional kernels can be stacked instead of a single large kernel to enhance the network's nonlinear representation and reduce the number of parameters. The activation function can be ReLU, LeakyReLU, or GELU, depending on the smoothness requirements of the feature distribution in different scenarios. Residual connections add a direct connection between every two convolutional modules, forming a short-circuit path and alleviating the vanishing gradient problem. During normalization, a per-pixel Sigmoid activation can be applied to the output image to ensure that each pixel value lies between zero and one. If the application requires integer image output, the Sigmoid output can be multiplied by a fixed scale and rounded. To adapt to different data scenarios, such as differences in input image size and number of image channels, the upsampling module and number of convolution kernels can be dynamically adjusted during the decoding process. Grayscale images in the healthcare field can use a single-channel output mode, while financial documents or contract images require support for three-channel color output.
[0093] Example: In the healthcare business, for encrypted MRI data, target image data is generated based on multi-scale fusion features, which can restore the organ microstructure details layer by layer while maintaining the consistency of the overall spatial layout of the organ. This allows the decrypted image to not only accurately restore local tissue details, but also support clinical diagnosis needs at the overall anatomical structure level.
[0094] In the field of financial technology, when faced with encrypted scans of loan contracts, generating target image data based on multi-scale fusion features can meticulously restore tiny details such as text boundaries and watermark patterns, while ensuring the correct restoration of the entire contract layout, thereby improving the accuracy and reliability of subsequent automatic contract review, clause extraction, and compliance audits.
[0095] This embodiment generates target image data based on multi-scale fusion features, which can restore the structural coherence of the mid-range area and the spatial consistency of the global layout while preserving local texture details, avoiding the problems of detail blurring and structural deformation that are prone to occur in traditional single-path decoding methods, greatly improving the visual quality and semantic usability of the decrypted image, and providing a high-quality input foundation for subsequent downstream tasks such as medical image analysis, bill recognition, contract review and other applications.
[0096] The present invention relates to the field of image processing technology and can be applied to business scenarios such as financial technology and medical health. A data processing method, device, equipment and medium based on cross-domain feature alignment are disclosed, including: obtaining image data to be processed and performing block processing to generate a local block sequence; converting the local block sequence through linear projection and position encoding to generate a latent space representation; inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features and decomposing them into short-range local detail features, medium-range regional features and long-range global structure features; aligning the short-range local detail features, medium-range regional features and long-range global structure features with the corresponding scale feature distribution of the pre-trained source domain through domain adaptation parameters in the temporal processing unit; fusing the aligned short-range local detail features, medium-range regional features and long-range global structure features through a hierarchical attention weight allocation mechanism to generate multi-scale fused features; and generating target image data based on the multi-scale fused features. The present invention combines local block segmentation processing with linear projection modeling of latent space features, uses a temporal processing unit to extract and decompose features of different scales, and aligns the distribution of short-range, medium-range and long-range scale features based on domain adaptation parameters. It adaptively allocates feature weights and fuses multi-scale information through a hierarchical attention mechanism to finally generate target image data, thereby effectively improving the accuracy of decrypted images in local detail restoration and global structural consistency, and solving the problems of local detail loss and global structure mismatch caused by insufficient multi-scale feature alignment in the prior art.
[0097] In one embodiment, the above step S10 includes:
[0098] S101, performing an edge filling operation on the image data to be processed to generate filled image data;
[0099] S102, setting the size parameters and step parameters of the sliding window;
[0100] S103, dividing the padded image data by the sliding window to generate an initial block sequence;
[0101] S104, arranging the initial block sequence according to the time sequence of the image data to be processed to generate a spatiotemporal block sequence;
[0102] S105 , performing normalization processing on the spatiotemporal block sequence to generate a standardized local block sequence.
[0103] In this embodiment, obtaining the image data to be processed refers to receiving the original image data from a storage device, an image acquisition device, or a data interface. The image data to be processed can be an encrypted image, a compressed image, or a data format that has undergone specific preprocessing. In different scenarios, it can include medical images, scanned copies of financial contracts, security monitoring data, etc. Edge filling of the image data to be processed refers to supplementing pixel values around the boundaries of the original image according to a specified strategy, so that subsequent block operations do not result in missing boundary data due to window movement. Edge filling can be performed using mirror filling, constant filling, or reflection filling. Mirror filling generates new pixels by copying the symmetrical values of boundary pixels, taking into account continuity and natural image transition effects.
[0104] Generating padded image data means adding a peripheral extension area to the original image. This operation ensures that all windows contain valid pixels during the sliding window division process, eliminating the need for special processing of edge anomalies. Setting the size and step parameters of the sliding window means defining the spatial range of each local block and the step interval of the window sliding. The size parameter is usually expressed as a fixed number of pixels to determine the width and height of the block, and the step parameter determines the degree of overlap between adjacent blocks. A smaller step size facilitates fine-grained feature extraction, while a larger step size improves computational efficiency. The step size and size can be dynamically adjusted to adapt to input images of different resolutions or compression ratios.
[0105] The padded image data is partitioned using a sliding window to generate an initial block sequence. This involves systematically scanning the image horizontally and vertically according to the set size and step size, intercepting local block areas, and organizing them into data sets in the order of capture. The initial block sequence maintains the spatial locality of the image and is suitable for subsequent spatiotemporal modeling. Arranging the initial block sequence according to the chronological order of the image data to be processed to generate a spatiotemporal block sequence means rearranging each group of local blocks along the temporal dimension based on the inter-frame timestamps or acquisition order in the presence of a data stream or time frame sequence, forming a feature stream that preserves local spatial information while reflecting dynamic changes in time.
[0106] Normalizing a spatiotemporal block sequence to generate a standardized local block sequence involves applying a uniform numerical normalization operation to the pixel values or feature values within each local block. This process, such as through mean normalization, maximum-minimum scaling, or zero-mean-unit-variance normalization, ensures a uniform numerical scale across different blocks, improving the stability and efficiency of subsequent feature extraction and learning. This normalization process can be performed based on statistics within the entire block or on a channel-by-channel basis to adapt to different types of data distribution characteristics.
[0107] In practical applications, appropriate edge padding strategies can be selected based on the source and characteristics of the image data being processed. For example, when processing medical images, mirror padding is preferred to maintain anatomical continuity. In financial contract or bill processing, reflection padding can be used to avoid artificial boundaries. The padding width can be dynamically calculated based on the sliding window size and is typically set to half the window size to ensure complete border coverage. The sliding window size parameter can be adaptively adjusted based on the input image resolution and the target block granularity. For high-resolution images, a larger window size can be set to reduce the number of blocks, while for low-resolution images, a smaller size can be used to preserve more local details. The step size parameter can be optimized based on the application scenario. For example, in image decryption tasks, a small step size with high overlap is preferred to improve recovery accuracy. When partitioning the initial block sequence, a row-first or column-first scanning order can be used. For streaming image input scenarios, a block sequence generation mode based on a sliding time window can be designed to update the local block set in real time. When arranging the spatiotemporal block sequence, blocks can be sorted based on time tags in the image metadata or using an external synchronization signal to ensure temporal consistency, which is particularly suitable for continuous frame acquisition scenarios. Normalization can be performed using a unified normalization layer for batch processing, or by embedding a data normalization module in real time during the block partitioning process to improve overall system processing efficiency. Normalization parameters can be preset based on global data statistics or dynamically updated through online learning to adapt to the distribution changes of different data batches.
[0108] This embodiment can effectively overcome the problem of feature loss caused by image boundary truncation by performing edge filling, sliding window division, time sequence arrangement and normalization on the processed image data, improve the consistency and integrity of local blocks, and at the same time, through the establishment of spatiotemporal block sequences, enhance the model's ability to capture the implicit temporal dynamic features in the encrypted image, and provide a unified, standardized and spatiotemporally continuous input data basis for subsequent feature extraction, cross-domain alignment and decryption recovery, significantly improving the decryption accuracy and robustness of the overall system.
[0109] In one embodiment, the above step S20 includes:
[0110] S201, performing a vectorization operation on each block element in the local block sequence to generate a block vector set;
[0111] S202, performing a linear transformation on the block vector set using a learnable projection matrix to generate a projection vector sequence;
[0112] S203, generating a corresponding position coding vector based on the spatiotemporal position information of each block element;
[0113] S204, adding the projection vector sequence and the position encoding vector element by element to generate a fusion feature sequence;
[0114] S205 , performing a layer normalization operation on the fused feature sequence to generate a normalized latent space representation.
[0115] In this embodiment, performing a vectorization operation on each block element in a local block sequence refers to flattening a two-dimensional or three-dimensional local image block into a one-dimensional vector. Each block element can be a fixed-size pixel matrix, including the channel dimension in the case of color images. The vectorization operation flattens the original spatially arranged pixel information into a one-dimensional array according to a fixed rule. This flattening operation can be performed in row-major or column-major order. In specific implementations, mapping can be performed based on the natural order of pixel arrangement within the block or according to a specific predefined order to meet the input requirements of a specific neural network architecture.
[0116] Generating a block vector set means that after all local block elements are vectorized, they are sequentially combined into a batched vector list. Each item represents the feature vector corresponding to a local block. The block vector set retains local detail features and provides a unified data format for subsequent linear mapping and spatiotemporal modeling.
[0117] Applying a linear transformation to a set of block vectors using a learnable projection matrix involves performing a linear mapping operation on the input block vectors, using a weight matrix whose parameters can be updated during training, to transform the original high-dimensional space into a fixed latent space. This learnable projection matrix is typically implemented using a fully connected layer or matrix multiplication unit defined in deep learning frameworks. Its parameters are trained using gradient backpropagation, enabling it to adaptively capture discriminative features within local blocks. The goal of this linear transformation is to compress the original feature space while extracting more discriminative feature representations.
[0118] Generating a corresponding position encoding vector based on the spatiotemporal position information of each block element involves generating a set of encoding vectors describing the relative or absolute position of the local block based on its two-dimensional position in the overall image and its temporal position in the time series. Position encoding vectors can use fixed position encoding, such as generating periodic vectors of varying frequencies based on sine and cosine functions, or learnable position encoding, which automatically learns the optimal encoding representation for each position during training. The introduction of position encoding provides the model with spatial and temporal prior knowledge, helping it understand the structural relationships and temporal changes between local blocks.
[0119] Element-wise addition of the projection vector sequence and the position encoding vector generates a fused feature sequence. This involves performing an element-by-element addition operation on each projection vector and its corresponding position encoding vector, resulting in a comprehensive feature representation that combines both local features and position information. This element-wise addition ensures that local content features and spatial position information are co-expressed on the same feature dimension, facilitating the capture of dynamic dependencies between blocks in subsequent temporal modeling.
[0120] Layer normalization of the fused feature sequence generates a standardized latent space representation. This involves independently performing mean normalization and variance normalization within each fused feature vector, ensuring that the numerical distribution of each fused feature vector has a uniform mean and variance. Layer normalization can eliminate inconsistent feature distributions between blocks due to differences in content complexity, improving the convergence speed and stability of subsequent model training. Layer normalization is typically implemented using a parameter-free normalization layer or a normalization layer with learnable scaling and offset parameters.
[0121] When implementing vectorized operations, direct flattening can be used to arrange the two-dimensional pixel matrix of each local block into a one-dimensional array in row-major order. When processing color images, this can be flattened in channel-major or alternating channel order. If the input block is large, preprocessing techniques such as downsampling or principal component analysis can be combined to further reduce the dimensionality and computational complexity. During the construction of the learnable projection matrix, the matrix dimensions can be determined based on the local block size and the target latent space dimensions. For example, if the local block is 1,000-dimensional after flattening, the projection matrix can be set to 1,000 rows and 500 columns, representing a mapping of 1,000-dimensional features to a 500-dimensional latent space. The projection matrix can be initialized using a normal distribution and then gradually optimized during training. For positional encoding, absolute position encoding can be used. Each position number is frequency-mapped using sine and cosine functions to construct a fixed encoding vector. The frequency range can be dynamically adjusted based on the number of blocks and the length of the time series. Alternatively, a learnable position encoding strategy can be employed, presetting a learnable vector for each possible position, initially randomly generated and adaptively adjusted during training. When generating the fused feature sequence, the projection vector and the position encoding vector of the corresponding position can be directly added together in each feature dimension. Alternatively, additional scaling factors can be introduced to attenuate or amplify the position information to account for the varying importance of position information in different scenarios. During the layer normalization phase, the mean and standard deviation of each fused feature vector can be independently calculated and normalized. This normalization process can be integrated into the input interface of subsequent network modules or inserted separately after the fusion phase to improve data consistency across the training and inference phases.
[0122] This embodiment can effectively capture local content features and spatiotemporal structure information in the latent space by vectorizing, linearly projecting, and fusing position encoding of local block sequences and performing normalization processing, providing unified, high-quality, time-series-aware, and scale-consistent feature input for subsequent time series modeling and domain adaptation processing, significantly improving the decryption model's ability to learn implicit dynamic features and local detail changes in encrypted images, thereby improving the accuracy and robustness of overall decryption recovery.
[0123] In one embodiment, the above step S30 includes:
[0124] S301, in the pre-trained temporal processing unit, performing a convolution operation on the latent space representation through a local attention window to generate target domain features;
[0125] S302, dividing the target domain features according to a preset short-range time window size, a medium-range time window size, and a long-range time window size to generate short-range time window features, medium-range time window features, and long-range time window features;
[0126] S303, performing a local detail enhancement operation on the short-range time window feature to generate a short-range local detail feature;
[0127] S304, performing a regional attention aggregation operation on the mid-range time window feature to generate a mid-range regional feature;
[0128] S305: Perform global self-attention analysis on the long-range time window features to generate long-range global structural features.
[0129] In this embodiment, in the pre-trained temporal processing unit, the convolution operation is performed on the latent space representation through the local attention window, which means that the latent space representation is used as the input feature, the local receptive field range is dynamically determined according to the local attention mechanism, and the convolution operation is performed in this local area to extract the target domain features. The local attention window can adaptively adjust the size and shape according to the spatiotemporal locality of the feature map to achieve sensitive modeling of dynamic changes in the local neighborhood. The convolution operation adopts a weight sharing mechanism and introduces local attention weights to assign different importance scores to different spatial positions, thereby improving the adaptability of feature extraction to local perturbations of the encrypted image.
[0130] The target domain features are divided according to the preset short-range time window size, medium-range time window size, and long-range time window size to generate short-range time window features, medium-range time window features, and long-range time window features. This means that the target domain features are divided into multiple sub-segments on the time axis based on the pre-set time scale boundaries. Short-range time window features mainly cover local dynamic information within a very small time span, such as subtle disturbances or small changes; medium-range time window features cover medium time spans and model continuous changes at the regional structural level; long-range time window features cover the overall trend of a large time span and are used to extract global dependencies. The time window size can be dynamically configured according to the task scenario. For example, in medical sequences or financial bill flows, an appropriate division threshold can be set based on the sampling rate or recording frequency.
[0131] Performing local detail enhancement on short-range temporal window features amplifies local dynamic detail changes within a short timescale, enhancing the representation of fine-grained features such as edge variations, local disturbances, and small movements. This enhancement is achieved through methods such as high-frequency component extraction modules, edge-enhanced convolution, and local attention weighting. This allows short-range local detail features to better highlight microscopic changes in encrypted data, improving the accuracy of detail reconstruction during the decryption phase.
[0132] Regional attention aggregation of mid-range temporal window features assigns attention weights based on the spatial correlation of features within a local region within a medium time span. This is done through weighted summation, aggregation, or feature compression within the region to extract stable feature patterns at the regional level. Regional attention aggregation effectively captures the local continuity and regional consistency of encrypted images within the medium-range temporal sequence, generating mid-range regional features and improving the ability to model local pattern migration and deformation.
[0133] Performing global self-attention analysis on long-range temporal window features involves calculating the similarity between any two positions within the long-range temporal window over a large time scale through the self-attention mechanism to extract global dynamic dependency information. Global self-attention analysis can be implemented using standard self-attention layers, such as the Transformer self-attention unit based on dot-product similarity. This can model long-range, cross-temporal dependencies and extract long-range global structural features, providing critical global semantic information for subsequent reconstruction of the overall image structure.
[0134] When implementing local attention window convolution, a deformable convolution unit can be used to adaptively adjust the sampling position of the convolution kernel, thereby flexibly capturing local perturbations in the encrypted image. Dynamic convolution techniques can also be combined to adaptively adjust the convolution weight distribution based on local content to accommodate variations in feature strength across different regions. During time windowing, thresholds for the short, medium, and long time windows can be set based on the application scenario. For example, in healthcare scenarios, the short-range time window can be set to two to four frames, the medium-range time window to five to ten frames, and the long-range time window to more than ten frames. In fintech scenarios, the time windowing scale can be adjusted based on the temporal density of transaction records or the frame rate of the bill scanning stream. Local detail enhancement can be achieved by designing high-frequency convolution kernels with small receptive fields to enhance features targeting local textures or subtle edge changes. Local noise suppression mechanisms can also be introduced to improve local feature clarity. Regional attention aggregation can be achieved by constructing a local self-attention submodule to generate an attention weight matrix based on feature similarity within each medium-range time window. Weighted feature summation is then performed to obtain a comprehensive feature representation at the regional level. Global self-attention analysis can achieve global feature modeling across time by stacking multiple layers of a standard Transformer Encoder structure. Each long-range time window feature is input as an input token to the self-attention module, which calculates the attention score matrix to achieve feature aggregation and capture global correlations.
[0135] This embodiment introduces local attention convolution, sub-scale time window division, local detail enhancement, regional attention aggregation and global self-attention analysis into the pre-trained timing processing unit. It can not only fully capture the micro-perturbations and local continuity features in the encrypted image, but also effectively extract the global dependency information of long-term cross-temporal dynamic changes. Ultimately, the extracted multi-scale target domain features are significantly superior to traditional single-scale feature extraction methods in spatiotemporal dynamic perception and structural alignment capabilities, providing richer and more hierarchical feature support for subsequent cross-domain alignment and decryption reconstruction.
[0136] In one embodiment, the above step S40 includes:
[0137] S401, activating a short-range domain adaptation parameter, a medium-range domain adaptation parameter, and a long-range domain adaptation parameter in the time series processing unit;
[0138] S402, performing discriminator forward analysis on the short-range local detail features, the medium-range regional features, and the long-range global structure features based on the short-range domain adaptation parameters, the medium-range domain adaptation parameters, and the long-range domain adaptation parameters, respectively, to generate a short-range discrimination result, a medium-range discrimination result, and a long-range discrimination result;
[0139] S403, determining, based on the gradient penalty coefficient in the short-range domain adaptation parameter, the gradient penalty coefficient in the medium-range domain adaptation parameter, and the gradient penalty coefficient in the long-range domain adaptation parameter, the gradient constraint conditions between the short-range local detail feature and the pre-trained source domain short-range feature distribution, the gradient constraint conditions between the medium-range regional feature and the pre-trained source domain medium-range feature distribution, and the gradient constraint conditions between the long-range global structure feature and the pre-trained source domain long-range feature distribution;
[0140] S404: randomly interpolating the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding source domain features according to the interpolation sample generation strategy in the domain adaptation parameters to generate cross-domain interpolation samples and verify their continuity;
[0141] S405, based on the short-range discrimination results, medium-range discrimination results, long-range discrimination results and corresponding gradient constraints, update the discriminator weight matrix in the short-range domain adaptation parameters, medium-range domain adaptation parameters and long-range domain adaptation parameters to achieve alignment processing of the short-range local detail features, medium-range regional features and long-range global structure features with the corresponding scale feature distribution of the pre-trained source domain.
[0142] In this embodiment, activating the short-range domain adaptation parameters, medium-range domain adaptation parameters, and long-range domain adaptation parameters in the timing processing unit refers to loading the domain alignment parameter sets learned during the pre-training process for the short-range local detail features, medium-range regional features, and long-range global structural features extracted at different time scales. Each set of domain adaptation parameters corresponds to a feature distribution correction task at one scale, mainly including the weight matrix, bias term, and gradient penalty correlation coefficient of the discriminator. Through the activation operation, these parameters are loaded into the timing processing unit, so that subsequent feature discrimination and alignment calculations can be independently optimized and adjusted according to the distribution characteristics of features at different scales.
[0143] Using features of corresponding scales as input, the discriminant score for the domain (source or target domain) to which the features belong is output through a forward inference process associated with a pre-set discriminator structure. The discriminator structure is typically a multi-layer perceptron, convolutional neural network, or lightweight Transformer unit. Its forward analysis process includes feature mapping, nonlinear activation, and output classification scores to evaluate the degree of match between features of different scales and the pre-trained source domain distribution.
[0144] To prevent discriminator gradient explosion or collapse when training the discriminator or performing alignment optimization, the discriminator score output for the input features is subjected to gradient norm control based on the gradient penalty coefficient defined in the domain adaptation parameters at different scales. The gradient penalty condition is set according to the principle of stability in adversarial learning and is typically applied to the discriminator gradient via an L2 regularization term to ensure stable model convergence during feature alignment and continuous differentiability of the feature map.
[0145] Based on the random linear interpolation strategy in the Wasserstein GAN, the target features are linearly combined with the source domain features at the corresponding scale in the feature space to generate new cross-domain interpolated feature samples. This interpolation method establishes a continuous mapping relationship between the source and target domain feature distributions. The discriminator is further used to detect the consistency and coherence of the interpolated samples, thereby promoting a smooth transition of the overall feature space at multiple scales and narrowing the gap between the source and target domain feature distributions.
[0146] Based on the short-range discrimination results, medium-range discrimination results, long-range discrimination results and corresponding gradient constraints, the discriminator weight matrices in the short-range domain adaptation parameters, medium-range domain adaptation parameters and long-range domain adaptation parameters are updated to achieve alignment of short-range local detail features, medium-range regional features and long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain. This refers to performing backpropagation and optimizing the domain adaptation discriminator parameters for each scale based on the difference between the discriminant score obtained by forward analysis and the true label, combined with the gradient penalty condition. Through this parameter update mechanism, the feature distribution differences between the source domain and the target domain of different scale features are gradually converged, ultimately achieving alignment of short-range, medium-range and long-range scale features in cross-domain scenarios, providing stable multi-scale domain consistency support for subsequent feature fusion and image reconstruction.
[0147] This embodiment introduces mechanisms such as multi-scale domain adaptation parameter activation, discriminator sub-scale forward analysis, sub-scale gradient constraint setting, and interpolation sample continuity verification within the timing processing unit. This can effectively alleviate the distribution differences between encrypted data and unencrypted source domain data at three scales: local details, mid-range areas, and long-range structures. It can achieve layer-by-layer alignment of dynamic features at different scales, significantly improve the continuity and stability of cross-domain feature migration, and provide a solid foundation for detail restoration and global structure recovery of decrypted images.
[0148] In one embodiment, the above step S50 includes:
[0149] S501, respectively performing linear transformation on the aligned short-range local detail features, medium-range regional features, and long-range global structure features using pre-trained short-range attention weight matrices, medium-range attention weight matrices, and long-range attention weight matrices to generate short-range attention weight values, medium-range attention weight values, and long-range attention weight values;
[0150] S502, normalizing the short-range attention weight value, the medium-range attention weight value, and the long-range attention weight value to generate a normalized short-range weight, a normalized medium-range weight, and a normalized long-range weight;
[0151] S503, multiplying the aligned short-range local detail features by the normalized short-range weights by channel to generate weighted short-range features;
[0152] S504, multiplying the aligned mid-range region features by the normalized mid-range weights by channel to generate weighted mid-range features;
[0153] S505, multiplying the aligned long-range global structural features by the normalized long-range weights by channel to generate weighted long-range features;
[0154] S506 , performing channel concatenation and linear compression on the weighted short-range features, weighted mid-range features, and weighted long-range features to generate multi-scale fusion features.
[0155] In this embodiment, before fusion, independent linear mappings are applied to the feature sequences at each scale based on the differences in feature importance at each scale to form a preliminary attention response. The pre-trained short-range, medium-range, and long-range attention weight matrices are learned through the offline training phase to capture the importance patterns of features at different time scales. The linear transformation operation is usually implemented as matrix multiplication, mapping the original features to an attention weight space to facilitate subsequent normalization and weighting processing.
[0156] Normalization is performed on the attention weights generated at each scale, for example using the Softmax function, L2 normalization, or maximum normalization, to bring each set of weights to a uniform range. This ensures comparability across scales and avoids feature fusion bias caused by inconsistent scales. Normalization not only ensures numerical stability during fusion but also enhances the dynamic balance between features at different scales.
[0157] The aligned short-range local detail features are multiplied by the normalized short-range weights on a channel-by-channel basis to generate weighted short-range features. This involves performing a per-channel weighting operation on each channel of the short-range local detail features based on the corresponding normalized short-range weights. This emphasizes the short-term local details that are most critical to overall feature recovery and suppresses redundant or noisy features. Channel multiplication can be implemented through a broadcast mechanism or dot product, allowing each channel feature map to carry different importance tags during the fusion process.
[0158] By performing channel-wise multiplication with the normalized mid-range weights, the importance weights of each region in the mid-range features are dynamically adjusted, emphasizing feature components with significant regional correlations or key structural information. In this way, mid-range features can more accurately reflect the structural continuity of local regions in the image.
[0159] At the long-range feature level, channel weighting is performed based on normalized long-range weights to highlight key dimensions in global structural relationships and suppress unnecessary long-range interference information. This weighting mechanism helps preserve complete global semantic coherence in the final feature fusion.
[0160] After weighting the features at each scale, the three scale features are first concatenated in the channel dimension to form a unified high-dimensional feature vector. Subsequently, a linear compression operation, typically a fully connected layer or a one-dimensional convolutional layer, is applied to map the high-dimensional features to a preset fusion feature dimension. This avoids the waste of computational resources caused by an excessive number of channels and promotes information interaction and complementation between features at different scales. This fusion mechanism enhances overall structural modeling capabilities while preserving the integrity of local details, ultimately providing a rich and coordinated feature foundation for target data reconstruction.
[0161] This embodiment uses multi-scale independent weight generation, normalization, weighted fusion and linear compression mechanisms to give full play to the complementary effects of different time scale features in local details, medium-range structure and long-range semantics, achieve coordinated optimization of short-range, medium-range and long-range information, and significantly improve the expressiveness and semantic coherence after feature fusion, thereby better maintaining the local texture details and global structural integrity of the image during the target data reconstruction process, and alleviating the common problems of detail loss and structural dislocation when fusing features of different scales.
[0162] In one embodiment, the above step S60 includes:
[0163] S601, performing a layer-by-layer upsampling operation on the multi-scale fusion features to generate a preliminary upsampling feature map;
[0164] S602, transferring the local block sequence to the corresponding upsampling layer via a skip connection;
[0165] S603, splicing the preliminary up-sampled feature map and the local block sequence transferred by the skip connection according to channels to generate an enhanced feature map;
[0166] S604, performing a residual convolution operation on the enhanced feature map, sequentially passing through a convolution layer, an activation function layer and a residual connection to generate a reconstructed feature map;
[0167] S605: Perform pixel-by-pixel normalization on the reconstructed feature map to generate standardized target image data.
[0168] In this embodiment, after the multi-scale fusion features are extracted, a layer-by-layer upsampling operation is first performed. Layer-by-layer upsampling generally refers to the continuous improvement of the spatial resolution of the feature map, and each layer of upsampling corresponds to a scale magnification. The upsampling operation can be implemented through a variety of means, such as deconvolution, bilinear interpolation, or sub-pixel convolution technology. The deconvolution operation achieves controllable feature restoration through a learnable convolution kernel, bilinear interpolation improves spatial continuity through a fixed weighted calculation, and sub-pixel convolution uses a specific data rearrangement strategy to reduce the number of model parameters. During the implementation process, the number of upsampling layers can be set according to the expected output size of the target image data. After each layer of upsampling, the size of the feature map gradually approaches the size of the original image, forming a preliminary upsampled feature map.
[0169] As low- and mid-level features retained during the encoding process, local block sequences need to be introduced during the upsampling phase of the decoder to compensate for the loss of detail information that may occur during the layer-by-layer upsampling process. The skip connection mechanism allows the local block sequence to be directly connected to the preliminary upsampled feature map of the corresponding scale. The skip connection can use element-level fusion methods such as channel splicing or element addition, where channel splicing is more conducive to preserving the original information. During the skip connection transmission process, adaptive mapping is required to keep the feature dimensions of the local block sequence consistent with the upsampled feature map. One-dimensional convolution or linear mapping matrix can be used for dimensional adjustment.
[0170] The initial upsampled feature map is fused with the sequence of local blocks transferred via skip connections through channel-wise splicing to generate an enhanced feature map. Channel-wise splicing can be understood as superimposing feature tensors from two different sources along the feature channel dimension to form a more informative fused feature. The spliced enhanced feature map maintains its spatial structure while increasing its feature expressiveness, providing a more comprehensive information foundation for subsequent reconstruction of the target image data.
[0171] The enhanced feature map is further optimized through residual convolution. This operation consists of a convolutional layer, an activation function layer, and a residual connection. The convolutional layer extracts local features, while the activation function layer (such as ReLU and GELU) introduces nonlinear transformations to enhance representation. The residual connection directly superimposes the input and the convolution output to alleviate the vanishing gradient problem in deep networks. In practice, a set of standard convolution kernels can be used for local perception, followed by element-by-element addition to achieve feature preservation and fine-grained enhancement.
[0172] The resulting reconstructed feature map needs to be converted into standardized target image data through pixel-by-pixel normalization. Pixel-by-pixel normalization typically uses an activation function, such as the Sigmoid function, to map feature values to a fixed range (e.g., 0 to 1) to match the numerical format requirements of the output image data. Normalization not only ensures that the output data meets downstream usage standards but also improves model stability and reduces the impact of numerical anomalies on image quality.
[0173] For example, the system first receives the encrypted image input, divides it into overlapping local block sequences through the input preprocessing layer and maps it to the latent space to form a spatiotemporal sequence representation. These sequences are then input into the enhanced temporal processing unit (TPU), which captures the local temporal dependencies through local attention windows, while using the global domain adaptation module to achieve distribution alignment of non-encrypted temporal data and encrypted image features. The multi-scale attention mechanism inside the TPU fuses features of different time scales to generate a representation with rich semantic information. Finally, the processed features are used to reconstruct the decrypted image through the residual convolution decoder, which uses a U-Net architecture and retains high-frequency details through jump connections to achieve high-quality image reconstruction. The entire process does not rely on prior knowledge or keys of the encryption algorithm, and decryption is achieved entirely through adaptive learning of the neural network.
[0174] The input preprocessing layer is responsible for converting the encrypted image ( represents the set of real numbers) into overlapping local blocks Each block Corresponding to a time step in the spatiotemporal sequence. Where H represents the image height, W represents the width, C represents the number of channels, k is the block size, and T is the total number of blocks. Through the learnable embedding matrix (where d = k 2 C represents the flattened block dimension, m represents the embedding dimension) Linearly project the block into a latent vector:
[0175] z t =E T vec(P t )+p t
[0176] Where vec(·) represents the operation of flattening the block into a one-dimensional vector, () T represents the transpose of the matrix, represents position encoding, and represents matrix multiplication operation, which is used to preserve the spatial relative position information of the blocks and ensure that the model can understand the spatial structure of the image.
[0177] The enhanced TPU is one of the components of the system and consists of the following three submodules:
[0178] Local attention window module: Calculates query-key-value attention within a sliding window τ to capture local dependencies in the time series:
[0179]
[0180] Where W Q and is the learnable projection matrix, α (t,j) represents the attention weight of time step j to t, z t The feature vector at time t is the representation of the input feature in the time series processing unit, z j Represents the eigenvector at time j, which is the same as z t Form the target feature of attention calculation, W Q and are respectively the query matrix and key-value matrix that can be learned, m represents the feature dimension, (z t W Q ) represents the time t feature z t By querying the matrix W Q Convert to query vector; (z j W K ) represents the time j feature z j Through the key matrix W K Convert to key vector; is a scaling factor used to prevent large values from affecting gradient stability; j∈[t-τ,t+τ] indicates that the attention window covers the time range from t-τ to t+τ; and softmax(·) represents a normalization function that converts vector elements into probability distributions. This local attention mechanism effectively captures temporal patterns in encrypted images.
[0181] Global Domain Adaptation Module: This module uses Wasserstein GAN to align the distribution of unencrypted temporal data (source domain S) with encrypted image features (target domain T).
[0182]
[0183] Where, L (adv) represents the adversarial loss function, which is used to measure the difference between the feature distributions of the source domain and the target domain; Represents the source domain feature z s The expected output of the discriminator D is the source domain feature z s is sampled from the source domain data distribution S, D(z s ) represents the discriminant’s discrimination result on the source domain features, and the mathematical expectation E represents the average value of all source domain features, which measures the overall output of the discriminator on the source domain.
[0184] Represents the target domain feature zt The expected output of the discriminator D is the target domain feature z t is sampled from the target domain data distribution T, D(z t ) represents the discriminator’s discrimination result on the target domain features, and the mathematical expectation E represents the average value of all target domain features, which measures the overall output of the discriminator on the target domain.
[0185] represents a random interpolation point between the source domain and the target domain, Represents a variable The operation of finding the gradient, λ is a hyperparameter that controls the intensity of the gradient penalty, and ||·||2 represents the L2 norm (Euclidean norm). This adversarial training method can effectively narrow the feature distribution difference between the source domain and the target domain, promoting cross-domain knowledge transfer.
[0186] Multi-scale attention fusion module: defining hierarchical attention weights β l Control the fusion of features at different time scales:
[0187]
[0188] in, Represents the output features after multi-scale fusion at time t; ∑(l=1) L It represents the summation of different time scales l, where L is the number of multiple scales (such as short-range, medium-range, and long-range scales), and L = 3 represents the three time scales of short / medium / long; Represents a specific time segment feature selected from the input feature sequence, where: z is the input feature sequence; t is the current time point; Δl represents the time window offset of the lth scale, t-Δ l +Δ l Ensure that features are acquired within this time window to capture multi-scale temporal information; TPU l (·) corresponds to processing units of different scales, β l This multi-scale design can capture both local details and global structures, improving the expressive power of the model.
[0189] Introducing the adaptive convergence factor η under the ZNN framework t , to achieve dynamic gradient back propagation:
[0190]
[0191] in, Represents the matrix A at time t t The transpose of A t is a dynamic weight matrix that defines the eigenvector z tThe transformation relationship at each time point can be a learnable matrix or a predefined fixed matrix, depending on the application scenario. The transposition operation Indicates the direction of back propagation, often used for gradient calculation. t represents the bias vector at time t, which is used to introduce translation or offset in feature transformation to increase the flexibility of the model. t are usually learnable parameters.
[0192] A t ·z t Indicates that at time t, the eigenvector z t Through the matrix A t The linear transformation A t It is the feature transformation matrix, which defines the linear combination relationship between features.
[0193] dz t / dt means z t The derivative with respect to time t, describing z t The instantaneous rate of change over time, sgn(·) represents the sign function (output 1 when the input is positive, -1 when it is negative, and 0 when it is zero), η t By the attention weight α (t,j) Dynamic adjustment: This mechanism enables the model to adaptively adjust the optimization process and improve convergence efficiency.
[0194] The residual convolution decoder uses the U-Net architecture to fuse TPU outputs of different scales through skip connections:
[0195]
[0196] in, The target image data is the reconstructed image generated by multi-scale upsampling and feature fusion. G represents the final generation network or decoding network (such as a convolutional neural network, U-Net, or other custom network structure), which is used to combine the multi-scale upsampled features and further process them to generate the final high-resolution target image. G contains five layers of transposed convolution, each followed by InstanceNorm (instance normalization) and LeakyReLU activation function. Concat[·] represents the feature concatenation operation, UpSample(·) represents the upsampling operation, h1..., h T Represents multi-scale features extracted from feature maps of different scales (e.g., short-range, medium-range, and long-range features). These features can be derived from the features separated in the previous temporal processing unit, such as short-range local detail features, medium-range regional features, and long-range global structural features. This design effectively preserves high-frequency details, avoids the information loss problem of traditional decoders, and improves the quality of the decrypted image.
[0197] This embodiment gradually upsamples the multi-scale fusion features and introduces fine-grained features of local block sequences to supplement them, allowing the decoder to simultaneously maintain local detail information and global structural continuity when generating target image data. The introduction of channel splicing and residual convolution further enhances the feature expression capabilities of the enhanced feature map while retaining the original semantic information. The pixel-by-pixel normalization step ensures the standardization and usability of the output image data in terms of numerical range, improving the visual quality of the decrypted image data and its adaptability to subsequent processing.
[0198] In one embodiment, a data processing device based on cross-domain feature alignment is provided, and the data processing device based on cross-domain feature alignment corresponds one-to-one to the data processing method based on cross-domain feature alignment in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of a data processing device based on cross-domain feature alignment according to the present invention. It includes a block preprocessing module 10, a feature encoding module 20, a temporal feature extraction module 30, a feature alignment module 40, a feature fusion module 50, and an image generation module 60. Each functional module is described in detail below:
[0199] The block preprocessing module 10 is used to obtain the image data to be processed, and perform block processing on the image data to be processed to generate a local block sequence;
[0200] A feature encoding module 20 is used to transform the local block sequence through linear projection and position encoding to generate a latent space representation;
[0201] a temporal feature extraction module 30 for inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structural features;
[0202] A feature alignment module 40 is configured to align the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain using the domain adaptation parameters in the temporal processing unit;
[0203] A feature fusion module 50 is used to fuse the aligned short-range local detail features, medium-range regional features, and long-range global structural features through a hierarchical attention weight distribution mechanism to generate multi-scale fused features;
[0204] The image generation module 60 is configured to generate target image data based on the multi-scale fusion features.
[0205] In one embodiment, the block preprocessing module 10 is specifically configured to:
[0206] Performing an edge filling operation on the image data to be processed to generate filled image data;
[0207] Set the size and step parameters of the sliding window;
[0208] Dividing the padded image data by the sliding window to generate an initial block sequence;
[0209] Arranging the initial block sequence according to the time sequence of the image data to be processed to generate a spatiotemporal block sequence;
[0210] Normalization is performed on the spatiotemporal block sequence to generate a standardized local block sequence.
[0211] In one embodiment, the feature encoding module 20 is specifically configured to:
[0212] Performing a vectorization operation on each block element in the local block sequence to generate a block vector set;
[0213] Performing a linear transformation on the block vector set using a learnable projection matrix to generate a projection vector sequence;
[0214] Generate a corresponding position encoding vector based on the spatiotemporal position information of each block element;
[0215] Adding the projection vector sequence and the position encoding vector element by element to generate a fusion feature sequence;
[0216] A layer normalization operation is performed on the fused feature sequence to generate a normalized latent space representation.
[0217] In one embodiment, the temporal feature extraction module 30 is specifically configured to:
[0218] In the pre-trained temporal processing unit, performing a convolution operation on the latent space representation through a local attention window to generate target domain features;
[0219] Dividing the target domain features according to the preset short-range time window size, medium-range time window size, and long-range time window size to generate short-range time window features, medium-range time window features, and long-range time window features;
[0220] Performing a local detail enhancement operation on the short-range time window feature to generate a short-range local detail feature;
[0221] Performing a regional attention aggregation operation on the mid-range time window features to generate mid-range regional features;
[0222] Performing global self-attention analysis on the long-range time window features to generate long-range global structural features.
[0223] In one embodiment, the feature alignment module 40 is specifically configured to:
[0224] activating a short-range domain adaptation parameter, a medium-range domain adaptation parameter, and a long-range domain adaptation parameter in the time series processing unit;
[0225] Based on the short-range domain adaptation parameters, the medium-range domain adaptation parameters, and the long-range domain adaptation parameters, respectively, a discriminator forward analysis is performed on the short-range local detail features, the medium-range regional features, and the long-range global structure features to generate a short-range discrimination result, a medium-range discrimination result, and a long-range discrimination result, respectively;
[0226] Determining, according to the gradient penalty coefficient in the short-range domain adaptation parameter, the gradient penalty coefficient in the medium-range domain adaptation parameter, and the gradient penalty coefficient in the long-range domain adaptation parameter, the gradient constraint conditions between the short-range local detail feature and the pre-trained source domain short-range feature distribution, the gradient constraint conditions between the medium-range regional feature and the pre-trained source domain medium-range feature distribution, and the gradient constraint conditions between the long-range global structure feature and the pre-trained source domain long-range feature distribution;
[0227] According to the interpolation sample generation strategy in the domain adaptation parameters, the short-range local detail features, the medium-range regional features, and the long-range global structure features are randomly interpolated with the corresponding source domain features to generate cross-domain interpolation samples and verify continuity;
[0228] According to the short-range discrimination results, medium-range discrimination results, long-range discrimination results and corresponding gradient constraints, the discriminator weight matrices in the short-range domain adaptation parameters, medium-range domain adaptation parameters and long-range domain adaptation parameters are updated to achieve alignment processing of the short-range local detail features, medium-range regional features and long-range global structure features with the corresponding scale feature distribution of the pre-trained source domain.
[0229] In one embodiment, the feature fusion module 50 is specifically configured to:
[0230] Performing linear transformation on the aligned short-range local detail features, medium-range regional features, and long-range global structure features using pre-trained short-range attention weight matrices, medium-range attention weight matrices, and long-range attention weight matrices, respectively, to generate short-range attention weight values, medium-range attention weight values, and long-range attention weight values;
[0231] Normalizing the short-range attention weight value, the medium-range attention weight value, and the long-range attention weight value to generate a normalized short-range weight, a normalized medium-range weight, and a normalized long-range weight;
[0232] Multiplying the aligned short-range local detail features by the normalized short-range weights by channel to generate weighted short-range features;
[0233] Multiplying the aligned mid-range region features by the normalized mid-range weights by channel to generate weighted mid-range features;
[0234] Multiplying the aligned long-range global structural features by the normalized long-range weights by channel to generate weighted long-range features;
[0235] Channel concatenation and linear compression are performed on the weighted short-range features, weighted mid-range features, and weighted long-range features to generate multi-scale fusion features.
[0236] In one embodiment, the image generation module 60 is specifically configured to:
[0237] Performing a layer-by-layer upsampling operation on the multi-scale fusion features to generate a preliminary upsampling feature map;
[0238] Passing the local block sequence to the corresponding upsampling layer through a skip connection;
[0239] splicing the preliminary up-sampled feature map with the local block sequence transferred by the skip connection according to channels to generate an enhanced feature map;
[0240] Performing a residual convolution operation on the enhanced feature map, sequentially passing through a convolution layer, an activation function layer and a residual connection to generate a reconstructed feature map;
[0241] A pixel-by-pixel normalization operation is performed on the reconstructed feature map to generate standardized target image data.
[0242] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a data processing method based on cross-domain feature alignment.
[0243] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a data processing method based on cross-domain feature alignment.
[0244] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0245] Acquiring image data to be processed, and performing block processing on the image data to be processed to generate a local block sequence;
[0246] Transforming the local block sequence through linear projection and position encoding to generate a latent space representation;
[0247] Inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structure features;
[0248] Aligning the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain, respectively, through the domain adaptation parameters in the temporal processing unit;
[0249] Through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fusion features;
[0250] Target image data is generated based on the multi-scale fusion features.
[0251] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0252] Acquiring image data to be processed, and performing block processing on the image data to be processed to generate a local block sequence;
[0253] Transforming the local block sequence through linear projection and position encoding to generate a latent space representation;
[0254] Inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structure features;
[0255] Aligning the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain, respectively, through the domain adaptation parameters in the temporal processing unit;
[0256] Through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fusion features;
[0257] Target image data is generated based on the multi-scale fusion features.
[0258] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0259] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0260] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0261] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A data processing method based on cross-domain feature alignment, characterized in that: The following steps are involved: Acquiring image data to be processed, and performing block processing on the image data to be processed to generate a local block sequence; Transforming the local block sequence through linear projection and position encoding to generate a latent space representation; Inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structure features; Aligning the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain, respectively, through the domain adaptation parameters in the temporal processing unit; Through the hierarchical attention weight allocation mechanism, the aligned short-range local detail features, medium-range regional features and long-range global structure features are fused to generate multi-scale fusion features; Target image data is generated based on the multi-scale fusion features.
2. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: Acquiring image data to be processed and performing block processing on the image data to be processed to generate a local block sequence, including: Performing an edge filling operation on the image data to be processed to generate filled image data; Set the size and step parameters of the sliding window; Dividing the padded image data by the sliding window to generate an initial block sequence; Arranging the initial block sequence according to the time sequence of the image data to be processed to generate a spatiotemporal block sequence; Normalization is performed on the spatiotemporal block sequence to generate a standardized local block sequence.
3. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: The local block sequence is transformed through linear projection and position encoding to generate a latent space representation, including: Performing a vectorization operation on each block element in the local block sequence to generate a block vector set; Performing a linear transformation on the block vector set using a learnable projection matrix to generate a projection vector sequence; Generate a corresponding position encoding vector based on the spatiotemporal position information of each block element; Adding the projection vector sequence and the position encoding vector element by element to generate a fusion feature sequence; A layer normalization operation is performed on the fused feature sequence to generate a normalized latent space representation.
4. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: Inputting the latent space representation into a pre-trained temporal processing unit, extracting target domain features from the latent space representation, and decomposing the target domain features into short-range local detail features, mid-range regional features, and long-range global structural features, including: In the pre-trained temporal processing unit, performing a convolution operation on the latent space representation through a local attention window to generate target domain features; Dividing the target domain features according to the preset short-range time window size, medium-range time window size, and long-range time window size to generate short-range time window features, medium-range time window features, and long-range time window features; Performing a local detail enhancement operation on the short-range time window feature to generate a short-range local detail feature; Performing a regional attention aggregation operation on the mid-range time window features to generate mid-range regional features; Performing global self-attention analysis on the long-range time window features to generate long-range global structural features.
5. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: The short-range local detail features, the mid-range regional features, and the long-range global structure features are respectively aligned with the corresponding scale feature distributions of the pre-trained source domain using the domain adaptation parameters in the temporal processing unit, including: activating a short-range domain adaptation parameter, a medium-range domain adaptation parameter, and a long-range domain adaptation parameter in the time series processing unit; Based on the short-range domain adaptation parameters, the medium-range domain adaptation parameters, and the long-range domain adaptation parameters, respectively, a discriminator forward analysis is performed on the short-range local detail features, the medium-range regional features, and the long-range global structure features to generate a short-range discrimination result, a medium-range discrimination result, and a long-range discrimination result, respectively; Determining, according to the gradient penalty coefficient in the short-range domain adaptation parameter, the gradient penalty coefficient in the medium-range domain adaptation parameter, and the gradient penalty coefficient in the long-range domain adaptation parameter, the gradient constraint conditions between the short-range local detail feature and the pre-trained source domain short-range feature distribution, the gradient constraint conditions between the medium-range regional feature and the pre-trained source domain medium-range feature distribution, and the gradient constraint conditions between the long-range global structure feature and the pre-trained source domain long-range feature distribution; According to the interpolation sample generation strategy in the domain adaptation parameters, the short-range local detail features, the medium-range regional features, and the long-range global structure features are randomly interpolated with the corresponding source domain features to generate cross-domain interpolation samples and verify continuity; According to the short-range discrimination results, medium-range discrimination results, long-range discrimination results and corresponding gradient constraints, the discriminator weight matrices in the short-range domain adaptation parameters, medium-range domain adaptation parameters and long-range domain adaptation parameters are updated to achieve alignment processing of the short-range local detail features, medium-range regional features and long-range global structure features with the corresponding scale feature distribution of the pre-trained source domain.
6. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: Through the hierarchical attention weight distribution mechanism, the aligned short-range local detail features, medium-range regional features, and long-range global structural features are fused to generate multi-scale fusion features, including: Performing linear transformation on the aligned short-range local detail features, medium-range regional features, and long-range global structure features using pre-trained short-range attention weight matrices, medium-range attention weight matrices, and long-range attention weight matrices, respectively, to generate short-range attention weight values, medium-range attention weight values, and long-range attention weight values; Normalizing the short-range attention weight value, the medium-range attention weight value, and the long-range attention weight value to generate a normalized short-range weight, a normalized medium-range weight, and a normalized long-range weight; Multiplying the aligned short-range local detail features by the normalized short-range weights by channel to generate weighted short-range features; Multiplying the aligned mid-range region features by the normalized mid-range weights by channel to generate weighted mid-range features; Multiplying the aligned long-range global structural features by the normalized long-range weights by channel to generate weighted long-range features; Channel concatenation and linear compression are performed on the weighted short-range features, weighted mid-range features, and weighted long-range features to generate multi-scale fusion features.
7. The data processing method based on cross-domain feature alignment according to claim 1, characterized in that: Generating target image data based on the multi-scale fusion features includes: Performing a layer-by-layer upsampling operation on the multi-scale fusion features to generate a preliminary upsampling feature map; Passing the local block sequence to the corresponding upsampling layer through a skip connection; Splicing the preliminary up-sampled feature map with the local block sequence transferred by the skip connection by channel to generate an enhanced feature map; Performing a residual convolution operation on the enhanced feature map, sequentially passing through a convolution layer, an activation function layer and a residual connection to generate a reconstructed feature map; A pixel-by-pixel normalization operation is performed on the reconstructed feature map to generate standardized target image data.
8. A data processing device based on cross-domain feature alignment, characterized in that: The data processing device based on cross-domain feature alignment includes: A block preprocessing module is used to obtain image data to be processed, and perform block processing on the image data to be processed to generate a local block sequence; A feature encoding module, configured to transform the local block sequence through linear projection and position encoding to generate a latent space representation; a temporal feature extraction module, configured to input the latent space representation into a pre-trained temporal processing unit, extract target domain features from the latent space representation, and decompose the target domain features into short-range local detail features, medium-range regional features, and long-range global structural features; a feature alignment module, configured to align the short-range local detail features, the mid-range regional features, and the long-range global structural features with the corresponding scale feature distributions of the pre-trained source domain using the domain adaptation parameters in the temporal processing unit; The feature fusion module is used to fuse the aligned short-range local detail features, medium-range regional features, and long-range global structural features through a hierarchical attention weight distribution mechanism to generate multi-scale fused features; An image generation module is used to generate target image data based on the multi-scale fusion features.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a data processing program based on cross-domain feature alignment stored in the memory and capable of running on the processor. When the data processing program based on cross-domain feature alignment is executed by the processor, the steps of the data processing method based on cross-domain feature alignment as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a data processing program based on cross-domain feature alignment, and when the data processing program based on cross-domain feature alignment is executed by the processor, the steps of the data processing method based on cross-domain feature alignment according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Image watermarking method and system based on decoding guided by self-supervised visual model
CN121504703A
Image watermarking method and system based on self-supervised visual model-guided decoding
CN121504703B
Flow dividing and converging lane detection method and device
CN121708339A
Target attribute identification method, computer equipment and storage medium
CN121937805A
Intrusion detection method for communication security of networked vehicles
CN122120031A