A diffusion encoder representation training algorithm based on timing law perception

CN122819342APending Publication Date: 2026-09-25NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611015140.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

如果仅约束全局特征,模型可能丢失空间敏感信息;如果始终强约束局部结构,模型又可能在高噪声阶段拟合不可靠扰动

Benefits of technology

[0091]1)由稳定参照目标生成支路产生的有益效果:该支路以轻扰动参照视图为输入,经参照编码器、多阶段特征提取器、全局投影头及局部结构投影头处理后,通过停止梯度操作输出全局语义参照表征和局部结构参照表征;由于该支路参数不直接接收当前批次损失的反向传播梯度,而是通过指数滑动平均由推断支路参数缓慢更新,因此能够提供跨迭代稳定的参照目标,避免单批次中的随机噪声、强增强扰动或异常样本对目标表征造成剧烈波动,从而降低非语义扰动占用潜变量空间的概率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819342A_ABST
    Figure CN122819342A_ABST
Patent Text Reader

Abstract

The application discloses a diffusion encoder representation training algorithm based on timing law perception, and relates to the field of artificial intelligence. The application sets a stable reference target generation branch and a disturbance state representation inference branch: the former encodes a light disturbance reference view and provides a stable reference target across iterations through an exponential moving average update; the latter encodes a disturbed state constructed according to diffusion time steps to recover global semantic and local structure representations. The application adaptively allocates local structure constraint strength according to the signal-to-noise ratio of the diffusion state, and performs end-to-end optimization combined with a global matching loss, a local structure loss, a diffusion prediction loss and an optional prototype matching loss. Thus, the inference encoder reserved after training has both discriminative semantic ability and generative structure preservation ability, and is suitable for image pre-training, low-label recognition, content retrieval and image editing and reconstruction scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and specifically to a diffusion encoder representation training algorithm based on temporal pattern perception. Background Technology

[0002] In the field of visual representation learning, various techniques exist, including self-supervised learning, generative pre-training, and diffusion model pre-training. While consistency-based approaches typically obtain discriminative features through view matching of source samples, their training objectives do not directly constrain the image reconstruction or generation process. Variational autoencoders (VAEs), generative adversarial networks (GANs), and diffusion autoencoders (DiffAEs) can recover image details, but their latent variables may simultaneously include unstable factors such as color, texture, local perturbations, and differences in sampling trajectories. Therefore, in scenarios requiring simultaneous support for image understanding and image generation, existing solutions still suffer from insufficient functional synergy.

[0003] Some existing diffusion representation schemes employ a separate structure for the auxiliary semantic encoder and the diffusion generation path. While this structure can mitigate the impact of noise chains on semantic vectors to some extent, the semantic vectors are not fully involved in the parameterization of endpoints or reverse processes, resulting in a lack of stable correspondence between understood and generated variables. In reconstruction tasks, detailed information often requires additional path compensation; in downstream recognition tasks, the information learned in the generation path is difficult to directly utilize, thus limiting the unification of understanding and generation capabilities.

[0004] Another existing approach directly uses latent variables as conditions, boundary, or endpoint parameters in the diffusion process to improve the coupling between latent variables and the reconstruction result. While this approach helps maintain reconstruction completeness, when relying primarily on pixel-level reconstruction errors, the encoder may incorporate local textures, endpoint perturbations, and random noise into the latent variables. Since the observation space dimension is typically significantly higher than the semantic structure dimension, error signals related to non-semantic directions tend to consume a large amount of optimization resources during training, causing the low-dimensional bottleneck to gradually become burdened with redundant information.

[0005] Existing technologies suffer from a lack of coordination between data augmentation strategies, target weights, and the diffusion timestep. A fixed augmentation intensity fails to reflect the signal changes from near-data regions to near-noise regions during diffusion: in low-noise phases, excessive augmentation can disrupt reconstructable structures; in high-noise phases, insufficient augmentation may force the model to rely on local texture shortcuts for matching. Therefore, a training mechanism is needed that can adaptively configure perturbation intensity, state reliability weights, and representation scale based on the diffusion timestep.

[0006] Furthermore, existing diffusion-based representation schemes typically lack a clear coordination mechanism between global semantics and local structure. Global features are more suitable for attribute recognition, retrieval, and classification tasks, while local structural features are more important for reconstruction, editing, and fine-grained analysis. If only global features are constrained, the model may lose spatially sensitive information; if local structure is always strongly constrained, the model may fit unreliable perturbations during high-noise phases. Existing empirical weights are difficult to stably adapt to changes in the diffusion state.

[0007] Therefore, it is necessary to provide a novel diffusion encoder training algorithm that reduces the probability of non-semantic perturbations occupying latent variables while maintaining the synergy between the understanding and generation links, and preserves structural information useful for reconstruction. This invention achieves synergy between diffusion-based generation capabilities and discriminative semantic capabilities through stable reference target generation, perturbation state representation inference, diffusion time-step perturbation control, a Global-Local Collaborative Representation Head, a State-Reliability Matching Objective, and an optional Prototype Constraint Unit. Summary of the Invention

[0008] To overcome the aforementioned shortcomings, this invention does not primarily rely on simply increasing the number of network layers or stacking additional supervision. Instead, it uses the temporal position within the diffusion process as a training control variable. Specifically, during the training phase, this invention establishes a Stable Reference Target Generation Branch and a Perturbed-State Representation Inference Branch: the former forms a stable reference target across iterations from a slightly perturbed reference view, while the latter recovers the corresponding representation from the perturbed state that changes with time steps. Furthermore, the Diffusion-State Reliability Weight is used to emphasize structural preservation in near-data regions and weaken unreliable local constraints and enhance semantic consistency in near-noise regions. This results in a post-training inference encoder that possesses both discriminative semantics and reconstructable structure.

[0009] The key improvement of this invention lies in transforming the diffusion temporal location into an executable basis for view generation, target assignment, and representation constraints. Unlike conventional contrastive learning schemes with fixed enhancement strength, and unlike diffusion autoencoders that focus solely on reconstruction error, this invention completes diffusion state view construction, dual-path encoding, global-local collaborative representation, state reliability matching, optional prototype constraints, and diffusion prediction error optimization within a unified training process. Thus, the encoder reduces the occupation of the latent variable space by non-semantic redundancy while preserving structural details beneficial to reconstruction and editing.

[0010] This invention can be applied to scenarios such as image pre-training of basic models, low-annotation visual recognition, unsupervised face attribute understanding, image content retrieval, image editing and reconstruction, latent variable diffusion generation, and edge-side visual model deployment. Since the proposed solution can complete pre-training without relying on manual attribute labels, and the newly added components mainly consist of a Slow-Moving Reference Network, a Lightweight Projection Head, a Timestep Controller, a Reliability Target Assignment Function, and an optional Prototype Memory Unit, it possesses good engineering feasibility, replaceability, and industrial application prospects.

[0011] This invention provides a diffusion encoder representation training method based on time-series pattern awareness, comprising:

[0012] 1) Data processing flow:

[0013] 1.1) Read the raw, unlabeled image data of a training batch;

[0014] 1.2) Perform basic preprocessing on each unlabeled image data to obtain basic samples, denoted as the first sample. The basic samples are ,in, ; This represents the total size of the training batches.

[0015] 1.3) Establish a set of diffusion steps ,in The preset maximum number of diffusion steps;

[0016] The noise variance of each diffusion step is predefined, and the diffusion step is denoted as ______. The noise variance is ,in, ;

[0017] The value as The value increases unidirectionally;

[0018] Calculate the single-step signal retention coefficient and cumulative signal retention coefficient for each diffusion step:

[0019]

[0020]

[0021] in, Indicates diffusion step The single-step signal retention coefficient;

[0022] Indicates diffusion step The cumulative signal retention coefficient;

[0023] For each base sample, one diffusion step is extracted from the set of diffusion steps as the corresponding current diffusion step, where, The corresponding current diffusion step is denoted as ;

[0024] 1.4) To The corresponding slight perturbation reference view is obtained using weak enhancement. ;

[0025] The weak enhancement refers to a slight transformation that does not disrupt the main semantics and the overall spatial structure;

[0026] generate Corresponding perturbation inference view Specifically, it is expressed as:

[0027]

[0028] in, Indicates and Related perturbation operations, and The larger the value, the stronger the disturbance;

[0029] The disturbance type of the disturbance operation is a disturbance type that does not damage the identifiable subject;

[0030] 1.5) After inputting the reference encoder, the multi-level features of the reference branch are obtained. ;in Indicates a reference encoder. Indicates the parameters of the reference encoder;

[0031] Will After inputting into the multi-stage feature extractor, a unified format feature representation of the reference branch is obtained. ;

[0032] Will Global semantic reference representation is obtained by inputting the global projection head. ;

[0033] Will Input local structure projection head to obtain Local structural reference representation , ∈{1,2,……, };in, This represents the total number of local structures.

[0034] 1.6) Independently sample a standard Gaussian noise for each base sample, where, The corresponding Gaussian noise is

[0035] Based on the view inferred from the disturbance Generate the current diffusion state:

[0036]

[0037] in, for The corresponding current diffusion state;

[0038] 1.7) will After inputting the inference branch encoder, multi-level features of the inference branch are obtained. ;in, Indicates the inference branch encoder. This indicates the parameters of the inferred branch encoder;

[0039] The The input to the multi-stage feature extractor yields a unified format feature representation of the inference branches. ;

[0040] Will Global semantic inference representation is obtained by inputting the global projection head. ;

[0041] Will Input local structure projection head to obtain Local structure inference characterization , ∈{1,2,……, };

[0042] Will Input the diffuse noise prediction head to obtain the noise prediction results. ;

[0043] The inference branch encoder has the same network structure as the reference encoder;

[0044] 2) Parameter update method:

[0045] Based on the joint loss function, By exponential moving average Update the other parameters using backpropagation until the training termination condition is met;

[0046] Among them By exponential moving average The specific method for updating is as follows:

[0047]

[0048] in, Indicates the updated ;

[0049] Indicates the state before the update ;

[0050] Let be the momentum coefficient, satisfying

[0051] The joint loss function is:

[0052]

[0053] in, Represents the joint loss function;

[0054] This represents the global matching loss;

[0055] Indicates local structural loss;

[0056] Indicates the diffusion prediction loss;

[0057] Indicates prototype matching loss;

[0058] These are the weights corresponding to the loss;

[0059] The calculation method is as follows:

[0060]

[0061] in, For temperature parameters;

[0062] The normalized similarity function is defined as follows:

[0063]

[0064] The representation vector to be compared;

[0065] The calculation method is as follows:

[0066]

[0067] in Let be the state reliability function, and as... The value increases and decreases monotonically;

[0068] The calculation method is as follows:

[0069]

[0070] The calculation method is as follows:

[0071]

[0072] in, Indicates the preset first One prototype vector; The total number of prototype vectors; each prototype vector represents a typical semantic center or structural center in the latent space;

[0073] Indicates the first The prototype vector index number corresponding to each basic sample; The calculation method is as follows:

[0074]

[0075] Indicates taking the maximum value The value;

[0076] This represents the prototype temperature parameter.

[0077] Preferably, the basic preprocessing includes scale normalization and random cropping.

[0078] Preferably, the method for predefining the noise variance for each diffusion step is at least one of linear scheduling, cosine scheduling, piecewise scheduling, or learnable scheduling.

[0079] Preferably, the weak enhancement includes at least one of slight cropping, horizontal flipping, slight color perturbation, mild blurring, or brightness perturbation.

[0080] Preferably, the perturbation type includes at least one of clipping perturbation, color perturbation, and occlusion perturbation.

[0081] Preferred, The definition of is:

[0082]

[0083] in, For the Sigmoid function;

[0084] The signal-to-noise ratio function is defined as follows:

[0085]

[0086] These are either preset parameters or learnable parameters, and a > 0;

[0087] This is the preset zero correction amount.

[0088] The present invention also provides a deep learning model, which is trained using the above-described training method.

[0089] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described training method.

[0090] Beneficial effects:

[0091] 1) Beneficial effects of the stable reference target generation branch: This branch takes a slightly perturbed reference view as input, and after processing by a reference encoder, a multi-stage feature extractor, a global projector head, and a local structure projector head, it outputs a global semantic reference representation and a local structure reference representation by stopping gradient operations. Since the parameters of this branch do not directly receive the backpropagation gradient of the current batch loss, but are slowly updated by the inference branch parameters through exponential moving average, it can provide a stable reference target across iterations, avoiding the drastic fluctuations in the target representation caused by random noise, strong enhancement perturbations, or anomalous samples in a single batch, thereby reducing the probability of non-semantic perturbations occupying the latent variable space.

[0092] 2) Beneficial effects of the diffusion time step synchronous distribution mechanism: The current diffusion step is simultaneously distributed to the dual-view generation module, the diffusion state mixer, and the state reliability allocation function, so that the disturbance intensity, diffusion state mixing ratio, and local structure reliability weight of the lightly disturbed reference view and the disturbed inferred view are all controlled by the same time variable; thus, at low time steps, the disturbed state is close to the basic sample, and the system automatically enhances the local structure constraints; at high time steps, the disturbed state is close to the noise distribution, and the system automatically weakens unreliable local constraints and enhances global semantic consistency, realizing the adaptive linkage between disturbance intensity and diffusion state.

[0093] 3) Beneficial effects of global-local collaborative representation heads: The global projection head maps a unified feature representation into a fixed-dimensional global semantic vector through global average pooling and a multilayer perceptron, while the local structure projection head maps the same unified feature representation into K region-level local structure vectors through region / attention pooling and a multilayer perceptron. The two representations originate from the same backbone feature and are suitable for understanding tasks such as attribute recognition, retrieval, and classification, as well as generation tasks such as reconstruction and editing, respectively. This allows the inference encoder retained after training to balance discriminative semantics and reconstructable structure within a limited number of latent variable dimensions.

[0094] 4) Beneficial effects of matching state reliability targets: The local structure reliability weight is calculated based on the current diffusion state signal-to-noise ratio, and the state reliability weight is obtained through Sigmoid function transformation. When the signal-to-noise ratio is high, the local structure reliability weight is large, and the local structure loss has a strong constraint on the local structure projection head. When the signal-to-noise ratio is low, the local structure reliability weight is small, and the local structure constraint is automatically weakened. This avoids the fitting noise problem caused by strong constraints on unreliable local textures in the high noise stage, while preserving the ability to maintain spatial structure details in the low noise stage.

[0095] 5) Beneficial effects of the joint loss function: The joint loss function weights and sums the global matching loss, local structure loss, diffusion prediction loss, and optional prototype matching loss, enabling the inference branch to simultaneously learn semantic discriminativeness, local structure preservation, and diffusion state modeling capabilities. Among them, the diffusion prediction loss constrains the encoder to retain the ability to predict diffusion noise, ensuring that the latent variables not only have semantic discriminativeness but also support subsequent image reconstruction, editing, and generation tasks. The optional prototype matching loss further promotes the formation of compact clusters around the semantic center in the latent space, improving the separability of the latent space.

[0096] 6) The beneficial effects of the reference branch EMA maintenance and removability after training: The reference branch parameters are updated only through exponential moving average, and only a limited additional computational overhead is introduced during the training phase. After training, the reference branch can be completely removed, leaving only the inference branch and the required output head for downstream deployment. Therefore, this invention improves training stability without increasing the computational cost and model parameter quantity during the inference phase, and has good engineering feasibility and end-side deployment adaptability. Attached Figure Description

[0097] Figure 1 A plot of average precision for semantic attributes in FFHQ linear probing;

[0098] Figure 2 Ablation comparison chart for the objective function;

[0099] Figure 3 To stabilize the reference branch effectiveness ablation map;

[0100] Figure 4 A Pareto comparison plot of decoupling performance and product quality;

[0101] Figure 5 Figure showing the robustness results of controlled high-frequency disturbances;

[0102] Figure 6 Sensitivity analysis plot for key hyperparameters;

[0103] Figure 7 Comparison chart of the quality of prior modeling for latent variables;

[0104] Figure 8 For cost comparison chart;

[0105] Figure 9 A comparison chart of generalization metrics for non-face data;

[0106] Figure 10 A comparison chart of annotation efficiency for a small number of samples; Detailed Implementation

[0107] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0108] This embodiment of a diffusion encoder representation training method based on time-series pattern awareness includes:

[0109] 1) Data processing flow:

[0110] 1.1) Read the raw, unlabeled image data of a training batch;

[0111] 1.2) Perform basic preprocessing on each unlabeled image data to obtain basic samples, denoted as the first sample. The basic samples are ,in, ; This represents the total size of the training batches.

[0112] There are many basic preprocessing techniques in this field, which are used to standardize and normalize the raw data while preserving the main semantics and spatial structure of the original image.

[0113] In this embodiment, the basic preprocessing includes scale normalization and random cropping. Scale normalization is used to scale images of different sizes to a uniform resolution, facilitating batch parallel computation; random cropping is used to increase spatial location and scale diversity while preserving the semantics of the subject.

[0114] 1.3) Establish a set of diffusion steps ,in The preset maximum number of diffusion steps;

[0115] The noise variance of each diffusion step is predefined, and the diffusion step is denoted as ______. The noise variance is ,in, ;

[0116] The value as The value increases unidirectionally; in this embodiment, the method for predefining the noise variance of each diffusion step is at least one of linear scheduling, cosine scheduling, piecewise scheduling, or learnable scheduling.

[0117] Calculate the single-step signal retention coefficient and cumulative signal retention coefficient for each diffusion step:

[0118]

[0119]

[0120] in, Indicates diffusion step The single-step signal retention coefficient;

[0121] Indicates diffusion step The cumulative signal retention coefficient;

[0122] For each base sample, one diffusion step is extracted from the set of diffusion steps as the corresponding current diffusion step, where, The corresponding current diffusion step is denoted as ;

[0123] 1.4) To The corresponding slight perturbation reference view is obtained using weak enhancement. ;

[0124] The weak enhancement refers to a slight transformation that does not destroy the main semantics and overall spatial structure; in this embodiment, the weak enhancement includes at least one of slight cropping, horizontal flipping, slight color perturbation, slight blurring or brightness perturbation;

[0125] generate Corresponding perturbation inference view Specifically, it is expressed as:

[0126]

[0127] in, Indicates and Related perturbation operations, and The larger the value, the stronger the disturbance;

[0128] The perturbation type of the perturbation operation is a perturbation type that does not damage the identifiable subject. In this embodiment, the perturbation type includes at least one of cropping perturbation, color perturbation, and occlusion perturbation.

[0129] Both views originate from the same base sample. Therefore, they share the same semantic origin; however, their perturbation levels differ, allowing them to be used to construct homologous positive sample relationships and for diffusion state recovery tasks. When smaller, and The differences are small, and the system emphasizes local structural consistency; when When it is large, In subsequent diffusion state mixing, with the superposition of more noise, the system reduces local strong constraints and enhances its dependence on stable reference semantics. This ensures both the continuity of data sources for the two branches and the consistency of disturbance intensity with diffusion time and location.

[0130] 1.5) After inputting the reference encoder, the multi-level features of the reference branch are obtained. ;in Indicates a reference encoder. Indicates the parameters of the reference encoder;

[0131] Will After inputting into the multi-stage feature extractor, a unified format feature representation of the reference branch is obtained. ;

[0132] Will Global semantic reference representation is obtained by inputting the global projection head. ;

[0133] Will Input local structure projection head to obtain Local structural reference representation , ∈{1,2,……, };in, This represents the total number of local structures.

[0134] The reference encoder is used to provide basic features for subsequent multi-stage feature extraction; in this embodiment, the reference encoder is implemented by Patch Embedding (input Stem) and Transformer encoder.

[0135] The parameters of the reference encoder Instead of directly receiving the backpropagation gradient of the current batch loss, it updates the inferred branch parameters through exponential moving average.

[0136] The multi-stage feature extractor is used to extract feature representations in a uniform format. In this embodiment, the multi-stage feature extractor is implemented by multi-stage feature aggregation, feature pyramid fusion, and normalization, and the multi-stage feature aggregation and feature pyramid fusion can be turned off as needed.

[0137] The function of the global projection head is to... Mapped to a fixed-dimensional global semantic vector It is used to characterize the subject category, identity, attributes or overall semantic content of an image; in this embodiment, the global projection head is implemented in the form of GAP (Global Average Pooling) or MLP (Projection Head).

[0138] The function of the local structure projection head is to represent features in a uniform format. Mapped to Local structural reference representation , ∈{1,2,……, This is used to characterize reconstructable structures such as edge contours, textures, and spatial layouts. In this embodiment, the local structure projection head includes: RoI / Attention Pool and MLP (projection head); the RoI / Attention Pool specifically divides the feature map into K spatial regions, pools each region, and obtains K region-level feature vectors; the MLP (projection head) specifically maps each of the K region-level feature vectors and outputs... Local structural reference representation , ∈{1,2,……, }

[0139] 1.6) Independently sample a standard Gaussian noise for each base sample, where, The corresponding Gaussian noise is

[0140] Based on the view inferred from the disturbance Generate the current diffusion state:

[0141]

[0142] in, for The corresponding current diffusion state;

[0143] This formula clearly distinguishes between signal and noise components, enabling the encoder to observe continuous changes from near-data state to near-noise state. When When smaller, Approaching 1, It retains relatively clear edges, textures, and spatial layouts, thus making local structural targets highly reliable; when When it is large, When the value approaches 0, the proportion of noise components increases, and the reliability of local pixels and high-frequency textures decreases. Therefore, the reliability weight of subsequent states will automatically weaken the local structural constraints.

[0144] 1.7) will After inputting the inference branch encoder, multi-level features of the inference branch are obtained. ;in, Indicates the inference branch encoder. This indicates the parameters of the inferred branch encoder;

[0145] The The input to the multi-stage feature extractor yields a unified format feature representation of the inference branches. ;

[0146] Will Global semantic inference representation is obtained by inputting the global projection head. ;

[0147] Will Input local structure projection head to obtain Local structure inference characterization , ∈{1,2,……, };

[0148] Will Input the diffuse noise prediction head to obtain the noise prediction results. ;

[0149] The inference branch encoder is used to recover semantic and structural representations from the perturbed diffusion state. In this embodiment, the inference branch encoder is implemented by Patch Embedding (input Stem) and Transformer encoder, and the inference branch encoder has the same network structure as the reference encoder.

[0150] The diffusion noise prediction head is used to predict the Gaussian noise component contained in the current diffusion state; in this embodiment, the diffusion noise prediction head is implemented by a multi-layer convolution or Transformer structure;

[0151] The parameters of the inference branch encoder directly receive the gradient of the joint loss through backpropagation and are updated using AdamW or an equivalent adaptive optimizer.

[0152] 2) Parameter update method:

[0153] Based on the joint loss function, By exponential moving average (EMA) Update the other parameters using backpropagation until the training termination condition is met;

[0154] Among them By exponential moving average The specific method for updating is as follows:

[0155]

[0156] in, Indicates the updated ;

[0157] Indicates the state before the update ;

[0158] Let be the momentum coefficient, satisfying ;

[0159] The closer the value is to 1, the slower the reference branch updates and the more stable the reference target. Since the reference branch is not directly affected by the gradient of random noise, strong perturbations, or anomalous samples in the current batch, it can provide a smooth target representation across iterations.

[0160] The joint loss function is:

[0161]

[0162] in, Represents the joint loss function;

[0163] This represents the global matching loss;

[0164] Indicates local structural loss;

[0165] Indicates the diffusion prediction loss;

[0166] Indicates prototype matching loss;

[0167] These are the weights corresponding to the loss;

[0168] The weight parameters in the above joint objective can be set according to application requirements: when the task is more focused on attribute recognition, retrieval, or classification, the weight can be increased. and When the task is more focused on rebuilding, editing, or generating, it can improve performance. and When a task requires balancing understanding and generation, a more balanced weight configuration is used; alternatively, the corresponding loss can be turned off by setting the weight to 0, for example... It can be set to 0 to disable prototype matching loss. This design allows the method of the present invention to be adapted to different types of visual pre-training and downstream application scenarios.

[0169] The calculation method is as follows:

[0170]

[0171] in, For temperature parameters;

[0172] The normalized similarity function is defined as follows:

[0173]

[0174] The representation vector to be compared;

[0175] The calculation method is as follows:

[0176]

[0177] in Let be the state reliability function, and as... The value increases and decreases monotonically;

[0178] As an optional implementation method, The definition of is:

[0179]

[0180] in, For the Sigmoid function;

[0181] The signal-to-noise ratio function is defined as follows:

[0182]

[0183] These are either preset parameters or learnable parameters, and a > 0;

[0184] This is a preset zero-correction amount, typically a small positive number;

[0185] The calculation method is as follows:

[0186]

[0187] The calculation method is as follows:

[0188]

[0189] in, Indicates the preset first One prototype vector; The total number of prototype vectors; each prototype vector represents a typical semantic center or structural center in the latent space; this module does not rely on manual labels and spontaneously organizes the semantic clustering structure in the latent space through prototype vectors.

[0190] Indicates the first The prototype vector index number corresponding to each basic sample; The calculation method is as follows:

[0191]

[0192] Indicates taking the maximum value The value; that is It equals the prototype vector index number that maximizes the similarity;

[0193] Indicates the prototype temperature parameter;

[0194] After training, save the required modules as needed. Specifically, during the deployment phase, the stable reference target generation branch can be removed, retaining only the inference encoder and the required output head. If the downstream task is attribute recognition, classification, or retrieval, a global semantic representation can be output; if the downstream task is reconstruction, editing, or fine-grained analysis, both global semantic representation and local structural representation can be output simultaneously; if the downstream task is latent variable generation, it can be further connected to a latent variable prior model or a diffusion decoder.

[0195] Related comparison explanation:

[0196] (1) Semantic attributes and objective function ablation

[0197] This embodiment verifies the effectiveness of the method of the present invention under an unlabeled image pre-training setting. To avoid confusion, the proposed algorithm is referred to as "the method of the present invention" in the following experimental figures and tables, and different implementations of the bridging baseline are distinguished by "bridging diffusion encoder (deterministic)" and "bridging diffusion encoder (random)".

[0198] Figure 1 Demonstrates the average precision (AP) of semantic attributes for FFHQ linear probing. Figure 2 Further comparisons were made between the Normalized Mean Squared Error (NMSE) target and the matching target of this invention on multiple metrics, including FFHQ and CelebA. The results show that the method of this invention has better overall performance in terms of semantic density and attribute relevance.

[0199] Table 1 Ablation Results of Semantic Attributes and Objective Function

[0200]

[0201] (2) Stable reference branch and reconstruction quality

[0202] Figure 3 This demonstrates the crucial role of stable reference branches in preventing latent space degradation. Even after 500k training steps, the AP, latent variable variance, and SVD entropy remain significantly low in the no-reference setting; the method of this invention significantly improves latent variable dispersion and semantic density by using stable reference targets. This method outperforms LPIPS and SSIM, demonstrating a balance between perceptual quality and structure preservation.

[0203] Table 2 Reconstruction quality, generation quality, and stability results

[0204]

[0205] (3) Disentanglement, Robustness and Hyperparameter Analysis

[0206] Figure 4 This demonstrates the relationship between decoupling and generation quality from a Pareto perspective. Figure 5 Demonstrates the stability of latent variables under controlled high-frequency perturbations. Figure 6 The results demonstrate the sensitivity of key hyperparameters, showing that the default configuration is in an optimal range and that small-scale perturbations do not cause serious degradation.

[0207] (4) Latent variable priors, computational cost and generalization ability

[0208] Figure 7 Comparing FID under the same lightweight latent variable prior networks; Figure 8 Compare the number of parameters, training time, and GPU memory usage. Figure 9 Display the generalization results of non-face data; Figure 10 This demonstrates sample efficiency at a 1% annotation ratio. Overall, the method of this invention significantly improves semantic density and latent space modelability while incurring only limited computational overhead.

[0209] Table 5. Latent Variable Priors, Computational Costs, and Generalization Results

[0210]

[0211] In summary, the method of this invention, within a unified diffusion coding framework, achieves consistent advantages in semantic attribute prediction, perceptual reconstruction, decoupled generation, perturbation robustness, latent variable prior modeling, and cross-dataset generalization through time-step-driven perturbation construction, stable reference targets, global-local collaborative representation, and reliability control. The above embodiments are merely illustrative of preferred implementations of this invention and do not constitute limitations on network backbone, noise scheduling, data types, specific forms of loss functions, hardware platforms, or deployment methods.

[0212] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A representation training method for a diffusion encoder based on temporal pattern perception, characterized in that, include: 1) Data processing flow: 1.1) Read the raw, unlabeled image data of a training batch; 1.2) Perform basic preprocessing on each unlabeled image data to obtain basic samples, denoted as the first sample. The basic samples are ,in, ; This represents the total size of the training batches. 1.3) Establish a set of diffusion steps ,in The preset maximum number of diffusion steps; The noise variance of each diffusion step is predefined, and the diffusion step is denoted as ______. The noise variance is ,in, ; The value as The value increases unidirectionally; Calculate the single-step signal retention coefficient and cumulative signal retention coefficient for each diffusion step: in, Indicates diffusion step The single-step signal retention coefficient; Indicates diffusion step The cumulative signal retention coefficient; For each base sample, one diffusion step is extracted from the set of diffusion steps as the corresponding current diffusion step, where, The corresponding current diffusion step is denoted as ; 1.4) To The corresponding slight perturbation reference view is obtained using weak enhancement. ; The weak enhancement refers to a slight transformation that does not disrupt the main semantics and the overall spatial structure; generate Corresponding perturbation inference view Specifically, it is expressed as: in, Indicates and Related perturbation operations, and The larger the value, the stronger the disturbance; The disturbance type of the disturbance operation is a disturbance type that does not damage the identifiable subject; 1.5) After inputting the reference encoder, the multi-level features of the reference branch are obtained. ;in Indicates a reference encoder. Indicates the parameters of the reference encoder; Will After inputting into the multi-stage feature extractor, a unified format feature representation of the reference branch is obtained. ; Will Global semantic reference representation is obtained by inputting the global projection head. ; Will Input local structure projection head to obtain Local structural reference representation , ∈{1,2,……, };in, This represents the total number of local structures. 1.6) Independently sample a standard Gaussian noise for each base sample, where, The corresponding Gaussian noise is Based on the view inferred from the disturbance Generate the current diffusion state: in, for The corresponding current diffusion state; 1.7) will After inputting the inference branch encoder, multi-level features of the inference branch are obtained. ;in, Indicates the inference branch encoder. This indicates the parameters of the inferred branch encoder; The The input to the multi-stage feature extractor yields a unified format feature representation of the inference branches. ; Will Global semantic inference representation is obtained by inputting the global projection head. ; Will Input local structure projection head to obtain Local structure inference characterization , ∈{1,2,……, }; Will Input the diffuse noise prediction head to obtain the noise prediction results. ; The inference branch encoder has the same network structure as the reference encoder; 2) Parameter update method: Based on the joint loss function, By exponential moving average Update the other parameters using backpropagation until the training termination condition is met; Among them By exponential moving average The specific method for updating is as follows: in, Indicates the updated ; Indicates the state before the update ; Let be the momentum coefficient, satisfying ; The joint loss function is: in, Represents the joint loss function; Represents the global matching loss; Indicates local structural loss; Indicates the diffusion prediction loss; Indicates prototype matching loss; These are the weights corresponding to the loss; The calculation method is as follows: in, For temperature parameters; The normalized similarity function is defined as follows: The representation vectors to be compared; The calculation method is as follows: in Let be the state reliability function, and as... The value increases and decreases monotonically; The calculation method is as follows: The calculation method is as follows: in, Indicates the preset first One prototype vector; The total number of prototype vectors; each prototype vector represents a typical semantic center or structural center in the latent space; Indicates the first The prototype vector index number corresponding to each basic sample; The calculation method is as follows: Indicates taking the maximum value The value; This represents the prototype temperature parameter.

2. The diffusion encoder representation training method based on time-series pattern perception according to claim 1, characterized in that, The basic preprocessing includes scale normalization and random cropping.

3. The diffusion encoder representation training method based on time-series pattern perception according to claim 1, characterized in that, The method for predefining the noise variance for each diffusion step is at least one of linear scheduling, cosine scheduling, piecewise scheduling, or learnable scheduling.

4. The diffusion encoder representation training method based on time-series pattern perception according to claim 1, characterized in that, The weak enhancement includes at least one of slight cropping, horizontal flipping, slight color perturbation, mild blurring, or brightness perturbation.

5. The diffusion encoder representation training method based on time-series pattern perception according to claim 1, characterized in that, The perturbation type includes at least one of clipping perturbation, color perturbation, and occlusion perturbation.

6. The diffusion encoder representation training method based on time-series pattern perception according to claim 1, characterized in that, The definition of is: in, For the Sigmoid function; The signal-to-noise ratio function is defined as follows: These are either preset parameters or learnable parameters, and >0; This is the preset zero correction amount.

7. A deep learning model, characterized in that, It is trained using any one of the training methods of claims 1 to 6.

8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1 to 6.