A long sequence video reasoning method based on a variational energy framework and cognitive entropy constraint

CN122737814APending Publication Date: 2026-09-11HUNAN KUNLUNYUAN ARTIFICIAL INTELLIGENCE APPLICATION SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610829999.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

现有推理框架由特征提取、时序建模与语言推理模块拼接而成,各模块独立训练优化,缺乏跨模块协同约束,导致视频证据表征、时序隐变量与推理步骤无法精准对齐,出现特征传递断裂、视觉证据与语言推理脱节等问题,推理过程难以被前端视频表征有效约束

Benefits of technology

基于变分能量模型将视频特征提取、时序建模与多步骤推理全过程纳入统一的能量优化理论框架,利用能量函数凸性分析严格证明系统收敛性与稳定性,从根本上解决了传统模块化架构中目标割裂、特征对齐困难及推理约束失效的问题。其次,提出的对偶-谱混合注意力机制依托不确定性原理实现空域细节与频域全局结构注意力的自适应融合,在ImageNet上Top-1准确率提升4.1%,计算量仅为传统Softmax全局注意力的18%,达成了精度与效率的帕累托最优。同时,拓扑自适应专家网络层引入持久同调理论分析特征空间拓扑结构,动态构建专家激活路由,使匹配精度提升37%,显著增强了模型特征路由的合理性与可解释性。耗散型可学习残差连接将网络层间信息流建模为梯度流动力学系统,通过严格的能量耗散条件保障训练稳定性,在100层网络中梯度范数稳定性提升2.3倍,收敛速度加快40%。在推理层面,贝叶斯思维链将视频多步推理建模为隐变量后验推断序列,为每步赋予认知一致性得分(与答案正确率相关系数0.81),并结合信息瓶颈约束抑制推理幻觉,使推理准确率提升12.7%、幻觉率降低58%。在主流基准测试中,本发明在MLVU长视频理解上达到78.9%准确率(较最优方法提升6.6%),在TGIF-QA动态推理上达到63.4%(提升4.7%),处理1小时超长视频的延迟仅为对比模型的34%,兼具高精度、高稳定性与低算力成本,具备极高的实际落地应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122737814A_ABST
    Figure CN122737814A_ABST
Patent Text Reader

Abstract

The application provides a long-sequence video inference method based on a variational energy framework and cognitive entropy constraint, relates to the technical field of computer vision and multi-modal video inference, and comprises long-sequence image coding based on a variational energy constraint, two-stage contrastive learning and cognitive entropy knowledge distillation training, spatio-temporal variational inference video time sequence modeling and evidence uncertainty quantification, and Bayesian thought chain cognitive consistency controllable inference. The application can construct a long-sequence visual feature encoder through a topological adaptive expert network, a dual-spectrum hybrid attention mechanism and a dissipative learnable residual connection, and realize posterior modeling, uncertainty estimation and cognitive consistency inference on long-sequence video evidence by combining cognitive entropy minimization training, spatio-temporal variational inference and Bayesian thought chain inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and multimodal video reasoning technology, and in particular to a long sequence video reasoning method based on variational energy framework and cognitive entropy constraints. Background Technology

[0002] With the rapid development of multimedia technology, video has become a core information carrier on the internet, widely used in artificial intelligence fields such as autonomous driving, intelligent monitoring, video question answering, and event tracing. Compared to short videos, long video sequences contain continuous spatiotemporal dynamic information, cross-segment event correlations, and complex hidden semantic clues. Intelligent understanding of these sequences requires not only efficient spatiotemporal feature encoding but also accurate modeling of long-term temporal dependencies and event evolution patterns, and high-precision multi-step reasoning based on video evidence. Therefore, how to balance encoding efficiency, temporal expressiveness, reasoning accuracy, and result interpretability in long video sequences has become a key technical challenge in computer vision and multimodal intelligence, hindering performance improvements in related applications.

[0003] Current long-sequence video inference technology suffers from significant and fundamental flaws. First, the system modules are fragmented and lack unified constraints. Existing inference frameworks are composed of feature extraction, temporal modeling, and language inference modules, each trained and optimized independently, lacking cross-module collaborative constraints. This leads to inaccurate alignment between video evidence representation, temporal latent variables, and inference steps, resulting in problems such as broken feature transfer and disconnect between visual evidence and language inference. The inference process is difficult to be effectively constrained by the front-end video representation. Second, it fails to balance computational efficiency and fine-grained representation. The computational complexity of Transformer-based models increases quadratically with the sequence length, resulting in extremely high computational costs for long-video inference. While existing optimization techniques such as frame sampling, feature compression, and linear attention can reduce computational costs, they easily lose detailed features and key temporal information. Furthermore, hybrid attention schemes are merely simple branch combinations, lacking adaptive fusion mechanisms, making it difficult to balance efficiency and modeling accuracy. Finally, the inference process lacks uncertainty quantification and consistency control. Existing models often directly generate inference results, and thought chain reasoning relies excessively on language models. They cannot quantify the credibility of inference support from video evidence and find it difficult to distinguish between valid deductions and model guesses. At the same time, they are easily affected by language priors, resulting in unfounded inference results. Furthermore, the lack of inference path filtering and error correction mechanisms leads to the continuous accumulation of single-step errors, significantly reducing the accuracy of inference.

[0004] Current industry solutions are mostly localized optimizations, improving performance from single dimensions such as temporal memory enhancement, video token compression, and language reasoning optimization. However, they have not yet established a comprehensive collaborative mechanism across the entire chain, encompassing video evidence encoding, uncertainty estimation, likelihood calculation of inference evidence, and inference path control. This prevents unified constraint optimization across multiple modules and fails to address the root causes of existing technological deficiencies. Therefore, the industry urgently needs to build an integrated long-sequence video reasoning framework. This framework should enable stable encoding and posterior modeling of long-sequence video evidence within a unified system. It should leverage video evidence uncertainty constraint reasoning generation, consistency assessment, and adaptive path control to overcome existing technological bottlenecks and improve the accuracy, reliability, and interpretability of multi-step reasoning in long videos. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a long sequence video reasoning method based on variational energy framework and cognitive entropy constraints. It can realize the construction of a long sequence visual feature encoder by using a topological adaptive expert network, a dual-spectral hybrid attention mechanism and dissipative learnable residual connections, and combined with cognitive entropy minimization training, spatiotemporal variational inference and Bayesian thought chain reasoning, to achieve posterior modeling, uncertainty estimation and cognitive consistency reasoning of long sequence video evidence.

[0006] This invention provides a long-sequence video inference method based on a variational energy framework and cognitive entropy constraints. The method includes: S1: Constructing a long-sequence image encoder based on a variational energy model. The encoder includes a topologically adaptive expert network layer (TAoE), a dual-spectral hybrid attention mechanism (DSHA), and a dissipative learnable residual connection (DLRC), used to obtain structure-aware, spatial-frequency domain fusion, and stable propagation visual feature representations under unified energy constraints; wherein, the variational energy model is used to impose energy constraints on the visual feature encoding process, and its energy functional representation is: in, Represented as latent variables, For input data, For model parameters, Reconstructing energy from data As a priori constraint energy, The energy is used for parameter regularization. The inference process of the model corresponds to energy minimization: The training process corresponds to parameter optimization: The Topology Adaptive Expert Network (TAoE) layer analyzes the topological structure of the input features through persistent cohomology analysis and dynamically constructs an expert activation graph based on the topological structure. , where the vertex To gather experts, on the side The weights are determined by the topological connectivity of the feature space, used to implement expert routing related to the input video structure; the dual-spectral hybrid attention mechanism DSHA calculates the attention response in the original feature space and frequency domain space respectively, and adaptively fuses them according to the uncertainty of the two types of attention responses, to take into account both local spatial details and global frequency domain structure; the dissipative learnable residual connection DLRC The information flow between network layers is modeled as a gradient flow dynamics process, and the inter-layer state update is controlled by dissipation constraints to improve the stability of the deep visual feature propagation process; S2: Based on the cognitive entropy minimization objective, the long sequence image encoder is trained by two-stage contrastive learning and knowledge distillation; wherein, the two-stage contrastive learning is used to enhance the semantic alignment ability and multi-scale robustness of visual feature representation, and the knowledge distillation is used to enable the student encoder to learn the output distribution, feature representation, and cognitive entropy representation of the teacher model; S3: The video encoder is initialized using the trained long sequence image encoder parameters, and temporal modeling is performed through spatiotemporal variational inference; wherein, the input long sequence video is represented as temporal data composed of multiple video segments or frame sequences, the posterior distribution of latent variables at the video segment level is obtained based on the video encoder, and video evidence representation and video evidence uncertainty are constructed based on the posterior distribution of latent variables; S4: A video inference model based on Bayesian thought chain and cognitive entropy constraints is constructed. Specifically, the video reasoning process is modeled as a latent variable posterior inference sequence consisting of multiple reasoning steps. When generating each reasoning step, the posterior probability of that reasoning step is determined based on the linguistic prior and the likelihood of the video evidence. A cognitive consistency score is calculated based on the uncertainty of the video evidence and the degree of matching between the reasoning step and the video evidence. When the cognitive consistency score is lower than a preset threshold, rejection sampling, regeneration, or supplementary video evidence retrieval is performed on the corresponding reasoning step or reasoning path. When the uncertainty of the video evidence is higher than a preset threshold, the reasoning depth is adaptively adjusted or relevant video segments are reselected, ultimately generating a video reasoning output that includes the reasoning process and the final conclusion.

[0007] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: Based on the variational energy model, the entire process of video feature extraction, temporal modeling, and multi-step inference is incorporated into a unified energy optimization theoretical framework. The convergence and stability of the system are rigorously proven using energy function convexity analysis, fundamentally solving the problems of target fragmentation, feature alignment difficulties, and inference constraint failure in traditional modular architectures. Secondly, the proposed dual-spectral hybrid attention mechanism, relying on the uncertainty principle, achieves adaptive fusion of spatial detail and frequency domain global structure attention, improving Top-1 accuracy by 4.1% on ImageNet while requiring only 18% of the computational cost of traditional Softmax global attention, achieving Pareto optimality in both accuracy and efficiency. Simultaneously, the topology-adaptive expert network layer introduces persistent cohomology theory to analyze the feature space topology and dynamically constructs expert activation routes, improving matching accuracy by 37% and significantly enhancing the rationality and interpretability of model feature routing. Dissipative learnable residual connections model the information flow between network layers as a gradient flow dynamics system, ensuring training stability through strict energy dissipation conditions. In a 100-layer network, gradient norm stability is improved by 2.3 times, and convergence speed is accelerated by 40%. At the reasoning level, Bayesian Mind Chain models multi-step video reasoning as a latent variable posterior inference sequence, assigning a cognitive consistency score to each step (correlation coefficient with answer accuracy of 0.81), and combining information bottleneck constraints to suppress reasoning illusions, resulting in a 12.7% improvement in reasoning accuracy and a 58% reduction in the illusion rate. In mainstream benchmark tests, this invention achieves 78.9% accuracy in MLVU long video understanding (a 6.6% improvement over the best method) and 63.4% accuracy in TGIF-QA dynamic reasoning (a 4.7% improvement). The latency for processing 1-hour long videos is only 34% of that of the comparison model, combining high precision, high stability, and low computational cost, making it highly valuable for practical applications. Attached Figure Description

[0008] Figure 1 This is a flowchart of a long sequence video reasoning method based on variational energy framework and cognitive entropy constraints provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of the variational energy model framework of a long sequence video reasoning method based on variational energy framework and cognitive entropy constraint provided in an embodiment of the present invention; Figure 3 This is a topological adaptive expert network layer structure diagram of a long sequence video reasoning method based on variational energy framework and cognitive entropy constraints provided in an embodiment of the present invention; Figure 4 This is a flowchart of the dual-spectral hybrid attention mechanism calculation for a long sequence video reasoning method based on variational energy framework and cognitive entropy constraints provided in this embodiment of the invention. Figure 5This is a schematic diagram of the dynamics of a dissipative residual connection in a long-sequence video inference method based on a variational energy framework and cognitive entropy constraints, provided by an embodiment of the present invention. Figure 6 This is a diagram of the Bayesian thought chain reasoning process with cognitive entropy constraints, provided by an embodiment of the present invention, for a long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints. Detailed Implementation

[0009] This invention provides a long-sequence video reasoning method based on a variational energy framework and cognitive entropy constraints, such as... Figure 1 The flowchart shown is a long sequence video inference method based on variational energy framework and cognitive entropy constraints. The processing flow of this method can include the following steps: S1: Construct a long sequence image encoder based on a variational energy model. The encoder includes a topological adaptive expert network layer TAoE, a dual-spectral hybrid attention mechanism DSHA, and a dissipative learnable residual connection DLRC, which is used to obtain a structure-aware, spatial-frequency domain fusion and stable propagation visual feature representation under a unified energy constraint. The variational energy model is used to impose energy constraints on the visual feature encoding process, and its energy functional is expressed as: in, Represented as latent variables, For input data, For model parameters, Reconstructing energy from data As a priori constraint energy, The energy is used for parameter regularization. The inference process of the model corresponds to energy minimization: The training process corresponds to parameter optimization: The Topology Adaptive Expert Network (TAoE) layer analyzes the topological structure of the input features through persistent cohomology analysis and dynamically constructs an expert activation graph based on the topological structure. , where the vertex To gather experts, on the side The weights are determined by the topological connectivity of the feature space and are used to implement expert routing related to the structure of the input video. The dual-spectral hybrid attention mechanism DSHA calculates the attention response in the original feature space and frequency domain space respectively, and performs adaptive fusion based on the uncertainty of the two types of attention responses, in order to take into account both local spatial details and global frequency domain structure. The dissipative learnable residual connection (DLRC) models the inter-layer information flow as a gradient flow dynamics process and controls the inter-layer state update through dissipative constraints, thereby improving the stability of the deep visual feature propagation process.

[0010] It should be understood that this invention constructs a long-sequence image encoder based on the Variational Energy Model (VEM) for the stable encoding of frame-level or segment-level visual information in long-sequence videos. This encoder introduces a Topology Adaptive Expert Network (TAoE), a Dual-Spectral Hybrid Attention Mechanism (DSHA), and Dissipative Learnable Residual Connections (DLRC) within the variational energy framework. These are used to achieve topology-aware expert routing, attention fusion between the original feature space and the frequency domain space, and stable state updates in deep networks, respectively. Through this structure, the image encoder can provide visual evidence representations with structural information, spatiotemporal expressive capabilities, and uncertainty representation foundations for subsequent video temporal modeling and Bayesian thought chain inference.

[0011] It should be noted that the specific implementation of the Topology Adaptive Expert Network (TAoE) layer in step S1 includes: Constructing input features The topological complex, which adopts the Vietoris-Rips complex: in, It is the diameter of a simplex shape. These are the filtering parameters.

[0012] The k-th order persistent homology group is calculated based on the Vietoris-Rips complex. Extracting topological feature descriptors from barcodes: in, For the first Betti number of order, For the first The birth-death interval of a topological feature. Indicates the first The persistence of topological features.

[0013] Dynamically constructing expert routing matrix based on topology features , of which The expert and the first The routing weights among experts are represented as follows: in, Measurement experts and The number of connected components corresponding to the feature subspace. Measuring the persistence of a one-dimensional topological cycle For feature dimensions.

[0014] Based on the expert routing matrix Calculate the activation probability of each expert, the activation probability being determined by topological constraints. get: in, For the first The gating parameters corresponding to each expert For experts Pre-calculated activation scores, For topology routing weights, , For temperature coefficient, For the number of experts.

[0015] Select a Top-K expert set based on the activation probabilities. The output of the topology adaptive expert network layer is calculated as follows: in: For the first Transformation results of input features by an expert This is the learnable residual scaling factor.

[0016] Therefore, the Topology Adaptive Expert Network (TAoE) adjusts the expert routing weights by considering the topological connectivity of the input features and the one-dimensional persistent topological structure, so that the expert activation results correspond to the topological structure of the input features.

[0017] In this embodiment, the state of the system is defined by implicit variables. and observable variables Composition, model parameters are The energy functional of a system is defined as follows: in: Reconstructing energy from data For the decoder, it measures the ability of the latent variable representation H to reconstruct the input data X; To constrain energy a priori, the posterior distribution of latent variables is constrained using KL divergence. Approximate prior distribution To prevent overfitting; For parameter regularization energy, This corresponds to L2 regularization. This corresponds to L1 sparse regularization.

[0018] Given input data X and model parameters θ, the model inference process corresponds to minimizing the energy of the latent variable representation H: This is equivalent to solving a fixed-point equation: For deep networks, the above energy minimization process is approximated by inter-layer recursion, with each network layer corresponding to a state update with respect to the hidden variable representation: in, The learning rate is specific to a layer (corresponding to the learnable scaling factor between network layers).

[0019] The training process corresponds to the training process on the dataset. Optimization parameters : The framework's unity lies in the fact that feature extraction, temporal modeling, and inference can all be viewed as different optimization problems on the same energy landscape. Information transfer between modules corresponds to the flow of energy gradients, and the system's convergence and stability can be rigorously proven through the convexity analysis of the energy function. Traditional Mixture-of-Experts (MoE) networks typically employ fixed or input-independent expert topologies, with different input samples sharing similar routing mechanisms, making it difficult to fully utilize the potential topological differences in video features. To address this, this invention proposes a Topology-Adaptive Autonomy-of-Experts (TAoE) layer. This layer analyzes the topology of the input feature space through persistent cohomology analysis and dynamically constructs expert activation graphs based on these topological features, thereby adapting expert routing to the structural features of the input visual evidence. Given input features... Construct the Vietoris-Rips complex: With filter parameters Increment from 0 to The complex gradually grows from a discrete set of points to a complete simplex. During this process, homology groups of various orders are calculated. The evolution of topological features (connected components, cycles, holes, etc.) records their birth time. and time of death The k-th order continuous graph is represented as: The importance of topological features is measured by their persistence: Among them, topological features with high persistence are used to characterize relatively stable structural information in the input feature space, while topological features with low persistence can correspond to noise or local perturbations.

[0020] Further extract the Betti number as a topological descriptor: Betti number Indicates the number of connected components. Represents the number of one-dimensional rings. This represents the number of two-dimensional holes.

[0021] Expert routing based on topology features: Divide the input feature space into Each subspace corresponds to a specific expert. (Computational expert) and Topological similarity of corresponding subspaces: in, The Wasserstein-2 distance measures the similarity between two persistent graphs.

[0022] Routing matrix The topological relationships between experts are encoded. Based on this matrix, the expert activation probability depends not only on the local response to the input features but also on the global topological structure: in, For the first The gating parameters corresponding to each expert For the next level of experts The activation score enables cross-layer topology information propagation.

[0023] Expert selection for differentiable features using Gumbel-Softmax: The TAoE layer output is: in, This represents the set of Top-K experts obtained based on the probabilities of expert selection. Let be the transformation result of the i-th expert on the input feature x. Learnable residual scaling factor Compared with traditional MoE, TAoE has the following advantages: it captures the global topological structure of the feature space through persistent cohomology to guide expert routing; the routing decision has theoretical interpretability, and highly persistent topological features correspond to more important semantic concepts; cross-layer topological propagation enables deep experts to utilize topological information discovered in shallow layers.

[0024] It should be further explained that, in step S1, the dual-spectral hybrid attention mechanism DSHA includes: Define dual attention as primitive spatial attention With frequency domain attention The coupling involves the frequency domain attention being mapped back to the original feature space via inverse discrete Fourier transform and then fused with the original feature space attention. in, and These are the Discrete Fourier Transform and its inverse transform, respectively. For querying the matrix, For coupling functions based on uncertainty; The frequency domain attention is calculated in Fourier space: in, The key matrix, For value matrices, This represents the Hermitian transpose. The dimension of the key vector.

[0025] The original feature space attention includes a linear attention branch. and Softmax attention branch And adaptive fusion is achieved through the gating parameter λ: in, For kernel mapping function The coupling function Adaptive fusion is performed based on the uncertainties of the frequency domain branch output and the original feature space branch output. Let: The coupling function is then expressed as: in, and Frequency domain branch outputs and the original feature space branch output variance To prevent constants with zero denominators, inverse variance weighting is used to give higher weights to branches with smaller output variances in the fusion result.

[0026] The gate parameters Based on the current hidden layer state Adaptive adjustment of frequency domain energy ratio: middle, and As a learnable parameter, when the proportion of frequency domain energy is high, Increase the value to enhance the computational efficiency of the linear attention branch in long sequence modeling.

[0027] In this embodiment, existing attention mechanisms mostly compute in the original feature space, which underutilizes the frequency domain structure of the video signal. For long video sequences, the original feature space is beneficial for preserving local spatial details, while the frequency domain is beneficial for characterizing global structure, periodic changes, and long-range correlation patterns. To address this, this invention proposes a Dual-Spectral Hybrid Attention (DSHA) mechanism, which computes attention responses in both the original feature space and the frequency domain, and adaptively fuses them based on the uncertainty of the branch outputs, thereby improving the expressive power and stability of long-sequence visual feature modeling.

[0028] Frequency domain attention: Perform Discrete Fourier Transform (DFT) on the query matrix Q, key matrix K, and value matrix V respectively: Compute the attention response in the frequency domain: in, This represents the Hermitian transpose. The physical meaning of frequency domain attention is that low-frequency components correspond to global structure (such as the overall shape of an object), while high-frequency components correspond to local details (such as texture and edges). By computing attention in the frequency domain, the model can explicitly distinguish the contributions of different frequency components.

[0029] Mapping frequency domain attention back to the spatial domain: Airspace attention: Spatial attention includes linear and softmax branches: in, These are gating parameters used to adjust the linear attention branch and... The contribution of the attention branch.

[0030] Linear attention branches reduce the computational overhead in long sequence modeling through kernel mapping: in, The kernel function has a computational complexity of O(n log n). .

[0031] Softmax branches preserve precise spatial relationships: Spatial-frequency domain fusion based on uncertainty: To adaptively fuse spatial and frequency domain attention responses under different input conditions, this invention estimates the uncertainty based on the variance of the two branch outputs. Let the frequency domain branch output be... The output of the space frequency branch is Their corresponding output variances are respectively and The smaller the variance, the more stable the output of that branch and the lower the uncertainty.

[0032] Based on inverse variance weighting, the dual-spectral hybrid attention output is defined as: Through the above inverse variance fusion method, the attention branch with more stable output and smaller variance obtains higher weight in the fusion result; when the input features depend more on local spatial relationships, the contribution of the spatial domain branch increases, and when the input features depend more on global or periodic structures, the contribution of the frequency domain branch increases.

[0033] Gating parameters Adaptive adjustment via frequency domain energy ratio: Here, γ and b are learnable parameters. When the proportion of feature frequency domain energy is high, λ is increased to enhance the computational efficiency of the linear attention branch in long sequence modeling; when spatial domain structural information is dominant, λ is decreased to enhance the ability of the Softmax attention branch to express fine-grained spatial relationships.

[0034] In this way, DSHA utilizes both spatial details and frequency structure information during the long sequence visual feature encoding process, and fuses them according to the uncertainty of the branch output, providing a more stable visual feature representation for subsequent video evidence a posteriori modeling.

[0035] It should be understood that in step S1, the dissipative learnable residual connection DLRC models the network dynamics as a gradient flow system: Definition of the first The hidden state of the layer is The forward propagation process of a network is represented as an energy functional. Regarding hidden states Gradient flow update process: in, This represents the state update term along the direction of energy decrease. Indicates a cross-layer dissipative connectivity term. For the first Hidden state of layer for the first Learnable connectivity coefficients for layer hidden state updates; Discretizing the gradient flow update process yields the inter-layer update form of dissipative learnable residual connections: in, For the first The update step size corresponding to the layer, For the energy functional relative to the first The gradient of the hidden state of the layer, the learnable connectivity coefficients Satisfy dissipation conditions: Through the aforementioned dissipation constraint, The cross-layer residual terms constitute a non-negative weighted update of the historical hidden states and suppress characteristic oscillations caused by excessive inter-layer state update amplitudes; when satisfying the dissipation constraints and updating along the energy descent direction, the inter-layer state update satisfies the following energy constraint relationship: parameter Learn by solving constrained optimization problems: in, The regularization coefficient is used, and the constraint is implemented through the projective gradient descent method to enable the dissipative learnable residual connection to achieve constrained cross-layer information transmission in the deep network.

[0036] In this embodiment, traditional residual connections typically employ $H_{l+1} = H_l + F(H_l)$ to alleviate the gradient vanishing problem in deep network training. However, their inter-layer state updates lack explicit constraints on energy changes and the magnitude of cross-layer information transfer. Therefore, this invention proposes Dissipative Learnable Residual Connection (DLRC), which represents network forward propagation as a gradient flow update process of energy functionals. It also controls hidden state updates through learnable cross-layer dissipative connections to improve propagation stability during deep visual encoding.

[0037] Gradient flow modeling: Hidden state Considered as continuous depth The function on the network forward propagation corresponds to the energy functional. Regarding hidden states Gradient flow: in, This represents the state update term along the direction of energy decrease. This is a residual join term. A standard residual join corresponds to: The DLRC proposed in this invention corresponds to cross-layer dissipative connections: in, Let represent the learnable connection coefficients of the hidden state of layer p updated by the hidden state of layer l.

[0038] Discretizing the gradient flow above yields the inter-layer update form: in, For the first Layer update step size, This represents the gradient of the energy functional with respect to the current hidden state.

[0039] Dissipation conditions and energy diminishing: To control the update magnitude of the cross-layer residual terms to the current hidden state, the learnable connectivity coefficients satisfy the following dissipation constraint: Under this constraint, the cross-layer residual term constitutes a non-negative weighted update of the historical hidden state, which can suppress the characteristic oscillations caused by excessive inter-layer state updates. The corresponding energy change relationship is expressed as: when By optimizing the learning process and suppressing the second term, the interlayer energy change can be made non-increasing, thereby ensuring... .

[0040] Parameter learning: Learn by solving constrained optimization problems: Here, μ is the regularization coefficient, used to limit the connection coefficients from becoming too large. This optimization problem is solved using the projective gradient descent method, where the connection coefficients are projected to the feasible region after each update. in, For the projection operator to the probabilistic simplex, Update step size for connection coefficients By incorporating interlayer residual connections into the state update process under variational energy constraints, the deep image encoder has a more stable cross-layer information propagation capability when extracting long-sequence visual features, and provides a stable visual evidence representation for subsequent video temporal modeling.

[0041] S2: Based on the cognitive entropy minimization objective, the long sequence image encoder is subjected to two-stage contrastive learning and knowledge distillation training; The two-stage contrastive learning is used to enhance the semantic alignment capability and multi-scale robustness of visual feature representations, and the knowledge distillation is used to enable the student encoder to learn the output distribution, feature representation and cognitive entropy representation of the teacher model.

[0042] It should be understood that the two-stage comparative learning and knowledge distillation training are based on the goal of minimizing cognitive entropy.

[0043] After constructing the long-sequence image encoder, this invention trains the image encoder based on the objective of minimizing cognitive entropy. This training process includes two-stage contrastive learning and knowledge distillation based on cognitive entropy matching, enabling the encoder to acquire visual-semantic alignment capabilities while simultaneously modeling the uncertainty of latent variable representations. The image encoder parameters obtained through this training process are used for subsequent video encoder initialization and provide a stable visual feature foundation for modeling the posterior distribution of video evidence.

[0044] It should be noted that the cognitive entropy minimization objective in step S2 includes: Define cognitive entropy Used to measure latent variable representation With input data The degree of information retention and the uncertainty in representation are expressed as follows: in, For mutual information, For conditional entropy, For the weighting factor; Through variational inference, the cognitive entropy is represented as an optimizable variational upper bound: in, Let be the variational posterior distribution represented by the latent variables. The prior distribution is a latent variable; In the first stage of the two-stage contrastive learning, based on positive sample images negative sample images and corresponding text We construct a visual-semantic alignment loss and introduce a cognitive entropy regularization term to obtain the first-stage training objective: in, For similarity function, For temperature coefficient, For cognitive entropy weights; In the second stage of the two-stage contrastive learning, a multi-scale enhancement transformation is introduced. and Furthermore, by constraining the latent variable representations and variational posterior distributions corresponding to different augmented views to remain consistent, the second-stage training objective is obtained: in, This is the consistency weight.

[0045] Thus, by enhancing the alignment between image features and text semantics through the first-stage training objective, and by constraining representation consistency and posterior distribution consistency under multi-scale enhancement conditions through the second-stage training objective, the long sequence image encoder obtains visual feature representations with semantic alignment, robust expression, and uncertainty representation capabilities.

[0046] In this embodiment, the present invention proposes cognitive entropy to measure the information relationship and representation uncertainty between the latent variable representation H and the input data X. The cognitive entropy is defined as the weighted sum of mutual information terms and conditional entropy terms: Mutual Information measure Captured about Information content, conditional entropy Measure a given back The uncertainty. When At this time, cognitive entropy degenerates into mutual information, and maximizing mutual information leads to Retain as much as possible Information (which may contain noise); when At that time, cognitive entropy equals Minimizing marginal entropy encourages A deterministic representation of .

[0047] Through variational inference, the cognitive entropy is represented as an optimizable variational upper bound: in, Let be the variational posterior distribution represented by the latent variables. The prior distribution of the latent variables is defined. By optimizing the variational upper bound, the degree of information retention and uncertainty level of the latent variable representation can be constrained during the visual feature learning process.

[0048] It should be further explained that in step S2, knowledge distillation is based on cognitive entropy matching: Teacher Model With student model For the input data respectively Encode the code to obtain the teacher model output. Student model output Teacher model latent variable representation and the latent variable representation of the student model Simultaneously, a distillation loss is constructed between the teacher model and the student model, wherein the distillation loss includes an output distribution matching term, a latent variable representation matching term, and a cognitive entropy matching term: in, Used to constrain student model output With teacher model output Consistency Used to constrain the latent variable representation of the student model Teacher model latent variable representation The third constraint ensures consistency between the cognitive entropy of the student model and the teacher model. and The cognitive entropy corresponding to the student model and the teacher model are respectively used to ensure the integrity of knowledge transfer; Through the cognitive entropy matching term, the student model learns not only the output results and latent variable representations of the teacher model, but also the teacher model's representation of uncertainty in the input data. The knowledge distillation also includes feature-level distillation, which uses the maximum mean difference (MMD) constraint to constrain the distribution of latent variable representations in both the teacher and student models. in, For kernel mapping, For the regenerating nucleus Hilbert space; The knowledge distillation also includes relation-level distillation, which is used to preserve the relative distance structure between samples, and its loss function is expressed as: in, and For training samples; Thus, the knowledge distillation enables the student model to acquire semantic expression capabilities, feature distribution structures, and uncertainty representation capabilities consistent with the teacher model through output distribution matching, latent variable representation matching, cognitive entropy matching, feature distribution matching, and sample relationship structure matching.

[0049] In this embodiment, the first stage is used to perform basic visual-semantic alignment. Given a positive sample image... negative sample images and corresponding text The InfoNCE loss is used to constrain the similarity between image features and text features, and a cognitive entropy regularization term is introduced to obtain the first-stage training objective: The second stage is used for multi-scale robust enhancement. Different random enhancement transformations are applied to the input image. and Furthermore, by constraining the latent variable representations and their posterior distributions to remain consistent across different augmented views, the second-stage training objective is obtained. in, To enhance consistency weights, this stage uses feature consistency constraints and posterior distribution consistency constraints to ensure that the image encoder maintains a stable representation under different scales, viewpoints, or perturbation conditions, thereby improving the robustness of subsequent long-sequence video coding processes.

[0050] To further enhance the representational capabilities of the image encoder, this invention employs knowledge distillation based on cognitive entropy matching. Let the teacher model be T and the student model be S, with both outputting the teacher model's prediction results. Student model prediction results and the corresponding implicit variable representation Distillation loss is defined as: in, Used to match the output distributions of the student model and the teacher model. This is used to ensure that the latent variable feature distributions of the student model and the teacher model are consistent. and Let α and δ represent the cognitive entropy of the student model and the teacher model, respectively, and α and δ be the weighting coefficients.

[0051] Among them, the MMD (maximum mean difference) loss ensures feature distribution alignment: Through the knowledge distillation of cognitive entropy matching, the student model not only learns the output results and feature distribution of the teacher model, but also learns the teacher model's representation of the uncertainty of the input data, thereby enabling the trained long sequence image encoder to have more stable semantic expression ability and uncertainty modeling ability.

[0052] S3: Initialize the video encoder using the trained long sequence image encoder parameters and perform temporal modeling through spatiotemporal variational inference; The input long video sequence is represented as time-series data composed of multiple video segments or frame sequences. The posterior distribution of latent variables at the video segment level is obtained based on the video encoder, and the video evidence representation and video evidence uncertainty are constructed based on the posterior distribution of latent variables.

[0053] It should be understood that the video encoder is initialized using a trained long-sequence image encoder, and temporal modeling is performed through spatiotemporal variational inference. After completing the training of the long-sequence image encoder, this invention uses the trained image encoder parameters to initialize the spatial feature extraction part of the video encoder, and performs temporal modeling of the long-sequence video using spatiotemporal variational inference. This process represents frames or video segments in the video as a sequence of latent variables that evolve over time, describes the uncertainty of video evidence through variational posterior distribution, and further aggregates them to obtain a video-level evidence representation, providing a video evidence foundation for subsequent Bayesian thought chain inference.

[0054] It should be further explained that in step S3, the video encoder performs temporal modeling through spatiotemporal variational inference: Video The model is constructed as a Hidden Markov Model, and corresponding frame-level latent variables are set for each frame or video segment. The posterior distribution of the frame-level latent variables is approximated by variational inference: in, This represents the sequence of historical latent variables preceding the t-th frame or the t-th video segment, with the mean... and variance The mean and variance parameters of the posterior distribution are represented by the outputs of the time-series inference network, respectively. in, The visual feature table of the t-th frame or t-th video segment extracted by the long sequence image encoder. It is a temporal reasoning network; By constraining the temporal consistency of the video latent variable sequence using variational lower bounds, the temporal modeling objective is obtained: in, Assuming a temporal prior distribution, it is assumed that the latent variables between frames change slowly, which is used to constrain the latent variables of adjacent frames or adjacent video segments to maintain a smooth temporal change. The video-level global representation is obtained through attention aggregation of the latent variable sequence: in, For query-key attention functions, It is the mean vector of the latent variable sequence and serves as the query vector in attention aggregation.

[0055] Therefore, by obtaining the posterior distribution of latent variables at the video segment level, temporal consistency constraints, and video-level global representation through the spatiotemporal variational inference, video evidence representation and its uncertainty basis are provided for subsequent Bayesian thought chain inference.

[0056] In this embodiment, the video encoder includes a spatial feature extraction layer and a temporal modeling layer. The spatial feature extraction layer inherits the parameters of the long sequence image encoder trained by S2, while the temporal modeling layer is randomly initialized. in, This represents the parameters of the spatial feature extraction layer in the video encoder. This represents the parameters of the long sequence image encoder after training. This represents the parameters of the temporal modeling layer in the video encoder. To initialize variance By initializing the parameters as described above, the video encoder can inherit the visual semantic expression ability and uncertainty representation ability learned by the image encoder, and learn the temporal dependencies in the video sequence based on this.

[0057] Represent the input video as a sequence of frames or segments: in, Let represent the t-th frame or the t-th video segment, and T represent the length of the video sequence. The video is modeled as a Hidden Markov Process, where ... Let be the latent variable corresponding to the t-th frame or the t-th video segment, and its generation process is expressed as follows: in, The prior variance of the time series is used to constrain the smooth changes between adjacent latent variables. For decoding mapping, For observation variance Variational inference is used, employing inference networks. The posterior distribution of latent variables at approximately frame or fragment levels is used by the temporal inference network based on the latent variables from the previous time step. and visual features of the current frame or segment Output posterior parameters: Temporal consistency is constrained by the Evidence Lower Bound (ELBO): The first term constrains the explanatory power of latent variables for the current frame or video segment, while the second term constrains the deviation between the variational posterior distribution and the temporal prior distribution, thereby maintaining the temporal continuity of the video latent variable sequence.

[0058] Video-level global representations are obtained by attention aggregation of the latent variable sequences: in, This is the mean lookup vector for the latent variable sequence. For attention scoring function, Let t be the attention weight corresponding to the t-th latent variable.

[0059] Through the above spatiotemporal variational inference, the video encoder obtains the posterior distribution of latent variables at the video segment level. Posterior variance parameter and video-level global representation The posterior variance parameter is used to characterize the uncertainty of video evidence, and the video-level global representation is used as the video evidence representation in subsequent Bayesian thought chain inference.

[0060] S4: Construct a video reasoning model based on Bayesian thought chain and cognitive entropy constraints.

[0061] Among them, the video reasoning process is modeled as a latent variable posterior inference sequence consisting of multiple reasoning steps. When generating each reasoning step, the posterior probability of the reasoning step is determined based on the language prior and the likelihood of video evidence. The cognitive consistency score is calculated based on the uncertainty of video evidence and the degree of matching between the reasoning step and the video evidence. When the cognitive consistency score is lower than a preset threshold, the corresponding reasoning step or reasoning path is rejected for sampling, regenerated, or supplemented with video evidence retrieval; when the uncertainty of the video evidence is higher than a preset threshold, the reasoning depth is adaptively adjusted or relevant video segments are reselected, and finally a video reasoning output containing the reasoning process and the final conclusion is generated.

[0062] It should be understood that, after obtaining the video-level global representation zvid and the uncertainty of video evidence, this invention further constructs a video reasoning model based on Bayesian Chain-of-Thought (BCoT) and cognitive entropy constraints. This model represents the video reasoning process as a latent variable posterior inference process consisting of multiple reasoning steps, ensuring that each reasoning step is simultaneously constrained by linguistic priors and the likelihood of video evidence, thereby reducing the occurrence of reasoning steps deviating from the video content.

[0063] It should be noted that in step S4, the video reasoning model based on Bayesian thought chains includes: The video inference process is represented as a latent variable sequence posterior inference process consisting of multiple inference steps. Given an input video V and a query q, the sequence of inference steps is... The conditional probability is expressed as: in, For the first The hidden state of step-by-step reasoning. For the first The sequence of inference states before the step For video, For querying; For the The reasoning process involves modeling its posterior probability as a joint constraint between the video evidence likelihood term and the language prior term: in, The video evidence likelihood term is used to assess the consistency between the current reasoning step and its historical reasoning states and the input video evidence. For language priors, it is used to represent the prior probability of the language model generating the current reasoning step under the conditions of query and historical reasoning states; By employing variational approximation and inference networks To approximate the posterior probability, a Bayesian thought chain training objective is constructed: The first term is used to improve the consistency between reasoning steps and video evidence, and the second term is used to constrain the deviation between the posterior distribution of the reasoning network output and the prior distribution of the language model. Therefore, the Bayesian thinking chain video reasoning model makes each reasoning step subject to both the likelihood of video evidence and linguistic priors, so as to reduce unfounded reasoning caused by relying solely on linguistic priors to generate reasoning steps.

[0064] In this embodiment, the thought chain reasoning is modeled as a posterior inference of a sequence of latent variables. Let the query be... The video is represented globally as The reasoning steps are as follows: Their joint distribution is: in, This represents the sequence of inference states up to step k. The inference process at step k corresponds to posterior sampling: in, This is the likelihood term for video evidence, used to assess the consistency between the current reasoning step and the video evidence. These are language priors, provided by the language model based on the query and historical reasoning states.

[0065] Using variational approximation and inference networks Approximating the posterior. The optimization objective is the lower bound of evidence (ELBO): The first term is used to improve the reasoning steps' ability to interpret video evidence, while the second term is used to constrain the deviation between the posterior distribution of the reasoning network output and the prior distribution of the language model.

[0066] Through the Bayesian thought chain reasoning mechanism described above, each reasoning step is no longer generated solely by the language model prior, but is simultaneously influenced by video-level evidence representation. Likelihood constraints enable the inference chain to maintain a stronger correlation with the input video content.

[0067] It should be further explained that the Bayesian thought chain reasoning uses the information bottleneck principle to constrain cognitive consistency: Define cognitive consistency as a reasoning step The cognitive consistency constraint objective between the video evidence V and the reasoning state Maintaining relevance to video evidence V while suppressing its over-reliance on prior linguistic information in query q, the objective is expressed as: in, This represents the mutual information between the k-th step reasoning state and the video evidence. This represents the mutual information between the k-th step reasoning state and the query. This represents the information bottleneck weighting coefficient. The cognitive consistency constraint objective is transformed into an optimizable objective by using a variable boundary: in, The first output of the inference network The posterior distribution of the inference state in the step-by-step reasoning process. Used to characterize the probability of reconstructing or interpreting video evidence from the current reasoning state. For query-based Cognitive bottleneck reference distribution; During the reasoning process, a visual re-attention mechanism is used to refocus on the video evidence corresponding to each reasoning step, in order to generate the first... The first reasoning token or the first reasoning token For each reasoning step, calculate: in, For the first The query vector corresponding to the step-by-step inference text state. and These represent the keys and values ​​corresponding to the video evidence. For the first The video evidence response corresponding to the step-by-step reasoning state; Constructing cognitive consistency loss based on video evidence responses of adjacent reasoning steps: in, The Frobenius norm is used to constrain the continuity of the attention distribution to video evidence in adjacent reasoning steps, thereby reducing the unfounded drift of attention to video evidence during reasoning. A cognitive score is constructed based on the likelihood and uncertainty of video evidence. Candidate reasoning paths are screened based on the cognitive scores, retaining only reasoning steps or paths whose cognitive scores meet preset threshold conditions, and filtering out reasoning paths with low likelihood or high uncertainty of video evidence. This allows the Bayesian thought chain reasoning process to be jointly controlled by the relevance of video evidence, linguistic prior constraints, and uncertainty of video evidence.

[0068] It should be understood that the modules achieve deep collaboration within the aforementioned variational energy framework: The Topology Adaptive Expert Network (TAoE) and the Dissipative Learnable Residual Connection (DLRC) work together in the long sequence image encoding process, wherein TAoE is based on topological features extracted from persistent coherence. and DLRC dynamically adjusts expert routes based on dissipation conditions. Ensuring energy decreases ensures that the expert routing process is constrained by the input feature topology, while the cross-layer information propagation process is constrained by energy updates, thereby improving the stability and convergence efficiency of the deep visual coding process. When the energy decrease condition is met, the energy convergence relationship of the coding process is expressed as: in, This is the initial hidden state. For the first Hidden state, Let this be the target hidden state corresponding to energy minimization. The smallest eigenvalue of the energy Hessian matrix is ​​. For learning rate, For network depth; The dual-spectral hybrid attention mechanism (DSHA) and cognitive entropy constraints work synergistically during feature fusion. DSHA captures global periodic patterns in long-sequence video features through a frequency domain branch, preserves local spatial details through a branch in the original feature space, and performs adaptive fusion based on the uncertainties of the two types of branch outputs. This allows the fused latent variable representation to retain valid information from the video evidence while reducing uncertainty. The fused cognitive entropy constraint relationship is expressed as follows: The video encoder and the Bayesian thought chain inference model BCoT work together during the inference phase, wherein the video encoder obtains the posterior distribution of video latent variables through spatiotemporal variational inference. As prior knowledge for reasoning within a thought chain, it is injected into each step of the reasoning process through cross-attention: in, This is a video-level aggregation feature; this coupling enables the inference model to adaptively adjust the inference depth based on the uncertainty of video evidence, especially when the uncertainty of video evidence is high ( Automatically increase the number of reasoning steps. This ensures that the reasoning chain of the final output is constrained by the uncertainty of video evidence and cognitive consistency.

[0069] In this embodiment, to ensure consistency between the Bayesian thought chain reasoning process and video evidence, the present invention introduces a cognitive consistency constraint based on the information bottleneck principle during the reasoning stage. This constraint is used to enhance the reasoning steps. With video evidence The correlation between them, while suppressing the influence of inference steps on the query. The over-reliance on prior information in Chinese language is represented by the following objective: in, This represents the mutual information between the k-th step reasoning state and the video evidence. This represents the mutual information between the k-th step reasoning state and the query. This represents the information bottleneck weighting coefficient. The objective is to encourage reasoning steps to fully utilize video evidence while reducing the risk of relying solely on linguistic priors to generate reasoning content.

[0070] By varying the boundary, the above objective is transformed into an optimizable form: in, This represents the probability of interpreting or reconstructing video evidence based on the current reasoning state. Let represent the posterior distribution of the k-th step inference state output by the inference network. The reference distribution for cognitive bottlenecks is usually determined by the prior output of the language model based on query q.

[0071] Visual attention mechanisms: When generating each inference token, this invention recalculates the attention response to video evidence based on the current inference text state, forcing the model to refocus on video features: in, Let be the query vector corresponding to the inference text state at step k. and Here, represents the key and value corresponding to the video evidence, and d represents the feature dimension. This indicates the video evidence response corresponding to the current reasoning step.

[0072] This visual re-attention mechanism enables each reasoning step to re-acquire video evidence responses relevant to the current reasoning state during the generation process, thereby reducing the likelihood that the reasoning chain will gradually deviate from the video content during multi-step generation.

[0073] Loss of cognitive consistency: Based on video evidence responses from adjacent reasoning steps, a cognitive consistency loss is constructed: in, The Frobenius norm is represented. The cognitive consistency loss is used to constrain the video evidence responses corresponding to adjacent reasoning steps to maintain continuity, reducing the possibility of unfounded drift in video attention during reasoning.

[0074] Rejection Sampling and Inference Path Selection: To filter out low-belief reasoning paths, this invention defines a cognitive score based on the likelihood and uncertainty of video evidence: in, This represents the likelihood expectation of the video evidence at the current reasoning step. To indicate the uncertainty of the likelihood of the video evidence, the model only retains cognitive scores above a threshold. The reasoning path is to discard reasoning steps that are of low quality or deviate from video evidence.

[0075] When the cognitive score of the candidate reasoning step or reasoning path Below the preset threshold When the reasoning step or path is not in question, the relevant video evidence is discarded, regenerated, or reselected. When the cognitive score meets the preset threshold, the corresponding reasoning step is retained and subsequent reasoning continues.

[0076] By incorporating the aforementioned cognitive consistency constraints, this invention introduces video evidence relevance, language prior constraints, and video evidence likelihood uncertainty into the Bayesian thought chain reasoning process, enabling the final reasoning chain to be continuously constrained by video evidence during the multi-step generation process.

[0077] Multi-stage training strategy To train the video reasoning model based on Bayesian thought chains and cognitive entropy constraints, this invention employs a multi-stage training strategy in its embodiments. This training strategy progressively completes visual-language adaptation, low-rank parameter fine-tuning, cognitive score-driven policy optimization, rejection sampling supervised fine-tuning, and preference alignment optimization, enabling the reasoning model to more fully utilize video evidence and reduce the impact of low-belief reasoning paths when generating reasoning chains.

[0078] Phase 1: Adapter Fine-tuning. In the first phase, the parameters of the language model and video encoder are frozen, and only the vision-language adapter is trained, enabling the video evidence representation to map to a representation space acceptable to the language model. The adapter training objective is: in, For the video input or video evidence representation of the m-th training sample, For the k-th token in the corresponding output sequence, This represents the historical output up to the k-th token.

[0079] Phase 2: LoRA Fine-tuning. In the second phase, the low-rank adaptation parameters in the language model are unfrozen, and some parameters are updated using low-rank increments: W = W_0 + BA in, A and B are pre-trained weights, and A and B are low-rank adaptation matrices.

[0080] To suppress excessive perturbation of the original model's capabilities by low-rank increments, a perturbation suppression regularization term is introduced: in, To monitor and fine-tune losses, and This is the regularization weight coefficient.

[0081] Phase 3: Reinforcement Learning Optimization Based on Cognitive Scores. In the third phase, changes in cognitive scores are used as reward signals to optimize the inference policy network. Optimize accordingly. The reward for the k-th reasoning step is defined as: in, The cognitive score represents the k-step reasoning state. The policy network is evaluated using PPO or GRPO. Optimize the model to make it more likely to generate inference steps with higher likelihood and lower uncertainty for video evidence.

[0082] Phase 4: Reject Sampling SFT. In the fourth phase, candidate inference chains generated by the model are filtered based on cognitive scores, retaining samples whose cognitive scores meet a preset threshold. These samples are then used as supervised fine-tuning data to further train the video inference model. This phase aims to enhance the model's ability to generate highly cognitively consistent inference chains.

[0083] Phase 5: DPO Preference Optimization. Collect preference data. Preference optimization is employed. in, For the current strategy model, For reference strategy model, Optimize the temperature coefficient to suit preferences.

[0084] Through the above multi-stage training strategy, the video reasoning model gradually acquires the ability to access visual evidence, the ability to generate reasoning chains under the constraints of video evidence, the ability to optimize paths based on cognitive scores, and the ability to align preferences, thereby improving the stability of the Bayesian thought chain reasoning process and the consistency of video evidence.

[0085] The above description is only an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A long-sequence video reasoning method based on a variational energy framework and cognitive entropy constraints, characterized in that, The method includes: S1: Construct a long sequence image encoder based on a variational energy model. The encoder includes a topological adaptive expert network layer TAoE, a dual-spectral hybrid attention mechanism DSHA, and a dissipative learnable residual connection DLRC, which is used to obtain a structure-aware, spatial-frequency domain fusion and stable propagation visual feature representation under a unified energy constraint. The variational energy model is used to impose energy constraints on the visual feature encoding process, and its energy functional is expressed as: in, Represented as latent variables, For input data, For model parameters, Reconstructing energy from data As a priori constraint energy, For parameter regularization energy, the model's inference process corresponds to energy minimization: The training process corresponds to parameter optimization: The Topology Adaptive Expert Network (TAoE) layer analyzes the topological structure of the input features through persistent cohomology analysis and dynamically constructs an expert activation graph based on the topological structure. , where the vertex To gather experts, on the side The weights are determined by the topological connectivity of the feature space and are used to implement expert routing related to the structure of the input video. The dual-spectral hybrid attention mechanism DSHA calculates the attention response in the original feature space and frequency domain space respectively, and performs adaptive fusion based on the uncertainty of the two types of attention responses, in order to take into account both local spatial details and global frequency domain structure. The dissipative learnable residual connection (DLRC) models the inter-layer information flow as a gradient flow dynamics process and controls the inter-layer state update through dissipative constraints, thereby improving the stability of the deep visual feature propagation process. S2: Based on the cognitive entropy minimization objective, the long sequence image encoder is subjected to two-stage contrastive learning and knowledge distillation training; The two-stage contrastive learning is used to enhance the semantic alignment capability and multi-scale robustness of visual feature representations, and the knowledge distillation is used to enable the student encoder to learn the output distribution, feature representation and cognitive entropy representation of the teacher model. S3: Initialize the video encoder using the trained long sequence image encoder parameters and perform temporal modeling through spatiotemporal variational inference; The input long sequence video is represented as time-series data composed of multiple video segments or frame sequences. The posterior distribution of latent variables at the video segment level is obtained based on the video encoder, and the video evidence representation and video evidence uncertainty are constructed based on the posterior distribution of latent variables. S4: Construct a video reasoning model based on Bayesian thought chain and cognitive entropy constraints; Among them, the video reasoning process is modeled as a latent variable posterior inference sequence consisting of multiple reasoning steps. When generating each reasoning step, the posterior probability of the reasoning step is determined based on the language prior and the likelihood of video evidence. The cognitive consistency score is calculated based on the uncertainty of video evidence and the degree of matching between the reasoning step and the video evidence. When the cognitive consistency score is lower than a preset threshold, the corresponding reasoning step or reasoning path is rejected for sampling, regenerated, or supplemented with video evidence retrieval; when the uncertainty of the video evidence is higher than a preset threshold, the reasoning depth is adaptively adjusted or relevant video segments are reselected, and finally a video reasoning output containing the reasoning process and the final conclusion is generated.

2. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 1, characterized in that, In step S1, the specific implementation of the Topology Adaptive Expert Network (TAoE) layer includes: Constructing input features The topological complex, which adopts the Vietoris-Rips complex: in, It is the diameter of a simplex shape. These are the filtering parameters; The k-th order persistent homology group is calculated based on the Vietoris-Rips complex. Extracting topological feature descriptors from barcodes: in, For the first Betti number of order, For the first The birth-death interval of a topological feature. Indicates the first The persistence of topological features; Dynamically constructing expert routing matrix based on topology features , of which The expert and the first The routing weights among experts are represented as follows: in, Measurement experts and The number of connected components corresponding to the feature subspace. Measuring the persistence of a one-dimensional topological cycle, For feature dimensions; Based on the expert routing matrix Calculate the activation probability of each expert, the activation probability being determined by topological constraints. get: in, For the first The gating parameters corresponding to each expert For experts Pre-calculated activation scores, For topology routing weights, , For temperature coefficient, For the number of experts; Select a Top-K expert set based on the activation probabilities. The output of the topology adaptive expert network layer is calculated as follows: in: For the first Transformation results of input features by an expert For learnable residual scaling factors; Therefore, the Topology Adaptive Expert Network (TAoE) adjusts the expert routing weights by considering the topological connectivity of the input features and the one-dimensional persistent topological structure, so that the expert activation results correspond to the topological structure of the input features.

3. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 2, characterized in that, In step S1, the dual-spectral hybrid attention mechanism DSHA includes: Define dual attention as primitive spatial attention With frequency domain attention The coupling involves the frequency domain attention being mapped back to the original feature space via inverse discrete Fourier transform and then fused with the original feature space attention. in, and These are the Discrete Fourier Transform and its inverse transform, respectively. For querying the matrix, For coupling functions based on uncertainty; The frequency domain attention is calculated in Fourier space: in, The key matrix, For value matrices, This represents the Hermitian transpose. The dimension of the key vector; The original feature space attention includes a linear attention branch. and Softmax attention branch And adaptive fusion is achieved through the gating parameter λ: in, For kernel mapping function The coupling function Adaptive fusion is performed based on the uncertainties of the frequency domain branch output and the original feature space branch output. Let: The coupling function is then expressed as: in, and Frequency domain branch outputs and the original feature space branch output variance To prevent constants with zero denominator, inverse variance weighting is used to give higher weight to branches with smaller output variance in the fusion result. The gate parameters Based on the current hidden layer state Adaptive adjustment of frequency domain energy ratio: middle, and As a learnable parameter, when the proportion of frequency domain energy is high, Increase the value to enhance the computational efficiency of the linear attention branch in long sequence modeling.

4. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 3, characterized in that, In step S1, the dissipative learnable residual connection DLRC models the network dynamics as a gradient flow system: Definition of the first The hidden state of the layer is The forward propagation process of a network is represented as an energy functional. Regarding hidden states Gradient flow update process: in, This represents the state update term along the direction of energy decrease. Indicates a cross-layer dissipative connectivity term. For the first Hidden state of layer for the first Learnable connectivity coefficients for layer hidden state updates; Discretizing the gradient flow update process yields the inter-layer update form of dissipative learnable residual connections: in, For the first The update step size corresponding to the layer, For the energy functional relative to the first The gradient of the hidden state of the layer, the learnable connectivity coefficients Satisfy dissipation conditions: Through the aforementioned dissipation constraint, The cross-layer residual terms constitute a non-negative weighted update of the historical hidden states and suppress characteristic oscillations caused by excessive inter-layer state update amplitudes; when satisfying the dissipation constraints and updating along the energy descent direction, the inter-layer state update satisfies the following energy constraint relationship: parameter Learn by solving constrained optimization problems: in, The regularization coefficient is used, and the constraint is implemented through the projective gradient descent method to enable the dissipative learnable residual connection to achieve constrained cross-layer information transmission in the deep network.

5. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 1, characterized in that, In step S2, the objective of minimizing cognitive entropy includes: Define cognitive entropy Used to measure latent variable representation With input data The degree of information retention and the uncertainty in representation are expressed as follows: in, For mutual information, For conditional entropy, For the weighting factor; Through variational inference, the cognitive entropy is represented as an optimizable variational upper bound: in, Let be the variational posterior distribution represented by the latent variables. The prior distribution is a latent variable; In the first stage of the two-stage contrastive learning, based on positive sample images negative sample images and corresponding text We construct a visual-semantic alignment loss and introduce a cognitive entropy regularization term to obtain the first-stage training objective: in, For similarity function, For temperature coefficient, For cognitive entropy weights; In the second stage of the two-stage contrastive learning, a multi-scale enhancement transformation is introduced. and Furthermore, by constraining the latent variable representations and variational posterior distributions corresponding to different augmented views to remain consistent, the second-stage training objective is obtained: in, Consistency weight; Thus, by enhancing the alignment between image features and text semantics through the first-stage training objective, and by constraining representation consistency and posterior distribution consistency under multi-scale enhancement conditions through the second-stage training objective, the long sequence image encoder obtains visual feature representations with semantic alignment, robust expression, and uncertainty representation capabilities.

6. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 5, characterized in that, In step S2, knowledge distillation is based on cognitive entropy matching: Teacher Model With student model For the input data respectively Encode the code to obtain the teacher model output. Student model output Teacher model latent variable representation and the latent variable representation of the student model Simultaneously, a distillation loss is constructed between the teacher model and the student model, wherein the distillation loss includes an output distribution matching term, a latent variable representation matching term, and a cognitive entropy matching term: in, Used to constrain student model output With teacher model output Consistency Used to constrain the latent variable representation of the student model Teacher model latent variable representation The third constraint ensures consistency between the cognitive entropy of the student model and the teacher model. and The cognitive entropy corresponding to the student model and the teacher model are respectively used to ensure the integrity of knowledge transfer; Through the cognitive entropy matching term, the student model learns not only the output results and latent variable representations of the teacher model, but also the teacher model's representation of uncertainty in the input data. The knowledge distillation also includes feature-level distillation, which uses the maximum mean difference (MMD) constraint to constrain the distribution of latent variable representations in both the teacher and student models. in, For kernel mapping, For the regenerating nucleus Hilbert space; The knowledge distillation also includes relation-level distillation, which is used to preserve the relative distance structure between samples, and its loss function is expressed as: in, and For training samples; Thus, the knowledge distillation enables the student model to acquire semantic expression capabilities, feature distribution structures, and uncertainty representation capabilities consistent with the teacher model through output distribution matching, latent variable representation matching, cognitive entropy matching, feature distribution matching, and sample relationship structure matching.

7. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 1, characterized in that, In step S3, the video encoder performs temporal modeling through spatiotemporal variational inference: Video The model is constructed as a Hidden Markov Model, and corresponding frame-level latent variables are set for each frame or video segment. The posterior distribution of the frame-level latent variables is approximated by variational inference: in, This represents the sequence of historical latent variables preceding the t-th frame or the t-th video segment, with the mean... and variance The mean and variance parameters of the posterior distribution are represented by the outputs of the time-series inference network, respectively. in, The visual feature table of the t-th frame or the t-th video segment extracted by the long sequence image encoder. It is a temporal reasoning network; By constraining the temporal consistency of the video latent variable sequence using variational lower bounds, the temporal modeling objective is obtained: in, Assuming a temporal prior distribution, it is assumed that the latent variables between frames change slowly, which is used to constrain the latent variables of adjacent frames or adjacent video segments to maintain a smooth temporal change. The video-level global representation is obtained through attention aggregation of the latent variable sequence: in, For query-key attention functions, It is the mean vector of the latent variable sequence and serves as the query vector in attention aggregation; Therefore, by obtaining the posterior distribution of latent variables at the video segment level, temporal consistency constraints, and video-level global representation through the spatiotemporal variational inference, video evidence representation and its uncertainty basis are provided for subsequent Bayesian thought chain inference.

8. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 1, characterized in that, In step S4, the video reasoning model based on Bayesian thought chains includes: The video inference process is represented as a latent variable sequence posterior inference process consisting of multiple inference steps. Given an input video V and a query q, the sequence of inference steps is... The conditional probability is expressed as: in, For the first The hidden state of step-by-step reasoning. For the first The sequence of inference states before the step For video, For querying; For the The reasoning process involves modeling its posterior probability as a joint constraint between the video evidence likelihood term and the language prior term: in, The video evidence likelihood term is used to assess the consistency between the current reasoning step and its historical reasoning states and the input video evidence. For language priors, it is used to represent the prior probability of the language model generating the current reasoning step under the conditions of query and historical reasoning states; By employing variational approximation and inference networks To approximate the posterior probability, a Bayesian thought chain training objective is constructed: The first term is used to improve the consistency between reasoning steps and video evidence, and the second term is used to constrain the deviation between the posterior distribution of the reasoning network output and the prior distribution of the language model. Therefore, the Bayesian thinking chain video reasoning model makes each reasoning step subject to both the likelihood of video evidence and linguistic priors, so as to reduce unfounded reasoning caused by relying solely on linguistic priors to generate reasoning steps.

9. The long-sequence video reasoning method based on variational energy framework and cognitive entropy constraints as described in claim 8, characterized in that, The Bayesian thought chain reasoning uses the information bottleneck principle to constrain cognitive consistency. Define cognitive consistency as a reasoning step The cognitive consistency constraint objective between the video evidence V and the reasoning state Maintaining relevance to video evidence V while suppressing its over-reliance on prior linguistic information in query q, the objective is expressed as: in, This represents the mutual information between the k-th step reasoning state and the video evidence. This represents the mutual information between the k-th step reasoning state and the query. This represents the information bottleneck weighting coefficient. The cognitive consistency constraint objective is transformed into an optimizable objective by using a variable boundary: in, The first output of the inference network Posterior distribution of the inference state in the step. Used to characterize the probability of reconstructing or interpreting video evidence from the current reasoning state. For query-based Cognitive bottleneck reference distribution; During the reasoning process, a visual re-attention mechanism is used to refocus on the video evidence corresponding to each reasoning step, in order to generate the first... The first reasoning token or the first reasoning token For each reasoning step, calculate: in, For the first The query vector corresponding to the step-by-step inference text state. and These represent the keys and values ​​corresponding to the video evidence. For the first The video evidence response corresponding to the step-by-step reasoning state; Constructing cognitive consistency loss based on video evidence responses of adjacent reasoning steps: in, The Frobenius norm is used to constrain the continuity of the attention distribution to video evidence in adjacent reasoning steps, thereby reducing the unfounded drift of attention to video evidence during reasoning. A cognitive score is constructed based on the likelihood and uncertainty of video evidence. Candidate reasoning paths are screened based on the cognitive scores, retaining only reasoning steps or paths whose cognitive scores meet preset threshold conditions, and filtering out reasoning paths with low likelihood or high uncertainty of video evidence. This allows the Bayesian thought chain reasoning process to be jointly controlled by the relevance of video evidence, linguistic prior constraints, and uncertainty of video evidence.

10. A long-sequence video reasoning method based on a variational energy framework and cognitive entropy constraints according to any one of claims 1-9, characterized in that, Each module achieves deep collaboration within the aforementioned variational energy framework: The Topology Adaptive Expert Network (TAoE) and the Dissipative Learnable Residual Connection (DLRC) work together in the long sequence image encoding process, wherein TAoE is based on topological features extracted from persistent coherence. and DLRC dynamically adjusts expert routes based on dissipation conditions. Ensuring energy decreases ensures that the expert routing process is constrained by the input feature topology, while the cross-layer information propagation process is constrained by energy updates, thereby improving the stability and convergence efficiency of the deep visual coding process. When the energy decrease condition is met, the energy convergence relationship of the coding process is expressed as: in, This is the initial hidden state. For the first Hidden state, Let be the target hidden state corresponding to energy minimization. The smallest eigenvalue of the energy Hessian matrix is ​​. For learning rate, For network depth; The dual-spectral hybrid attention mechanism (DSHA) and cognitive entropy constraints work collaboratively during feature fusion. DSHA captures global periodic patterns in long-sequence video features through a frequency domain branch, preserves local spatial details through a branch in the original feature space, and performs adaptive fusion based on the uncertainties of the two types of branch outputs. This allows the fused latent variable representation to retain valid information from the video evidence while reducing uncertainty. The fused cognitive entropy constraint relationship is expressed as follows: The video encoder and the Bayesian thought chain inference model BCoT work together during the inference phase, wherein the video encoder obtains the posterior distribution of video latent variables through spatiotemporal variational inference. As prior knowledge for reasoning within a thought chain, it is injected into each step of the reasoning process through cross-attention: in, This is a video-level aggregation feature; this coupling enables the inference model to adaptively adjust the inference depth based on the uncertainty of video evidence, especially when the uncertainty of video evidence is high ( Automatically increase the number of reasoning steps. This ensures that the reasoning chain of the final output is constrained by the uncertainty of video evidence and cognitive consistency.