A flood video generation method based on flood forecast quantitative result driving

CN122783705APending Publication Date: 2026-09-18NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610930489.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

然而,此类模型主要通过大规模数据驱动学习图像或视频的统计分布规律,并未显式建模水文物理过程,因此在洪涝灾害场景生成中仍存在明显局限性,例如水体边界漂移、洪水扩散路径不符合地形约束、水流方向错误以及淹没时序不一致等问题

Benefits of technology

(1)本方法构建了基于数据与物理双驱动的极速推演机制:利用物理信息傅里叶神经算子,突破传统二维水动力学模型计算瓶颈,将一维洪水预报定量数值映射为二维空间掩码与光流图,实现了将一维预报定量数值(流量、水位)在极低延迟下,高保真地映射为界定洪水扩散边界的二维空间掩码与引导水体运动的时空光流图。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122783705A_ABST
    Figure CN122783705A_ABST
Patent Text Reader

Abstract

This invention discloses a flood video generation method driven by quantitative flood forecast results, belonging to the field of flood control emergency auxiliary decision-making technology. The method includes: parsing discrete-time time-series data, using physical information Fourier neural operators for deduction, and extracting continuous two-dimensional spatial mask sequences and optical flow map sequences; combining a specific watershed hydrological expert knowledge base and employing a multi-head cross-attention mechanism to embed expert knowledge; sequentially executing physical state reasoning, visual strategy mapping, and physical consistency verification based on a multi-stage thought chain to generate visual cue words; constructing a spatiotemporal control network and introducing a masking bias matrix into the self-attention mechanism, injecting the mask and optical flow map into a large-scale video generation model, and performing rendering output. This invention, using the above-mentioned flood video generation method driven by quantitative flood forecast results, generates highly realistic flood disaster evolution videos with spatiotemporal consistency and physical rationality, providing intuitive reference for flood control decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of flood control emergency auxiliary decision-making technology, and in particular to a flood video generation method driven by quantitative flood forecast results. Background Technology

[0002] Flood simulation primarily involves calculating and simulating the evolution of acquired rainfall data using appropriate hydrological and hydrodynamic mathematical models. Its core theoretical foundations include the shallow water equation and the Navier-Stokes equations, which govern continuous media. By numerically solving the rainfall generation, runoff, and water collection processes, it achieves an accurate characterization of the flood evolution process. Existing flood simulation methods typically output information such as inundation extent, water depth distribution, and velocity fields in two-dimensional or three-dimensional numerical results. However, these results are often presented in professional layers or numerical data formats, which are costly to understand for non-professional users and difficult to directly use for emergency decision-making.

[0003] To improve the interpretability of flood simulation results, researchers introduced computer graphics and Geographic Information System (GIS) technologies to visualize the flood inundation area. Water depth information can be displayed through color grading or raster rendering, thus achieving a spatial representation of flood risk. With the development of Virtual Reality (VR) technology, it has been widely applied in water conservancy engineering and disaster simulation. By constructing immersive flood disaster scenarios, users can perceive the flood evolution process in advance within a virtual environment, thereby improving emergency response and decision-making capabilities.

[0004] However, although virtual reality technology can provide a strong sense of immersion and interactive experience, it still relies on high-precision 3D modeling and complex scene rendering. In practical applications, it is constrained by computing resources and data acquisition costs, making it difficult to dynamically construct large-scale flood scenes in a short period of time. At the same time, its production cycle is long and its maintenance cost is high, which limits its real-time application capability in sudden flood events.

[0005] In the field of flood numerical simulation and visualization, early studies mainly relied on two-dimensional GIS maps to spatially display the flood inundation process, expressing the inundation depth and risk level of different areas through color gradients or raster hierarchies. For example, Liu et al. used GIS technology combined with a seed spread algorithm to calculate flood-inundated areas to improve the efficiency of inundation simulation; Yang et al. constructed a flood inundation analysis method based on a digital elevation model (DEM) and combined it with Arc Engine to visualize the inundated area; for sudden flood events such as dam breaks, Gao et al. constructed a flood diffusion model based on cellular automata to simulate and analyze the evolution process of dam-break floods. However, the above two-dimensional representation methods are essentially static or time-sliced ​​visualizations, which are difficult to fully depict the continuous evolution process and dynamic propagation mechanism of floods in the time dimension.

[0006] With the development of 3D computer graphics, flood simulation is gradually evolving from 2D representation to 3D visualization. Han et al. constructed a 3D topographic model of a watershed based on GIS data and combined it with OpenGL to realize a 3D visualization system for flood evolution; Li et al. used the MIKE hydrodynamic model to simulate the evolution of reservoir floods and combined it with BIM and GIS technologies to realize the construction of 3D flood scenes; Lai et al. proposed a method based on 3D mapping to couple hydrodynamic simulation results with a 3D topographic model to achieve a high-precision visualization of the dynamic evolution of floods; Zhao et al. constructed a 1D-2D coupled hydrodynamic model based on HEC-RAS and combined it with BIM+GIS technology to realize a 3D display of flood inundation downstream of a reservoir; Tang et al. used a coupled SWMM and OpenFOAM model to simulate urban flooding and reconstructed flood scenes using the glTF 3D data structure. Although 3D visualization methods have significantly improved in expressive power, they still rely on complex modeling processes and high computing resources.

[0007] Building upon this foundation, with the further development of virtual reality and digital twin concepts, researchers have begun to conduct dynamic simulation research on flood disasters based on 3D engines such as Unity3D and UE5. By integrating particle systems, real-time fluid simulation, and multi-source remote sensing data, they have achieved more realistic disaster scene construction. Wang et al. used drones to acquire terrain data and constructed a dam-break flood simulation scene based on 3ds Max; Wang et al. introduced shutter-based 3D technology into the Unity3D platform to achieve high-resolution dynamic flood simulation; Shi et al. introduced particle systems in UE5 to simulate rainfall processes and combined them with shallow water equations to achieve real-time water flow simulation; Lu et al. fused remote sensing and point cloud data based on UE5 to achieve highly realistic reconstruction of complex flood scenes. However, these simulation methods based on game engines typically rely on extensive manual modeling and high-performance computing power, making it difficult to quickly generate usable results in sudden disaster scenarios, and the development cycle is long and costly.

[0008] On the other hand, with the rapid development of AI generative models, text-to-image and text-to-video generation technologies based on diffusion models have gradually matured. Methods such as Stable Diffusion, video diffusion models, and text-to-video generation models can generate high-quality visual content based on natural language descriptions. Furthermore, generative models based on control conditions (such as ControlNet) introduce structured conditional inputs to achieve spatial constraints and structural control over the generation process. However, these models primarily learn the statistical distribution patterns of images or videos through large-scale data-driven learning, without explicitly modeling hydrological and physical processes. Therefore, they still have significant limitations in generating flood disaster scenarios, such as water boundary drift, flood diffusion paths not conforming to terrain constraints, incorrect water flow direction, and inconsistent inundation timing.

[0009] While existing flood simulation and visualization technologies have made some progress in numerical simulation, GIS representation, 3D visualization, and virtual reality, they still generally suffer from the following problems: On the one hand, traditional hydrodynamic models are computationally complex and struggle to meet the demands for real-time performance and rapid representation; on the other hand, 3D modeling and virtual reality methods are costly and time-consuming, making them unsuitable for sudden disaster scenarios; furthermore, existing video generation methods based on generative artificial intelligence lack hydrophysical constraints, making it difficult to guarantee consistency between the generated results and quantitative flood forecasts. Therefore, how to achieve the generation of dynamic flood disaster videos with physical consistency from quantitative flood forecast results has become a key technical problem that urgently needs to be solved in the field of flood simulation and disaster visualization. Summary of the Invention

[0010] The purpose of this invention is to provide a flood video generation method driven by quantitative flood forecast results, thereby solving the problems mentioned in the background art.

[0011] To achieve the above objectives, this invention provides a flood video generation method driven by quantitative flood forecast results, comprising the following steps: S1. Obtain the time-series quantitative results output by the flood forecasting system, quickly deduce the spatiotemporal physical field of floods based on the hydrodynamic surrogate model, and generate a two-dimensional spatial mask sequence and an optical flow map sequence reflecting water flow motion to control the video boundary; S2. Analyze key quantitative indicators of floods and map them into structured semantics, construct a knowledge base of experts for specific watersheds, embed expert knowledge through attention mechanism, and integrate hydrological and physical constraints into the multimodal latent features of prompt words; S3. For flood scenarios, guide the large language model to perform multi-stage thinking chain reasoning, map physical state, expert knowledge and visual strategy, and output high-quality visual control instructions after consistency verification. S4. Using a spatiotemporal control network, the physical field constraints generated in S1 and the visual instructions generated in S3 are jointly injected into the large video generation model to generate a flood evolution video that conforms to physical laws and quantitative forecast results.

[0012] Preferably, S1 includes: S11. Obtain the discrete time series data output by the hydrological forecasting system, perform tensor analysis on it, construct the time-varying boundary conditions that drive the deduction of the two-dimensional spatial physical field, and then construct the input feature tensor of the operator network. S12. Introduce a Fourier neural operator with physical information to construct a hydrodynamic surrogate model. Through the hydrodynamic surrogate model, perform rapid dynamic deduction on the input feature tensor and output a continuous two-dimensional submerged water depth matrix. S13. Set a water depth threshold and binarize the derived two-dimensional submerged water depth matrix to obtain a continuous two-dimensional spatial mask sequence for the control video large model. S14. Simultaneously extract the velocity component from the output tensor of the physical information Fourier neural operator to form a velocity vector. Map the velocity vector to the optical flow color space to generate a high-resolution optical flow map sequence.

[0013] Preferably, the continuous two-dimensional spatial mask sequence in S13 is as follows: ; In the formula, It is a continuous two-dimensional spatial mask sequence. This is a two-dimensional submerged water depth matrix. This is the water depth threshold.

[0014] Preferably, the high-resolution optical flow map sequence in S14 is shown in the following formula: ; ; In the formula, for directional flow velocity, for directional flow velocity, For flow velocity vectors, Let the magnitude of the flow velocity vector be denoted as . The direction of the flow velocity vector.

[0015] Preferably, S2 includes: S21. Receive the extracted set of quantitative indicators. ,in, To submerge the water depth, The cross-sectional average velocity is... The sediment concentration is represented by an attribute mapping function. The continuous quantitative feature values ​​are discretized and mapped into a structured natural language description that can be understood by the large language model. The mapping includes flow state determination and mapping, and the mapping results are concatenated to generate the structured initial prompt word text. S22. Construct an expert knowledge base for the target watershed that includes prior laws of hydrology and fluid mechanics. A pre-trained text encoder is used to transform the extracted physical constraint rule text into a high-dimensional prior knowledge feature matrix. ,in, The sequence length of the expert rule. For feature embedding dimension; S23. Input the initial prompt text generated in S21 into the text encoder to extract the latent feature matrix. ,in, The sequence length of the prompt words; S24. Employ a multi-head cross-attention mechanism, using the cue word feature matrix. As a query matrix, the expert knowledge feature matrix As the key matrix and value matrix, residual connections and layer normalization are introduced to output a multimodal fusion feature tensor after expert knowledge embedding.

[0016] Preferably, the multimodal fusion feature tensor in S24 is shown in the following formula: ; In the formula, For multimodal fusion feature tensors, For layer normalization function, For the prompt word feature matrix, This represents the attention weight.

[0017] Preferably, the multi-stage thought chain reasoning in S3 includes: S31, Physical State Reasoning Stage: The disaster text input by the user is decomposed into fine-grained physical features, and the unstructured text is transformed into a text representation that conforms to the objective laws of fluid mechanics, guiding the large language model to gradually perform causal inference. The logic of causal deduction is shown in the following formula: ; In the formula, This represents unstructured disaster situation text input by the user; This represents the spatial cross-section and local terrain attributes deduced from the context. Hydrodynamic characteristics; The reservoir is in a water storage and regulation state; Represents the macroscopic dynamic behavior of fluids; Represents the physical state derivation function; This is the set of physical state representations for the output; S32, Visual Strategy Mapping Stage: After obtaining a set of fine-grained physical state representations containing spatial cross-sections and reservoir status, the physical state representation set is transformed into the camera language of the large video generation model through an external enhancement mechanism, combined with a pre-built multimodal knowledge base, and the initial video generation prompts are output. S33. Consistency Verification Stage: The generated initial video prompts are combined with preset flood physics constraints as input to guide the large language model to perform self-logic checks, identify potential logical conflicts and descriptions that violate physical facts, and form a set of reflection results, as shown in the following formula: ; In the formula, The initial prompt words generated; Constrained by pre-set physical principles of flooding; This represents a reflection and verification function; For the collection of reflection results; If a visual illusion that violates the physical rules of flood control is detected in the generated instructions, a rewrite mechanism is triggered, and the final video generation prompt is output, as shown in the following formula: ; In the formula, Represents the correction function for a large language model; Generate prompts for the final high-realism video output.

[0018] Preferably, S32 includes: S321. Dense Retrieval: Using a pre-trained semantic embedding model, the target query and candidate documents are mapped to a high-dimensional continuous vector space, and the cosine similarity score between them is calculated, as shown in the following formula: ; In the formula, Indicates an embedding encoding operation; The second norm of a vector; For the target query statement; Candidate documents in the knowledge base; S322, Sparse Search: To address the need for precise matching between hydrological terminology and specific numerical values, the BM25 algorithm is introduced to calculate sparse search scores. The focus is on matching the section name with the flow rate value; S323, Weighted Normalization Fusion: The Min-Max normalization method is used to map the dual-path recall scores to intervals, and a weighting factor is introduced for linear fusion to construct an initial candidate fragment set; The linear fusion formula is shown below: ; In the formula, Represents the normalization function. The weighting factor for dense retrieval branches; S324, Relevance Rearrangement Based on Cross-Encoder: Rearranging the Query The candidate fragments from the initial recall are concatenated with the input of the cross-encoder, which outputs the final matching probability. The fragments are then sorted in descending order of their matching probability scores, and the first fragments are truncated. The highest-scoring segments are used to form a reference visual strategy set, from which the optimal two-dimensional space mask path is extracted; S325, Multi-dimensional Fusion Generation of Prompt Words: Construct a thought chain fusion context template to guide the large language model to deeply fuse the physical state representation set with the reference visual strategy set; the fusion process is decomposed into three dimensions: camera position setting, lighting and atmosphere rendering, and fluid material mapping, and finally outputs the initial video generated prompt words.

[0019] Preferably, S4 includes: S41. Collect drone aerial photography and monitoring videos of real flood disasters, extract and label the visual representation features of water bodies with high sediment content, and construct a text-video fine-tuning dataset; S42. Freeze the backbone network parameters of the pre-trained large model, and inject a trainable low-rank dimensionality-reduced matrix into the bypass of the projection matrices of the self-attention layer and cross-attention layer of the Diffusion Transformer module. and the increasing dimension matrix We employ low-rank adaptive techniques combined with a fine-tuning dataset to perform domain fine-tuning on a video generation model based on the diffusion Transformer architecture. The forward propagation calculation formula during the domain fine-tuning process is as follows: ; In the formula, For input implicit variables, For the output vector, To scale scalars, To freeze the backbone network parameters of the pre-trained large model, This is the weight increment matrix during the fine-tuning process; By fine-tuning the low-rank dimensionality reduction matrix and the increasing dimension matrix This enables large models to efficiently learn and fit the texture, color, and dynamic turbulence characteristics of water bodies in specific watersheds; S43. Concatenate the generated two-dimensional spatial mask sequence with the optical flow map sequence in the feature channel dimension to construct a comprehensive spatiotemporal conditional tensor; S44. Construct a parallel spatiotemporal control network, inputting the comprehensive spatiotemporal conditional tensor into a convolutional layer with initial weights of zero, and perform feature fusion, as shown in the following equation: ; In the formula, These are the features of the intermediate hidden layers in the backbone network. For conditional feature encoders, To integrate the spatiotemporal condition tensor, The control feature residual is the output of mapping encoded spatiotemporal conditional features with initial weights and zero convolutional layers; The fused tensor As a residual term injected into the latent space denoising block of the large video model; S45. Introduce a masking bias matrix into the self-attention mechanism of the diffusion Transformer architecture to impose a mandatory pixel generation constraint on the non-submerged region. S46. The large video generation model receives visual cues in the latent space, combines the feature injection of the spatiotemporal control network with the joint control of the masking bias matrix, executes a multi-step back-diffusion denoising inference process, and maps the latent space tensor into the decoder of the variational autoencoder back to the pixel space, outputting a video of the evolution of flood disaster.

[0020] Preferably, the masking bias matrix in S45 is shown in the following formula: ; In the formula, For the currently generated number The state values ​​of each target pixel in a continuous two-dimensional spatial mask sequence. For the self-attention calculation process and the target pixel The first step in calculating the correlation degree The state values ​​of each reference pixel in a continuous two-dimensional spatial mask sequence; Controlled attention calculation is updated as follows: ; In the formula, The projection matrix representing the latent space. To scale the dimensions, This is the penalty control coefficient.

[0021] Therefore, the present invention employs the above-mentioned flood video generation method driven by quantitative flood forecast results, which has the following beneficial effects: (1) This method constructs a rapid inference mechanism based on data and physics: by using the physical information Fourier neural operator, it breaks through the calculation bottleneck of the traditional two-dimensional hydrodynamic model, and maps the one-dimensional flood forecast quantitative values ​​into a two-dimensional spatial mask and optical flow map. This realizes the high-fidelity mapping of the one-dimensional forecast quantitative values ​​(flow rate, water level) into a two-dimensional spatial mask that defines the flood diffusion boundary and a spatiotemporal optical flow map that guides the movement of water bodies with extremely low delay.

[0022] (2) This method proposes an expert knowledge embedding method based on multi-head cross attention: innovatively constructs a hydrological expert knowledge base for a specific watershed, introduces an expert knowledge embedding mechanism in the semantic feature fusion stage, constrains the fluid dynamics of the large model when generating flood scenarios from the bottom-level representation, and effectively eliminates visual physical illusion.

[0023] (3) This method designs a multi-stage thinking chain reasoning architecture for flood scenarios: guides the large language model to perform reasoning according to the logical hierarchy of "physical state reasoning - visual strategy mapping - physical consistency verification", and accurately translates abstract quantitative indicators and professional terms into high-quality visual control instructions.

[0024] (4) This method developed a spatiotemporal joint control mechanism supported by multimodal data: based on the 6900 sets of high-fidelity “text-video-mask” multimodal flood disaster dataset, the video diffusion model was fine-tuned, and the mask space constraints and optical flow graph motion constraints were jointly injected into the model latent space through the spatiotemporal control network to realize the absolute control of high-fidelity video by quantitative values.

[0025] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0026] Figure 1 This is a flowchart of a flood video generation method driven by quantitative flood forecast results according to the present invention; Figure 2 This is a flowchart illustrating the process of obtaining a two-dimensional spatial mask sequence and a spatiotemporal hyperflow map sequence in a flood video generation method driven by quantitative flood forecast results according to the present invention. Figure 3 This is a schematic diagram of the physical information Fourier neural operator network structure of a flood video generation method driven by quantitative flood forecast results according to the present invention. Figure 4 This is a flowchart illustrating the multimodal feature fusion and knowledge embedding process of a flood video generation method driven by quantitative flood forecast results, as described in this invention. Figure 5 This is a flowchart of the dual-path hybrid retrieval mechanism of a flood video generation method driven by quantitative flood forecast results according to the present invention; Figure 6 This is a flowchart of the consistency verification chain for a flood video generation method driven by quantitative flood forecast results according to the present invention. Figure 7 This is a network architecture diagram of a controlled terrain boundary video model for a flood video generation method driven by quantitative flood forecast results, as described in this invention. Detailed Implementation

[0027] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0028] Example Please see Figures 1-7 This invention provides a flood video generation method driven by quantitative flood forecast results, comprising the following steps: S1. Obtain the time-series quantitative results output by the flood forecasting system, rapidly extrapolate the spatiotemporal physical field of flooding based on a hydrodynamic surrogate model, and generate a two-dimensional spatial mask sequence for controlling video boundaries and an optical flow map sequence reflecting water flow motion, such as... Figure 2 As shown, the specific steps include: Discrete-time series data output from the hydrological forecasting system are acquired, subjected to tensor quantization analysis, and time-varying boundary conditions are constructed to drive the derivation of the two-dimensional spatial physical field. Specifically, the input one-dimensional hydrological forecasting data sequence is defined as follows: ; In the formula, for The discharge flow rate at the upstream control section at all times. This represents the water level elevation at the corresponding cross-section. To estimate the total evolution duration, the static spatial feature matrix is ​​fused with the dynamic hydrological sequence to construct the input feature tensor of the operator network. As shown in the following formula: ; In the formula, Spatial coordinates; Topographic elevation matrix provided for digital elevation model (DEM); The Manning roughness matrix of the land surface is obtained by resolving the land cover type; The dynamic boundary condition matrix that maps one-dimensional flow rate and water level to the boundary of the computational domain; This represents a splicing operation along the feature channel dimension.

[0029] A hydrodynamic surrogate model is constructed by introducing the Physics-Informed Fourier Neural Operator (PI-FNO) to the input feature tensor. Perform rapid dynamic simulations, such as Figure 3 As shown, the specific steps include: Data-Physics Dual-Driven Network Construction and Expert Knowledge Embedding: Traditional data-driven models are prone to physical illusions. This method introduces an expert knowledge embedding mechanism at this step. The hydrological and sediment transport laws and shallow water dynamic equations of a specific watershed are used as physical constraints and embedded into the operator network training process in the form of a residual loss function. The mass and momentum conservation of the core two-dimensional shallow water equations (SWEs) controlling fluid motion are expressed as follows: ; ; ; In the formula, For water depth; They are respectively Flow velocity in the direction; It is the acceleration due to gravity; These are the bottom frictional stresses, which include Manning roughness and sand content correction factors. To reflect the physical constraints of the underlying topographic features, the specific calculation formula for the bottom frictional stress is defined as follows: ; In the formula, It is the acceleration due to gravity. This refers to the Manning roughness.

[0030] Calculate the physical residuals of the above equations. and the data fitting error Jointly optimize model weights.

[0031] Frequency Domain Global Convolution and Ultra-Fast Inference: The PI-FNO network extracts features through multiple Fourier operator layers. In the... In the operator layer, the feature update formula is as follows: ; In the formula, and These represent the Fast Fourier Transform (FFT) and the Inverse Fourier Transform (IFFT), respectively. This is the high-frequency filter weight matrix truncated in the frequency domain; Let be the linear transformation matrix in the spatial domain; It is a non-linear activation function. For the first A layer of Fourier operators in spatiotemporal coordinates The hidden feature tensor at the location. By learning the integral operator in Fourier space, the network can achieve fast mesh-invariant mapping. The model can directly output a continuous two-dimensional inundation depth matrix in a single forward propagation. .

[0032] Spatial mask binarization extraction: setting a water depth threshold The derived water depth matrix is ​​binarized to obtain a continuous two-dimensional spatial mask sequence for the large control video model, as shown in the following equation: ; The region with a mask value of 1 defines the absolute physical boundary of water flow diffusion in video generation.

[0033] To guide the large video model in generating dynamic water flow effects, an optical flow graph needs to be constructed simultaneously. This is done by simultaneously extracting the optical flow graph from the output tensor of the PI-FNO network. Directional flow velocity and Directional flow velocity . Flow velocity vector Mapping to the optical flow color space of computer vision standards. The direction of the flow velocity vector is mapped to hue, and the magnitude of the flow velocity is mapped to saturation, generating a high-resolution optical flow map. : ; ; In the formula, for directional flow velocity, for directional flow velocity, For flow velocity vectors, Let the magnitude of the flow velocity vector be denoted as . The direction of the flow velocity vector.

[0034] This optical flow map will serve as a dynamic guiding signal, and will be injected into the subsequent spatiotemporal control network along with the spatial mask generated in the previous steps.

[0035] S2. Analyze key quantitative indicators of flooding and map them into structured semantics, construct a specific watershed expert knowledge base, and embed expert knowledge through an attention mechanism to integrate hydrological and physical constraints into the multimodal latent features of prompt words. Addressing the issue of physical illusions easily generated by general-purpose large language models when planning visual strategies, this step introduces an expert knowledge embedding mechanism to ensure that the large-scale video generation model strictly follows the regional physical laws of the target watershed, such as... Figure 4 As shown, the specific steps include: Receive the set of quantitative indicators extracted in the previous stage, defined as: ; In the formula, Submerged water depth (unit: m). The average flow velocity across the cross section (unit: m / s). Sand content concentration (unit: kg / m³) 3 Using specific attribute mapping functions. This method discretizes continuous quantitative feature values ​​and maps them to a Structured Natural Language Description (SNLD) that can be understood by a large language model. The mapping rules are specifically based on empirical hydrodynamic formulas, and the mapping includes flow regime determination and mapping.

[0036] Flow regime determination and mapping: Introducing the dimensionless Froude number. To determine the flow pattern of water: ; in, This is the acceleration due to gravity. If (In slow-flow state), text mapping is performed. "The water surface is relatively calm, exhibiting continuous ripples and a slow current"; if (In rapid flow state), then text mapping "Rapid currents and surging waves, accompanied by strong hydrothermal phenomena and a large amount of white foam."

[0037] Sediment content characteristic mapping: setting the critical threshold for sediment suspension in the target watershed. .like Then text mapping "Yellowish-brown, highly turbid, sediment-laden water body with silt eddies on the surface"; conversely, a clearer description corresponds to a typical water body. These mapping results are then concatenated to generate a structured initial prompt text. .

[0038] Construct an expert knowledge base for the target watershed that contains clearly defined a priori laws of hydrology and fluid mechanics. This knowledge base stores specialized geological laws, such as sediment suspension and transport characteristics at specific flow velocities and hydraulic jump and surge patterns under complex topographic elevation differences, in the form of structured rule pairs. A pre-trained text encoder is used to transform the extracted physical constraint rule text into a high-dimensional prior knowledge feature matrix. ,in The sequence length of the expert rule. This is the feature embedding dimension. This matrix will serve as the physical knowledge base for subsequent multimodal fusion.

[0039] To achieve a deep integration of semantic features and physical laws, this invention performs expert knowledge embedding operations to correct the physical representation of the generated prompt word features within the latent space. Specifically, this includes the following steps: Initial semantic feature extraction: extracting the initial prompt text generated earlier. Using the same text encoder as input, its latent feature matrix is ​​extracted as follows: ; in, The sequence length of the prompt words.

[0040] Multi-Head Cross-Attention Fusion: Employs a multi-head cross-attention mechanism, using the cue word feature matrix... As a query matrix, the expert knowledge feature matrix The key and value matrices are shown below: ; In the formula, This is a learnable linear projection weight matrix. The calculation process for the attention weights is as follows: ; In the formula, The projection matrix representing the latent space. To scale the dimensions, This is the penalty control coefficient.

[0041] Output fused feature tensor: By introducing residual connections and layer normalization, the multimodal fused feature tensor after expert knowledge embedding is output, as shown in the following equation: ; Multimodal fusion feature tensor It not only contains the intuitive visual semantics of flood scenes, but is also forcibly bound to regional hydrological and physical constraints in the underlying representation, providing physical guidance for the subsequent multi-stage logical reasoning foundation.

[0042] S3, targeting flood scenarios, guides the large language model to execute multi-stage thought chain reasoning, maps physical states, expert knowledge and visual strategies, and outputs high-quality visual control instructions after consistency verification.

[0043] To address the technical issues that arise when general-purpose large language models directly convert unstructured disaster text into video generation prompts, resulting in the loss of physical laws and blurred visual representations (i.e., physical illusions), this paper proposes a multi-stage Chain-of-Thought (CoT) reasoning method for flood disaster video generation.

[0044] The multi-stage thought chain reasoning method specifically includes the following three cascaded processing stages: the physical state reasoning stage, which is responsible for parsing the user-inputted text using a fluid mechanics knowledge base; the visual strategy mapping stage, which performs precise mapping and planning based on the physical reasoning results from the previous stage; and the consistency verification stage, which, based on the initial prompts and physical states generated in the first two stages, guides the checking and correction of statements that violate objective physical laws. These three reasoning stages, through thought chains and related instructions, and through explicit step-by-step guidance, enable the large language model to extract physical attributes from complex disaster descriptions and, in conjunction with an external retrieval enhancement module, complete the mapping of professional visual strategies and self-reflection, generating high-quality video prompts that conform to fluid mechanics and the characteristics of real disasters.

[0045] The physical state reasoning chain includes: The system performs fine-grained physical feature decomposition on the disaster-related text input by users, transforming unstructured text into a text representation that conforms to the objective laws of fluid mechanics, guiding the large language model to gradually perform causal inference. The specific physical inference logic can be expressed by the following formula: ; In the formula, This represents unstructured disaster situation text input by the user; This represents the spatial cross-section and local terrain attributes deduced from the context. Hydrodynamic characteristics; The reservoir is in a water storage and regulation state; Represents the macroscopic dynamic behavior of fluids; Represents the physical state derivation function; This is the set of physical state representations for the output.

[0046] To improve the reliability of large-scale model inference, this step employs a few-shot prompting technique to design trigger templates and embed expert knowledge. Several standard hydrological and physical inference examples are pre-injected into the prompt words to guide the model to accurately extract core physical labels such as "section name" when processing unknown text, thereby providing strict physical boundary constraints for subsequent multimodal retrieval mapping.

[0047] The visual policy mapping chain includes: After acquiring a fine-grained set of physical state representations including spatial cross-sections and reservoir conditions, and combining it with a pre-built multimodal knowledge base, this is transformed into cinematic language easily understood by the large-scale video generation model through an external enhancement mechanism. This step includes: knowledge retrieval based on dual-path hybrid recall: A dual-path hybrid retrieval mechanism is constructed to address the characteristics of flood scene descriptions, which combine natural language and professional hydrological terminology. For example... Figure 5 As shown, the specific steps include: Dense Retrieval: Utilizes a pre-trained semantic embedding model to map the target query and candidate documents to a high-dimensional continuous vector space, and calculates their cosine similarity score. ; In the formula, Indicates an embedding encoding operation; The second norm of a vector; For the target query statement; These are candidate documents in the knowledge base.

[0048] Sparse Retrieval: Addressing the need for precise matching between hydrological terminology and specific numerical values, the BM25 algorithm is introduced to calculate sparse retrieval scores. The focus is on matching the cross-section name with the flow rate value.

[0049] Weighted Normalization Fusion: The Min-Max normalization method is used to map the dual-path recall scores to intervals, and a weighting factor is introduced for linear fusion to construct an initial candidate fragment set. The fusion formula is shown below: ; In the formula, Represents the normalization function. This is the weighting factor for dense search branches.

[0050] Relevance-based reordering based on cross-encoder: Query The fragments are concatenated with the initially recalled candidate fragments and input into a cross-encoder, which outputs the final matching probability. This is intended to filter out semantically relevant but logically inconsistent noisy samples. Sort by matching probability score in descending order, and truncate the first few results. The highest-scoring segments are used to form a reference visual strategy set, from which the optimal two-dimensional spatial mask path is extracted.

[0051] Multi-dimensional fusion generation of prompts: A thought chain fusion context template is constructed to guide the large language model to deeply fuse the set of physical state representations with the set of reference visual strategies. The fusion process is decomposed into three dimensions: camera position and perspective setting, lighting and atmosphere rendering, and fluid material mapping, ultimately outputting the initial video-generated prompts.

[0052] The consistency check chain includes: To ensure that the generated initial visual cues conform to objective fluid dynamics laws, a consistency verification closed-loop mechanism is introduced at the end of the visual cues generation process. For example... Figure 6 As shown, the specific steps include: Self-reflection: The initial prompts generated by the large language model are combined with pre-defined constraints based on common-sense knowledge of flood physics as input to guide the large language model to perform self-logical checks, identify potential logical conflicts and descriptions that violate physical facts, and form a set of reflection results, as shown in the following formula: ; in, The initial prompt words generated; Constrained by pre-set physical principles of flooding; This represents a reflection and verification function; For the set of reflection results.

[0053] Optimization and Correction: Based on the output reflection result set, the large language model optimizes and corrects the physical illusions existing in the initial prompts, outputting highly realistic final video-generated prompts that conform to fluid dynamics and real disaster characteristics, as shown in the following formula: ; in, Represents the correction function for a large language model; Generate prompts for the final high-realism video output.

[0054] S4. Using a spatiotemporal control network, the physical field constraints generated in S1 and the visual instructions generated in S3 are jointly injected into the large video generation model to generate a flood evolution video that conforms to physical laws and quantitative forecast results.

[0055] The generated 2D spatial mask and high-quality visual cues are jointly injected. Through domain fine-tuning and a spatiotemporal control network, the large video generation model is forced to perform high-fidelity flood rendering within a specified physical polygon range. By selecting a base model based on the Diffusion Transformer (DiT) architecture and combining it with domain fine-tuning, the model acquires the visual representation capability of water bodies with high sediment content. Based on this, a spatiotemporal control network is constructed, using the previously derived physical spatial mask and dynamic optical flow map as the underlying forced boundary. Simultaneously, a masking self-attention mechanism is introduced to penalize non-inundated areas. Finally, guided by multimodal semantic instructions, a high-fidelity flood disaster evolution video that strictly follows the hydrological quantitative forecast results is rendered. Figure 7 As shown, the specific steps include: First, fine-tuning of the domain to address the characteristics of flood evolution, including: Construction of flood feature dataset: Collect drone aerial photography and monitoring videos of real flood disasters, extract and label the visual representation features of water bodies with high sediment content, and construct a "text-video" fine-tuning dataset.

[0056] LoRA weight update mechanism: Freezes the backbone network parameters of the pre-trained large model. A trainable low-rank dimensionality-reduced matrix is ​​injected as a side path into the projection matrices (Query, Key, Value) of the self-attention and cross-attention layers of the DiT module. and the increasing dimension matrix The formula for calculating the forward propagation during the fine-tuning process is: ; in, For input implicit variables, For the output vector, To scale scalars, To freeze the backbone network parameters of the pre-trained large model, This is the weight increment matrix during fine-tuning. Fine-tuning the low-rank matrix... and This enables large models to efficiently learn and fit the texture, color, and dynamic turbulence characteristics of water bodies in specific watersheds.

[0057] Secondly, spatial conditional joint injection based on zero convolution and masked self-attention is performed, including: To ensure that the generated water flow does not overflow and conforms to the predicted flow velocity, the previously obtained continuous two-dimensional spatial mask sequence is used. With optical flow sequence Concatenate along the feature channel dimension to construct a comprehensive spatiotemporal conditional tensor. Specifically, it includes the following steps: Zero-convolution-based control network injection: Constructing a parallel spatiotemporal control network (SpatiotemporalControlNet). Injecting the spatiotemporal conditional tensor... Input with a convolutional layer with initial weights of zero This ensures that the prior capabilities of the pre-trained model are not compromised during the initial training phase. The feature fusion formula is: ; in, These are the features of the intermediate hidden layers in the backbone network. For conditional feature encoders, To integrate the spatiotemporal condition tensor, This refers to the control feature residuals output by mapping the encoded spatiotemporal conditional features to a convolutional layer with initial weights and zero values. The fused tensor... As residual terms, they are injected into the latent space of the large video model for denoising.

[0058] Self-attention penalty based on masking bias matrix: To impose a mandatory pixel generation constraint on non-submerged regions, a masking bias matrix is ​​introduced into the core self-attention mechanism of the DiT architecture. If the target pixel is generated If the pixel is located in a non-submerged area (corresponding to a mask value of 0), then the pixel of interest is the water body. The attention weights are penalized with an infinite amount, as shown in the following formula: ; In the formula, For the currently generated number The state values ​​of each target pixel in the continuous two-dimensional spatial mask sequence. For the self-attention calculation process and the target pixel The first step in calculating the correlation degree The state values ​​of the reference pixels in the continuous two-dimensional spatial mask sequence.

[0059] Controlled attention calculation is updated as follows: ; In the formula, The projection matrix representing the latent space. To scale the dimensions, This is a penalty control coefficient. Through this mechanism, the model is forced to prevent water pixels from spreading outside the spatial polygon range defined by the mask.

[0060] Finally, achieving controlled physical boundaries for video rendering and high-fidelity output includes: The large video generation model receives high-quality visual cues from the preceding output in the latent space. Simultaneously, the constructed zero-convolutional features and attention masking matrix jointly control the execution of a multi-step back-diffusion denoising inference process. After denoising, the latent space tensor is input into the decoder of the variational autoencoder, mapping the high-dimensional latent variables back to the high-resolution pixel space, and asynchronous rendering is performed. The final output is a high-fidelity flood disaster evolution video with high realism and water body dynamic evolution features that conform to the initial flood forecast quantitative results. .

[0061] The following examples further illustrate this method: The Xiaolangdi Reservoir and its downstream section in a certain river basin are key hubs for national flood control and drought relief. During the flood season, due to the typical high sediment content of the water in this basin, the evolution of flood peaks triggered by extreme rainfall is extremely complex. Traditional flood control decision-making relies heavily on one-dimensional cross-sectional water level and flow prediction values, making it difficult for commanders to intuitively predict the extent of flood overflow and the impact of high-turbidity jet streams on dams. Furthermore, commonly used video recording tools cannot accurately reflect the physical characteristics of this basin, such as "high sediment content" and "specific topographical surges," and tend to generate images of clear and calm water bodies.

[0062] This generation method has been innovatively applied in the Xiaolangdi Reservoir flood discharge and downstream flood control drills. Using this method, the system directly accesses the downstream flow process line and the predicted peak water level of the river section output by the hydrological forecasting center. Through the PI-FNO operator, a high-precision inundation mask is derived within an extremely short delay. Combined with the embedding of hydrological expert knowledge specific to this watershed, the flow regime evolution of high-turbidity water bodies is accurately reconstructed.

[0063] By repeatedly incorporating forecast values ​​under different operating conditions into video simulations, flood control command departments can intuitively and three-dimensionally observe the dynamic evolution of floodwaters, overflow boundaries, and realistic visual effects of water flow impacting weak sections of dikes under different discharge flows. Compared with traditional two-dimensional hydrodynamic rendering models, this method significantly reduces computation time; compared with large-scale pure text video models, this method ensures that the video is not only visually highly realistic but also strictly adheres to the quantitative results of flood forecasts in terms of physical boundaries and water flow characteristics, providing efficient and accurate visual decision support for the formulation of flood control and disaster relief plans and sand table simulations.

[0064] Therefore, the present invention adopts the above-mentioned flood video generation method driven by quantitative flood forecast results, which solves the problems of lack of physical constraints in existing video models and the time-consuming calculation of traditional hydrodynamic models, and generates highly realistic flood disaster evolution videos with spatiotemporal consistency and physical rationality, providing an intuitive reference for flood control decision-making.

[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A flood video generation method driven by quantitative flood forecast results, characterized in that, Includes the following steps: S1. Obtain the time-series quantitative results output by the flood forecasting system, quickly deduce the spatiotemporal physical field of floods based on the hydrodynamic surrogate model, and generate a two-dimensional spatial mask sequence and an optical flow map sequence reflecting water flow motion to control the video boundary; S2. Analyze key quantitative indicators of floods and map them into structured semantics, construct a knowledge base of experts for specific watersheds, embed expert knowledge through attention mechanism, and integrate hydrological and physical constraints into the multimodal latent features of prompt words; S3. For flood scenarios, guide the large language model to perform multi-stage thinking chain reasoning, map physical state, expert knowledge and visual strategy, and output high-quality visual control instructions after consistency verification. S4. Using a spatiotemporal control network, the physical field constraints generated in S1 and the visual instructions generated in S3 are jointly injected into the large video generation model to generate a flood evolution video that conforms to physical laws and quantitative forecast results.

2. The flood video generation method based on quantitative flood forecast results as described in claim 1, characterized in that, S1 includes: S11. Obtain the discrete time series data output by the hydrological forecasting system, perform tensor analysis on it, construct the time-varying boundary conditions that drive the deduction of the two-dimensional spatial physical field, and then construct the input feature tensor of the operator network. S12. Introduce a Fourier neural operator with physical information to construct a hydrodynamic surrogate model. Through the hydrodynamic surrogate model, perform rapid dynamic deduction on the input feature tensor and output a continuous two-dimensional submerged water depth matrix. S13. Set a water depth threshold and binarize the derived two-dimensional submerged water depth matrix to obtain a continuous two-dimensional spatial mask sequence for the control video large model. S14. Simultaneously extract the velocity component from the output tensor of the physical information Fourier neural operator to form a velocity vector. Map the velocity vector to the optical flow color space to generate a high-resolution optical flow map sequence.

3. The flood video generation method based on quantitative flood forecast results as described in claim 2, characterized in that, The continuous two-dimensional spatial mask sequence in S13 is shown in the following formula: ; In the formula, It is a continuous two-dimensional spatial mask sequence. This is a two-dimensional submerged water depth matrix. This is the water depth threshold.

4. The flood video generation method based on quantitative flood forecast results as described in claim 3, characterized in that, The high-resolution optical flow map sequence in S14 is shown in the following formula: ; ; In the formula, for directional flow velocity, for directional flow velocity, For flow velocity vectors, Let the magnitude of the flow velocity vector be denoted as . The direction of the flow velocity vector.

5. The flood video generation method based on quantitative flood forecast results as described in claim 4, characterized in that, S2 includes: S21. Receive the extracted set of quantitative indicators. ,in, To submerge the water depth, The cross-sectional average velocity is... The sediment concentration is represented by an attribute mapping function. The continuous quantitative feature values ​​are discretized and mapped into a structured natural language description that can be understood by the large language model. The mapping includes flow state determination and mapping, and the mapping results are concatenated to generate the structured initial prompt word text. S22. Construct an expert knowledge base for the target watershed that includes prior laws of hydrology and fluid mechanics. A pre-trained text encoder is used to transform the extracted physical constraint rule text into a high-dimensional prior knowledge feature matrix. ,in, The sequence length of the expert rule. For feature embedding dimension; S23. Input the initial prompt text generated in S21 into the text encoder to extract the latent feature matrix. ,in, The sequence length of the prompt words; S24. Employ a multi-head cross-attention mechanism, using the cue word feature matrix. As a query matrix, the expert knowledge feature matrix As the key matrix and value matrix, residual connections and layer normalization are introduced to output a multimodal fusion feature tensor after expert knowledge embedding.

6. The flood video generation method based on quantitative flood forecast results as described in claim 5, characterized in that, The multimodal fusion feature tensor in S24 is shown in the following equation: ; In the formula, For multimodal fusion feature tensors, For layer normalization function, For the prompt word feature matrix, This represents the attention weight.

7. The flood video generation method based on quantitative flood forecast results as described in claim 6, characterized in that, The multi-stage thought chain reasoning in S3 includes: S31, Physical State Reasoning Stage: The disaster text input by the user is decomposed into fine-grained physical features, and the unstructured text is transformed into a text representation that conforms to the objective laws of fluid mechanics, guiding the large language model to gradually perform causal inference. The logic of causal deduction is shown in the following formula: ; In the formula, This represents unstructured disaster situation text input by the user; This represents the spatial cross-section and local terrain attributes deduced from the context. Hydrodynamic characteristics; The reservoir is in a water storage and regulation state; Represents the macroscopic dynamic behavior of fluids; Represents the physical state derivation function; This is the set of physical state representations for the output; S32, Visual Strategy Mapping Stage: After obtaining a set of fine-grained physical state representations containing spatial cross-sections and reservoir status, the physical state representation set is transformed into the camera language of the large video generation model through an external enhancement mechanism, combined with a pre-built multimodal knowledge base, and the initial video generation prompts are output. S33. Consistency Verification Stage: The generated initial video prompts are combined with preset flood physics constraints as input to guide the large language model to perform self-logic checks, identify potential logical conflicts and descriptions that violate physical facts, and form a set of reflection results, as shown in the following formula: ; In the formula, The initial prompt words generated; Constrained by pre-set physical principles of flooding; This represents a reflection and verification function; For the collection of reflection results; If a visual illusion that violates the physical rules of flood control is detected in the generated instructions, a rewrite mechanism is triggered, and the final video generation prompt is output, as shown in the following formula: ; In the formula, Represents the correction function for a large language model; Generate prompts for the final high-realism video output.

8. The flood video generation method based on quantitative flood forecast results as described in claim 7, characterized in that, S32 includes: S321. Dense Retrieval: Using a pre-trained semantic embedding model, the target query and candidate documents are mapped to a high-dimensional continuous vector space, and the cosine similarity score between them is calculated, as shown in the following formula: ; In the formula, Indicates an embedding encoding operation; The second norm of a vector; For the target query statement; Candidate documents in the knowledge base; S322, Sparse Search: To address the need for precise matching between hydrological terminology and specific numerical values, the BM25 algorithm is introduced to calculate sparse search scores. The focus is on matching the section name with the flow rate value; S323, Weighted Normalization Fusion: The Min-Max normalization method is used to map the dual-path recall scores to intervals, and a weighting factor is introduced for linear fusion to construct an initial candidate fragment set; The linear fusion formula is shown below: ; In the formula, Represents the normalization function. The weighting factor for dense retrieval branches; S324, Relevance Rearrangement Based on Cross-Encoder: Rearranging the Query The candidate fragments from the initial recall are concatenated with the input of the cross-encoder, which outputs the final matching probability. The fragments are then sorted in descending order of their matching probability scores, and the first fragments are truncated. The highest-scoring segments are used to form a reference visual strategy set, from which the optimal two-dimensional space mask path is extracted; S325, Multi-dimensional Fusion Generation of Prompt Words: Construct a thought chain fusion context template to guide the large language model to deeply fuse the physical state representation set with the reference visual strategy set; the fusion process is decomposed into three dimensions: camera position setting, lighting and atmosphere rendering, and fluid material mapping, and finally outputs the initial video generated prompt words.

9. A flood video generation method based on quantitative flood forecast results as described in claim 8, characterized in that, S4 includes: S41. Collect drone aerial photography and monitoring videos of real flood disasters, extract and label the visual representation features of water bodies with high sediment content, and construct a text-video fine-tuning dataset; S42. Freeze the backbone network parameters of the pre-trained large model, and inject a trainable low-rank dimensionality-reduced matrix into the bypass of the projection matrices of the self-attention layer and cross-attention layer of the Diffusion Transformer module. and the increasing dimension matrix We employ low-rank adaptive techniques combined with a fine-tuning dataset to perform domain fine-tuning on a video generation model based on the diffusion Transformer architecture. The forward propagation calculation formula during the domain fine-tuning process is as follows: ; In the formula, For input implicit variables, For the output vector, To scale scalars, To freeze the backbone network parameters of the pre-trained large model, This is the weight increment matrix during the fine-tuning process; By fine-tuning the low-rank dimensionality reduction matrix and the increasing dimension matrix This enables large models to efficiently learn and fit the texture, color, and dynamic turbulence characteristics of water bodies in specific watersheds; S43. Concatenate the generated two-dimensional spatial mask sequence with the optical flow map sequence in the feature channel dimension to construct a comprehensive spatiotemporal conditional tensor; S44. Construct a parallel spatiotemporal control network, inputting the comprehensive spatiotemporal conditional tensor into a convolutional layer with initial weights of zero, and perform feature fusion, as shown in the following equation: ; In the formula, These are the features of the intermediate hidden layers in the backbone network. For conditional feature encoders, To integrate the spatiotemporal condition tensor, The control feature residual is the output of mapping encoded spatiotemporal conditional features with initial weights and zero convolutional layers; The fused tensor As a residual term injected into the latent space denoising block of the large video model; S45. Introduce a masking bias matrix into the self-attention mechanism of the diffusion Transformer architecture to impose a mandatory pixel generation constraint on the non-submerged region. S46. The large video generation model receives visual cues in the latent space, combines the feature injection of the spatiotemporal control network with the joint control of the masking bias matrix, executes a multi-step back-diffusion denoising inference process, and maps the latent space tensor into the decoder of the variational autoencoder back to the pixel space, outputting a video of the evolution of flood disaster.

10. A flood video generation method based on quantitative flood forecast results as described in claim 9, characterized in that, The masking bias matrix in S45 is shown in the following formula: ; In the formula, For the currently generated number The state values ​​of each target pixel in a continuous two-dimensional spatial mask sequence. For the self-attention calculation process and the target pixel The first step in calculating the correlation degree The state values ​​of each reference pixel in a continuous two-dimensional spatial mask sequence; Controlled attention calculation is updated as follows: ; In the formula, The projection matrix representing the latent space. To scale the dimensions, This is the penalty control coefficient.