A method, apparatus, equipment, medium, and product for assessing the health of hydropower equipment.

By using a unified time reference and a microsecond-level spatiotemporal dual compensation alignment mechanism, combined with graph attention networks and cross-modal temporal symbolization, the spatiotemporal semantic misalignment and semantic gap problems of multimodal data of hydropower equipment are solved, achieving high-precision and interpretable health assessment and eliminating the physical logic illusion of large models.

CN122490470APending Publication Date: 2026-07-31CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA THREE GORGES CORPORATION
Filing Date
2026-04-29
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as spatiotemporal semantic misalignment of multimodal data, semantic gap between heterogeneous modalities, and physical logic illusions generated by large language models when dealing with complex mechanical transmission chains and giant flexible hydropower equipment, making it difficult to achieve high-precision and interpretable health assessments.

Method used

Multi-source data is acquired by using a unified time reference. A topological mask matrix is ​​generated using a microsecond-level spatiotemporal dual compensation alignment mechanism and a graph attention network. By combining cross-modal temporal symbolization and feature manifold mapping network, a multimodal fusion semantic feature sequence is obtained. The device health assessment results are output using a Monte Carlo random deactivation mechanism, and physical constraint rules are embedded to suppress hallucinations.

Benefits of technology

It achieves high-precision, high-reliability, and interpretable health assessment of hydropower equipment, solves multimodal misalignment and semantic gap, and improves the stability and credibility of assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490470A_ABST
    Figure CN122490470A_ABST
Patent Text Reader

Abstract

This invention relates to the field of equipment condition monitoring technology, and discloses a method, device, equipment, medium, and product for assessing the health of hydropower equipment. Based on a unified time reference, it acquires the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of the hydropower equipment. After processing through a microsecond-level spatiotemporal dual compensation alignment mechanism, a multimodal perception dataset of the hydropower equipment is obtained. A graph attention network is used to encode the acquired dynamic spatiotemporal causal graph of the hydropower equipment and generate a topological mask matrix. Then, a cross-modal temporal symbolization and feature manifold mapping network is used to process the multimodal perception dataset to obtain a multimodal fused semantic feature sequence. This sequence is then input into a preset large model and processed using a Monte Carlo random deactivation mechanism to obtain equipment health assessment results that satisfy the PINN joint loss function of the neural network. This achieves high-precision, high-reliability, and interpretable continuous health assessment of hydropower equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of equipment condition monitoring technology, specifically to a method, apparatus, equipment, medium, and product for assessing the health of hydropower equipment. Background Technology

[0002] With the evolution of industrial internet and artificial intelligence technologies, equipment condition monitoring is developing from single signal analysis to multimodal comprehensive diagnosis. However, for heavy equipment with complex mechanical transmission chains and giant flexible body characteristics, existing technologies still face the following challenging technical problems in underlying data fusion and model inference mechanisms: 1. Physical-level misalignment of spatiotemporal semantics in multimodal data: There is an inherent difference in the sampling rate between macroscopic video surveillance images and high-frequency microscopic sensor signals (such as high-frequency vibration and sway), which can easily lead to temporal phase distortion. In addition, under variable operating conditions (such as drastic fluctuations in water head or active load), the small dynamic deformation of large racks can cause deviations in traditional static spatial calibration models, making it difficult to achieve high-precision spatiotemporal joint mapping.

[0003] 2. Semantic gap exists between heterogeneous modalities and large language models: Existing methods mostly use coarse-grained splicing at the numerical level. Large language models have difficulty directly parsing the dynamic mechanisms contained in one-dimensional high-frequency time-series signals. Visual features, temporal physical features and textual semantics cannot achieve deep and effective semantic alignment and joint reasoning within the manifold space of LLM.

[0004] 3. Purely data-driven large models are prone to physical logic "illusions": Existing large models are essentially generative networks based on statistical probability, lacking boundary constraints on the inherent physical laws of industrial equipment. When dealing with complex diagnostic tasks involving multivariate coupling, they are prone to generating erroneous feature associations and reasoning paths that violate the basic laws of mechanical dynamics, fluid mechanics, or electromagnetics, resulting in false alarms lacking physical interpretability and failing to meet the high reliability requirements of industrial control systems for diagnostic conclusions. Summary of the Invention

[0005] This invention provides a method, device, equipment, medium, and product for assessing the health of hydropower equipment, in order to solve the problems faced by existing technologies in the underlying data fusion and model reasoning mechanisms, such as the physical-level misalignment of spatiotemporal semantics of multimodal data, the semantic gap between heterogeneous modalities and large language models, and the tendency of pure data-driven large models to generate physical logic "illusions".

[0006] In a first aspect, the present invention provides a method for assessing the health of hydroelectric equipment, the method comprising: Based on a unified time reference, the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of the hydropower equipment are acquired. These are then processed using a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain a multimodal sensing dataset for the hydropower equipment. The dynamic spatiotemporal causal graph of the hydropower equipment and the PINN joint loss function of the physical information of the fused rotor dynamics equations of a pre-defined large model are obtained. The dynamic spatiotemporal causal graph is encoded using a graph attention network to generate a topological mask matrix. Based on the topological mask matrix, the multimodal sensing dataset is processed using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fused semantic feature sequence. This multimodal fused semantic feature sequence is input into a pre-defined large model and processed using a Monte Carlo random deactivation mechanism to obtain a device health assessment result containing a confidence interval that satisfies the PINN joint loss function.

[0007] The hydropower equipment health assessment method provided by this invention establishes a unified spatiotemporal benchmark for video, high-frequency sensing, and operating parameters by acquiring multi-source data based on a unified time reference, thus eliminating the asynchronous sampling problem of multi-source equipment. Furthermore, a microsecond-level spatiotemporal dual compensation alignment mechanism is used to align the acquired multi-source data, resolving the temporal phase distortion and spatial calibration deviation issues between video and high-frequency sensing signals, achieving physical-level precise matching of multimodal data. Further, by acquiring a dynamic spatiotemporal causal graph and using the PINN joint loss function, the physical mechanisms of hydropower equipment can be embedded into the model framework, establishing physical constraint rules and fundamentally suppressing large-model inference illusions. Finally, by generating a topological mask matrix through graph attention encoding, the physical transmission relationships of the equipment are structurally constrained, masking combinations of features without physical associations and ensuring that the inference path conforms to the laws of mechanical dynamics. Furthermore, under the constraint of the topological mask matrix, the multimodal perception dataset is processed using cross-modal temporal symbolization and feature manifold mapping networks to obtain a multimodal fusion semantic feature sequence. This transforms heterogeneous data into semantic features understandable by a large model, achieving deep fusion of visual, temporal, and operational information and eliminating the semantic gap between modalities. Further, by combining a pre-defined large model inference and a Monte Carlo random deactivation mechanism to process and output equipment health assessment results containing confidence intervals, single-point estimation is upgraded to probabilistic assessment, improving the stability of results under extreme operating conditions and providing highly reliable and interpretable health conclusions. Therefore, by implementing this invention, the three major problems of multimodal misalignment, semantic gap, and model physical illusion are solved, achieving high-precision, high-reliability, and interpretable continuous health assessment of hydropower equipment.

[0008] In one optional implementation, the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set are processed through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain a multimodal sensing dataset of the hydropower equipment, including: A low-light image enhancement network based on Retinex theory and self-attention mechanism is used to perform brightness equalization and contrast stretching on the initial monitoring video stream to obtain an initial visual feature map set. Using the exposure center time of the video keyframes corresponding to the initial visual feature map set as a reference, a fractional-order Sinc filter with a Hamming window truncation is used to resample the high-frequency sensor signal set to obtain a time-series signal set. Based on the real-time operating condition parameter set, displacement calculation is performed through a pre-set finite element proxy model, and a binary Gaussian spatial attention mask is constructed. Based on the binary Gaussian spatial attention mask and the initial visual feature map set, the target visual feature map set is determined. Based on the real-time operating condition parameter set, the time-series signal set, and the target visual feature map set, the multimodal sensing dataset of the hydropower equipment is determined.

[0009] The hydropower equipment health assessment method provided by this invention utilizes a low-light image enhancement network based on Retinex theory and a self-attention mechanism to perform brightness equalization and contrast stretching on the initial monitoring video stream. This improves the problems of low light, water mist, and oil mist interference in the water turbine room, enhances the texture of equipment surface defects, and improves the effectiveness of visual features. Furthermore, using the exposure center time of the video keyframe corresponding to the initial visual feature map as a benchmark, a fractional-order Sinc filter with a Hamming window truncation is used to resample the high-frequency sensor signal set, achieving microsecond-level time alignment, eliminating sub-sampling period phase errors, and avoiding time artifacts during fusion. Furthermore, by using a preset finite element proxy model for displacement calculation and constructing a binary Gaussian spatial attention mask, it is possible to compensate for the deformation of the unit's flexible body and camera vibration, achieving dynamic and accurate mapping from three-dimensional physical position to two-dimensional pixels. Furthermore, by combining the binary Gaussian spatial attention mask and the initial visual feature map and determining the target visual feature map, it is possible to guide the visual encoder to focus on the actual heating / vibration area, eliminating background interference and improving the accuracy of key feature extraction. Furthermore, by fusing trimodal data, a multimodal perception dataset with precise spatiotemporal alignment, clean features, and clear physical relationships can be formed. Therefore, by implementing this invention, through triple processing of time synchronization, spatial dynamic calibration, and image enhancement, microsecond-level temporal alignment and sub-pixel-level spatial alignment of multimodal data are achieved, providing high-quality underlying data for health assessment.

[0010] In one optional implementation, a low-light image enhancement network based on Retinex theory and a self-attention mechanism is used to perform brightness equalization and contrast stretching on the initial surveillance video stream to obtain an initial visual feature map set, including: The video frames in the initial monitoring video stream are preprocessed to obtain the target monitoring video stream. Using Retinex theory, the images in the target monitoring video stream are decomposed to obtain multiple initial reflection components and multiple initial illumination components. Based on a lightweight convolutional network containing a spatial self-attention mechanism, the multiple initial illumination components are dynamically balanced and smoothed to obtain multiple target illumination components. Based on the multiple target illumination components and multiple initial reflection components, an initial visual feature map set is generated.

[0011] The health assessment method for hydropower equipment provided by this invention can remove image noise and abnormal frames through preprocessing, improving the stability of subsequent decomposition and enhancement. Furthermore, by using Retinex theory to decompose and obtain multiple initial reflection components and multiple initial illumination components, it achieves the separation of the equipment's own texture features from the influence of ambient lighting, preserving true defect information. Furthermore, by optimizing the illumination components through a lightweight convolutional network incorporating a spatial self-attention mechanism, it can adaptively balance brightness, smooth uneven illumination, enhance details in dark areas without amplifying noise. Furthermore, by synthesizing multiple target illumination components and multiple initial reflection components to obtain an initial visual feature map set, it can highlight subtle defects such as blade cracks, cavitation, oil seepage, and deformation, improving visual feature recognition. Therefore, by implementing this invention, for the closed, low-light, and multi-interference visual acquisition environment of hydropower equipment, it can output clear, high signal-to-noise ratio, and defect-highlighting visual feature maps, ensuring the effective participation of visual modalities in the assessment.

[0012] In one optional implementation, displacement calculation is performed based on a real-time operating condition parameter set using a preset finite element proxy model, and a binary Gaussian spatial attention mask is constructed. This includes: calculating the displacement using a preset finite element proxy model based on the real-time operating condition parameter set to obtain the deformation offset of the hydropower equipment's three-dimensional spatial mesh relative to the static reference at the current moment; determining the camera vibration compensation amount using a visual odometer or the frame vibration frequency; obtaining the mapping center point coordinates based on the deformation offset and the camera vibration compensation amount through a dynamic spatial mapping equation; and constructing a binary Gaussian spatial attention mask based on the mapping center point coordinates.

[0013] The hydropower equipment health assessment method provided by this invention uses a pre-set finite element proxy model for displacement calculation, enabling rapid calculation of the flexible body displacement of the unit's frame and top cover under varying operating conditions, replacing the time-consuming traditional finite element calculation. Furthermore, by determining the camera vibration compensation amount through visual odometry or frame vibration frequency, it can correct camera pose offset caused by unit operating vibration, improving spatial mapping accuracy. Further, by processing and obtaining the coordinates of the mapping center point through a dynamic spatial mapping equation, it achieves real-time and accurate projection of the sensor's physical position onto image pixels, solving the problem of static calibration failure. Finally, based on the mapping center point coordinates, a binary Gaussian spatial attention mask is constructed, which can adaptively lock the excitation / heat source region, ensuring a strict correspondence between the model's receptive field and the physical measurement points. Therefore, by implementing this invention, dynamic deformation compensation and camera disturbance correction are achieved, solving the problem of spatial semantic misalignment under varying operating conditions of large units, thereby achieving sub-pixel-level spatial alignment accuracy.

[0014] In one optional implementation, based on a topological mask matrix, a multimodal sensing dataset is processed using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fused semantic feature sequence, including: Based on the multimodal sensing dataset, a temporal semantic symbol sequence is obtained through a temporal segmentation network based on physical spectrum envelope. Based on the topological mask matrix, a multi-head cross-attention fusion network is used to process the multimodal sensing dataset and the temporal semantic symbol sequence to obtain a multimodal fused semantic feature sequence.

[0015] The hydropower equipment health assessment method provided by this invention can transform one-dimensional high-frequency vibration and sway signals into semantic symbols that can be processed by large models through a time-series segmentation network based on physical spectrum envelopes, while preserving dynamic physical characteristics. Furthermore, under the constraint of a topological mask matrix, a multi-head cross-attention fusion network is used to process and obtain a multimodal fused semantic feature sequence. Multimodal feature association is achieved under physical topological constraints, avoiding physically meaningless feature combinations and improving the rationality of reasoning. Therefore, by implementing this invention, a high-fidelity mapping from industrial physical signals to the semantic space of large models is achieved, realizing cross-modal deep semantic alignment and joint reasoning, thus improving the accuracy and interpretability of the assessment.

[0016] In one optional implementation, the PINN joint loss function for obtaining physical information of the fused rotor dynamics equations of a preset large model includes: The simplified rotor dynamics differential equation with damping and gyroscopic effects, the cross-entropy loss and mean square error loss of the pre-set large model are obtained; latent variables characterizing the displacement or vibration trend of the shaft system are extracted from the hidden layer of the pre-set large model, and the second and first derivatives of the latent variables with respect to time are obtained using the automatic differentiation method of the deep learning framework; the latent variables, the second and first derivatives are input into the simplified rotor dynamics differential equation to obtain the physical residuals; based on the cross-entropy loss, mean square error loss and physical residuals, the PINN joint loss function of the neural network is constructed.

[0017] The health assessment method for hydropower equipment provided by this invention establishes a triple optimization objective of model training logic, health fitting, and physical laws by obtaining a simplified rotor dynamics differential equation with damping and gyroscopic effects, and pre-setting the cross-entropy loss and mean square error loss of a large model. Furthermore, by extracting latent variables and automatically differentiating them, the model output can be mapped to calculable physical quantities such as shaft displacement, velocity, and acceleration. Further, by inputting latent variables and their corresponding derivatives into the simplified rotor dynamics differential equation and obtaining physical residuals, the degree of deviation between model predictions and physical laws is quantified, which helps to identify illusory reasoning. Finally, by combining cross-entropy loss, mean square error loss, and physical residuals to construct a PINN joint loss function for a neural network, and by embedding physical constraints into the training process, the model can be forced to learn characteristic relationships that conform to mechanical dynamics. Therefore, by implementing this invention, by implanting hard physical constraints at the training objective level, illusory model outputs are significantly reduced, ensuring that the assessment results strictly follow the physical mechanisms of the equipment, and improving the usability in industrial scenarios.

[0018] In a second aspect, the present invention provides a health assessment device for hydroelectric equipment, the device comprising: The first acquisition module acquires the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of the hydropower equipment based on a unified time reference. The first processing module processes the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain a multimodal sensing dataset of the hydropower equipment. The second acquisition module acquires the dynamic spatiotemporal causal graph of the hydropower equipment and the PINN joint loss function of the physical information of the fused rotor dynamics equation of the preset large model. The encoding module encodes the dynamic spatiotemporal causal graph using a graph attention network and generates a topological mask matrix. The second processing module processes the multimodal sensing dataset based on the topological mask matrix using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fused semantic feature sequence. The third processing module inputs the multimodal fused semantic feature sequence into the preset large model and processes it using a Monte Carlo random deactivation mechanism to obtain a device health assessment result containing a confidence interval that satisfies the PINN joint loss function of the neural network.

[0019] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the water and electricity equipment health assessment method described in the first aspect or any corresponding embodiment thereof.

[0020] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the hydroelectric equipment health assessment method described in the first aspect or any corresponding embodiment thereof.

[0021] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the water and electricity equipment health assessment method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application scenario according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the method for assessing the health of hydroelectric equipment according to an embodiment of the present invention; Figure 3 This is a structural block diagram of a hydroelectric equipment health assessment device according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.

[0026] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0027] As an optional application scenario of this invention, considering the specific application environment architecture or specific hardware architecture upon which the hydropower equipment health assessment method depends, the specific application environment architecture or specific hardware architecture is described herein. For example... Figure 1 As shown, the architecture system may include at least one terminal device and at least one server. Figure 1 The system is illustrated in the example, which includes a computer 101, a mobile terminal 102, and a server 103, and the terminal devices such as the computer 101 and the mobile terminal 102 are connected to the server 103 through a network 110.

[0028] Specifically, the terminal device can be a smartphone, tablet, laptop, PDA, desktop computer, game console, smart TV, smart wearable device, in-vehicle terminal, VR (Virtual Reality) device, AR (Augmented Reality) device, etc. Server 103 can be a standalone physical server, a server cluster, a distributed system, or a cloud server providing cloud services. Network 110 can be a wired or wireless network, examples of which include, but are not limited to, the Internet, corporate intranet, local area network, wide area network, mobile communication network, and combinations thereof.

[0029] This invention provides a method for assessing the health of hydropower equipment, which solves three major problems: multimodal misalignment, semantic gap, and physical illusion of the model, and achieves high-precision, high-reliability, and interpretable continuous assessment of the health of hydropower equipment.

[0030] According to an embodiment of the present invention, a method for assessing the health of hydroelectric equipment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0031] This embodiment provides a method for assessing the health of hydroelectric equipment, which can be used on the aforementioned mobile terminals, such as mobile phones and tablets. Figure 2 This is a flowchart of a method for assessing the health of hydroelectric equipment according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps: Step S201: Based on a unified time reference, acquire the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of the hydropower equipment.

[0032] In one optional embodiment, the unified time reference means that an industrial-grade edge gateway supporting the IEEE 1588 PTP v2 protocol is used as the clock source to provide nanosecond-level hardware timestamps for all acquisition devices, thereby enabling video, sensor, and operational data to have a completely consistent time start and time scale, and eliminating the problem of asynchronous sampling of multiple devices at the hardware level.

[0033] In one optional embodiment, the initial monitoring video stream represents the raw video footage captured by high-definition monitoring cameras deployed in the water turbine room and generator floor, typically at a frame rate of 25fps / 30fps, used to capture visual defects such as surface cracks, oil seepage, deformation, cavitation, and oil / water mist on the equipment.

[0034] In an optional embodiment, the high-frequency sensing signal set represents vibration, sway, water pressure pulsation, and other signals collected by IEPE high-frequency sensors installed in key locations such as water guide bearings, thrust bearings, and stator frames, with a sampling rate of 10kHz to 50kHz, used to reflect the micro-dynamic state of the equipment.

[0035] In one optional embodiment, the real-time operating condition parameter set represents the unit operating status parameters collected in real time by the edge gateway, which are used to characterize the current load and hydrodynamic boundary conditions of the unit, and may include parameters such as active power, guide vane opening, and head height.

[0036] In one optional embodiment, visual images, high-frequency physical signals, and operating condition data of hydropower equipment can be collected synchronously under the constraint of a unified hardware clock.

[0037] For example, firstly, an industrial edge gateway supporting IEEE 1588 PTP v2 is deployed as the unified clock master node for the entire system. Then, a high-definition camera is connected to the gateway to capture video streams from the indoor water purifier equipment at a frame rate of 30Hz.

[0038] Furthermore, a high-frequency IEPE vibration / sway sensor is connected to the gateway, and the equipment dynamic signals are acquired at a sampling rate of 10kHz~50kHz. At the same time, the gateway is used to read the unit's SCADA system in real time and obtain operating parameters such as active power, guide vane opening, and head.

[0039] Furthermore, the gateway uniformly adds nanosecond-level hardware timestamps to video frames, sensor sampling points, and operating condition data.

[0040] Step S202: The initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set are processed by a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain a multimodal sensing dataset of hydropower equipment.

[0041] In one optional embodiment, the microsecond-level spatiotemporal dual compensation alignment mechanism represents an integrated collaborative alignment strategy that simultaneously performs bidirectional correction and compensation from both temporal and spatial dimensions to address issues such as clock offset, sampling phase difference, spatial position deviation, device deformation, and shooting disturbances present in multi-source heterogeneous acquisition devices.

[0042] Furthermore, this microsecond-level spatiotemporal dual compensation alignment mechanism uses microsecond-level time precision as a constraint to correct temporal misalignment of different modal data, and simultaneously completes the correction of spatial mapping deviation caused by equipment dynamic deformation and machine position vibration, thereby enabling precise coupling and matching of video, high-frequency sensing, and operating condition data in time axis and physical space.

[0043] In one optional embodiment, the multimodal sensing dataset represents a standardized fusion dataset formed after spatiotemporal dual compensation alignment. It integrates three types of heterogeneous information: visual image modality, high-frequency vibration sensing time-series modality, and unit operating condition parameter modality. The timestamps of each type of data correspond one-to-one, the spatial correlation is accurately matched, the data dimensions are complete, and the spatiotemporal benchmark is unified.

[0044] In one optional embodiment, based on a microsecond-level spatiotemporal dual compensation alignment mechanism, the time domain deviation compensation and spatial relationship correction of the three types of original heterogeneous data are uniformly completed, and the data misalignment problems caused by asynchronous acquisition of multiple devices, differences in installation positions, and dynamic operation disturbances of the unit are eliminated. In this way, a multimodal sensing dataset of hydropower equipment with high spatiotemporal matching and multi-source information correlation and coupling can be generated.

[0045] Step S203: Obtain the PINN joint loss function of the neural network for the physical information of the dynamic spatiotemporal causal graph of the hydropower equipment and the fusion rotor dynamics equation of the preset large model.

[0046] In one optional embodiment, the Dynamic Spatio-Temporal Causal Graph (DSTCG) represents a directed graph consisting of key component nodes and dynamic edge weights, with the mechanical structure and physical transmission path of the hydro-generator unit as priors. It is used to describe the mechanical, fluid, and electromagnetic transmission relationships between components and updates the correlation strength in real time according to the operating conditions, thus providing physical transmission prior constraints for large models.

[0047] In one optional embodiment, the preset large model represents a large language model with visual encoding, temporal encoding, multi-head cross-attention and text reasoning generation capabilities, which is used for health assessment in this embodiment.

[0048] In one alternative embodiment, the PINN joint loss function of the neural network represents a composite optimization objective function that integrates data-driven supervision loss and physical mechanism constraint loss, used to force the output of the large model to conform to the laws of mechanical dynamics during training and suppress inference illusion.

[0049] In one alternative embodiment, a dynamic spatiotemporal causal graph can be constructed based on the structure of the hydropower equipment and its real-time operating conditions.

[0050] For example, firstly, based on the actual mechanical transmission chain of the hydro-generator unit, a set of prior structural nodes for the unit is established by pre-setting key physical nodes. It may include nodes such as: generator upper guide, thrust bearing, stator frame, lower guide, water guide bearing, impeller, top cover, tailrace pipe, etc.

[0051] Furthermore, static transfer functions can be incorporated. Edge weights between nodes based on data-driven instantaneous transfer entropy dynamic calculation .

[0052] Among them, static transfer function It represents the mechanical model based on the equipment design phase, defining the inherent physical connections between components; the data-driven instantaneous transfer entropy calculates the dynamic correlation between nodes by monitoring data in real time.

[0053] Specifically, based on the mechanical model and shaft transmission topology of the hydroelectric equipment design phase, the inherent physical transmission relationships between components can be determined, and static foundation weights can be generated. Simultaneously, real-time multimodal data (vibration, sway, operating parameters) of the unit are accessed, and the instantaneous dynamic correlation between nodes is calculated through a data-driven algorithm to obtain the data-side weight, i.e., the dynamic instantaneous transfer entropy.

[0054] Furthermore, the static base weights By fusing with the dynamic instantaneous transfer entropy, dynamic edge weights that vary with time / operating conditions are obtained. .

[0055] Furthermore, combining the prior structure node set and dynamic edge weights A corresponding dynamic spatiotemporal causal graph can be constructed. Furthermore, by constructing a dynamic spatiotemporal causal graph, a layer of "physical constraint" can be added to the pre-defined large language model (LLM).

[0056] Step S204: Encode the dynamic spatiotemporal causal graph using a graph attention network and generate a topological mask matrix.

[0057] In one alternative embodiment, a Graph Attention Network (GAT) represents a deep learning model for processing graph-structured data. Its core is to dynamically learn the relationship weights between nodes and their neighbors through an attention mechanism, thereby more effectively aggregating neighbor information.

[0058] In an optional embodiment, the topology mask matrix is ​​a binary / weight matrix used in the multi-head cross-attention calculation of the Large Language Model (LLM) to force the weights of nodes with no direct physical relationship (e.g., between the upper guide bearing and the tailpipe) to be set to zero, thereby preventing the model from generating associations that violate physical common sense (i.e., "diagnostic illusions") from the source and ensuring that the reasoning path fully conforms to the industrial physical boundary.

[0059] In an optional embodiment, encoding the dynamic spatiotemporal causal graph using a graph attention network can transform the device physical topology rules into attention constraint rules that the model can execute, i.e., the output dimension can be directly embedded into the LLM attention layer. The topological mask matrix.

[0060] For example, taking equipment component nodes as input, the physical transmission weights between nodes are learned using GAT, and the weight distribution under grid connection / steady-state / variable operating conditions is dynamically updated. That is, the electromagnetic transmission weight increases during transients, and the mechanical weight dominates during steady-state, thereby outputting a characterization of the physical correlation strength between nodes.

[0061] Furthermore, construct according to the number of LLM tokens. The matrix is ​​further defined as follows: For node pairs with direct physical transmission paths, valid weights are retained; for node pairs without physical connections (such as the upper guide bearing and the tailrace pipe), the weights are forced to zero.

[0062] Furthermore, the mask matrix is ​​fed into the multi-head cross-attention module of LLM, and in the Attention calculation, invalid associations are filtered out using the mask, only physically reasonable feature interactions are retained, and the corresponding topological mask matrix is ​​finally output.

[0063] Step S205: Based on the topological mask matrix, the multimodal perception dataset is processed using cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fusion semantic feature sequence.

[0064] In one optional embodiment, the cross-modal temporal symbolization and feature manifold mapping network represents a dedicated network designed for multimodal data of hydropower equipment. It can transform one-dimensional high-frequency physical signals into high-dimensional semantic symbols that can be understood by large language models (LLMs), and achieve deep alignment and fusion of textual, visual, and temporal features in the same semantic space.

[0065] In one optional embodiment, the multimodal fusion semantic feature sequence represents a unified semantic token sequence that can be directly input into the LLM carrying all multimodal information of the device's health status and has physical consistency.

[0066] In an alternative embodiment, under the physical constraints of the topological mask matrix, a multimodal sensing dataset can be transformed into a unified multimodal fusion semantic feature sequence through cross-modal temporal symbolization and feature manifold mapping network.

[0067] Step S206: Input the multimodal fused semantic feature sequence into a preset large model and process it using the Monte Carlo random deactivation mechanism to obtain the device health assessment result containing the confidence interval that satisfies the PINN joint loss function of the neural network.

[0068] In one alternative embodiment, the Monte Carlo random deactivation mechanism is to keep Dropout enabled during the inference phase in the regression head of the large language model with a fixed deactivation probability (Drop rate = 0.1~0.3), perform multiple forward propagations on the same set of inputs, and use the statistical distribution of multiple outputs to characterize the uncertainty of the evaluation results, and finally output the health score with confidence interval, rather than a single point estimate.

[0069] In one optional embodiment, the multimodal fused semantic feature sequence is input into a preset physical constraint large language model, and then Monte Carlo random deactivation is enabled during inference. Combined with the PINN physical residual joint loss function, the final output of the hydropower equipment health assessment result with confidence interval and physical consistency can be achieved.

[0070] For example, the multimodal fused semantic feature sequence is input into a preset large model, and then Monte Carlo random deactivation is enabled in the regression head (health output layer) of the preset large model, that is, Dropout is set to 0.1~0.3, and Dropout is not turned off during the inference stage to maintain random deactivation.

[0071] Furthermore, perform M independent forward propagations (usually M=30) on the same set of input data, obtaining M health output values. Then, calculate the mean of the M outputs. With variance And obtain the probability distribution of health status.

[0072] Furthermore, confidence intervals can be generated at a 95% confidence level. It outputs the device health assessment results, which include health score, confidence interval, and physical consistency check, and satisfy the PINN joint loss function of the neural network.

[0073] The health assessment method for hydropower equipment provided in this embodiment solves three major problems: multimodal misalignment, semantic gap, and physical illusion of the model, and realizes high-precision, high-reliability, and interpretable continuous health assessment of hydropower equipment.

[0074] In some optional implementations, step S202 above includes: Step S2021: Using a low-light image enhancement network based on Retinex theory and self-attention mechanism, brightness equalization and contrast stretching are performed on the initial monitoring video stream to obtain an initial visual feature map set.

[0075] In one optional embodiment, the low-light image enhancement network based on Retinex theory and self-attention mechanism represents an image enhancement network designed for enclosed machine rooms of hydropower equipment, low-light, and high-moisture scenarios. By combining Retinex decomposition and spatial self-attention, it can highlight defects such as equipment cracks, oil seepage, and cavitation without amplifying noise.

[0076] Furthermore, the Retinex theory states that it is composed of the retina and the cortex, and aims to explain the human visual perception of the color of objects under different lighting conditions.

[0077] Specifically, step S2021 above includes: Step a1: Preprocess the video frames in the initial monitoring video stream to obtain the target monitoring video stream.

[0078] Step a2: Using Retinex theory, the image in the target surveillance video stream is decomposed to obtain multiple initial reflection components and multiple initial illumination components.

[0079] Step a3: Based on a lightweight convolutional network containing a spatial self-attention mechanism, multiple initial illumination components are dynamically balanced and smoothed to obtain multiple target illumination components.

[0080] Step a4: Generate an initial visual feature set based on multiple target illumination components and multiple initial reflection components.

[0081] In one optional embodiment, consecutive image frames of the initial monitoring video stream are read frame by frame, and batch preprocessing operations are performed to output a target monitoring video stream with consistent specifications and preliminary interference filtering. The batch preprocessing operations may include: cropping invalid edges, global weak noise reduction for water vapor and fog, normalizing pixel grayscale range, uniformly scaling resolution, removing random noise and invalid redundant areas during video acquisition, and unifying the image specifications and data distribution of all video frames.

[0082] Furthermore, Retinex theory can be used to perform hierarchical decomposition on a single frame image in the target surveillance video stream, and the corresponding initial reflection component and initial illumination component can be obtained, as shown in the following relationship (1): (1) In the formula: Represents the pixel values ​​of a single frame of the original image; Indicates the initial reflection component; This represents the initial illumination component.

[0083] Furthermore, by traversing the entire image using multi-scale Gaussian convolution, the ambient illumination distribution is estimated pixel-by-pixel, resulting in multiple initial illumination components at different scales. Further, based on color constrained inverse solving, the device's texture is preserved, and the corresponding initial reflection components are calculated.

[0084] Furthermore, the obtained initial illumination components are input into a lightweight convolutional network containing a spatial self-attention mechanism. Then, the local brightness distribution features of the illumination are extracted through lightweight depth convolution, and spatial self-attention weights are introduced to calculate the corresponding spatial attention weight matrix.

[0085] Furthermore, the spatial attention weight matrix can be used to model the global pixel illumination, and dynamic brightness equalization can be performed on overly bright, overly dark, and shadow areas. At the same time, the edge-preserving smoothing operator can be combined to perform noise reduction and smoothing of the illumination components, correct illumination distortion, and finally output the corresponding multiple target illumination components.

[0086] Furthermore, multiple initial reflection components and multiple corrected target illumination components are substituted into the above relation (1), and the enhanced image is reconstructed pixel by pixel. At the same time, local contrast is stretched and global brightness is balanced, thereby outputting the corresponding initial visual feature set.

[0087] Step S2022: Using the exposure center time of the video keyframe corresponding to the initial visual feature map set as a reference, the high-frequency sensor signal set is resampled using a fractional Sinc filter truncated by a Hamming window to obtain a time-series signal set.

[0088] In one optional embodiment, the exposure center time of the video keyframe is the midpoint of the video frame shutter exposure, which serves as the absolute time reference for multimodal alignment, enabling the high-frequency vibration signal to be strictly aligned with the image.

[0089] In an alternative embodiment, the Hamming window-truncated fractional-order Sinc filter represents a high-order finite-length filter that combines the Hamming window function truncation constraint with fractional-order interpolation characteristics to resolve sub-microsecond time misalignment between video (30Hz) and high-frequency sensing (10k-50kHz) and eliminate phase truncation error.

[0090] In an optional embodiment, the exposure center time of each video keyframe within the initial visual feature map is first extracted. This is used as the absolute reference point for timing alignment. Simultaneously, the raw discrete sampling sequence of the high-frequency sensor is read. Then, calculate the video reference time. The small time deviation between the current sampling point and the nearest sensor sampling point yields the subsampling period phase truncation error (fractional delay). It is used to correct time misalignment caused by sampling intervals that are not multiples of integers.

[0091] Furthermore, regarding the aforementioned subsampling period phase truncation error... By activating a fractional-order Sinc filter based on Hamming window truncation at the system edge, the discrete sampling sequence of the sensor is resampled, as shown in the following relation (2): (2) In the formula: This represents the resampled target signal value. That is, the calculated value relative to the exposure center time of the video keyframe. Sensor signal strengths that are perfectly aligned in time; Indicates that the sensor is in The raw values ​​actually collected at any given time; Indicates the fractional-order Sinc interpolation weights; This represents the weighting coefficient of the Hamming window.

[0092] Among them, utilizing Function to construct ideal band-limited interpolation kernel And introduce the Hamming window function. By truncating the filter in the time domain, the spectral leakage and Gibbs oscillation effect of the ideal Sinc filter can be suppressed, and a finite-length weighted interpolation kernel can be formed.

[0093] Furthermore, regarding the reference time The original sampling points before and after are weighted and summed to obtain the exposure center time of the video keyframe. Time-aligned resampled target signal values .

[0094] Furthermore, iterate through all the video keyframes corresponding to... At any given moment, the above interpolation and resampling process is repeated for all types of high-frequency sensor signals to precisely shift the time-domain phase of the one-dimensional high-frequency signal to the same microsecond moment when the video shutter opens, eliminating time artifacts during cross-modal fusion, and finally integrating all aligned signal sequences to form a corresponding time-series signal set.

[0095] Step S2023: Based on the real-time operating condition parameter set, displacement calculation is performed through a preset finite element proxy model, and a binary Gaussian spatial attention mask is constructed.

[0096] In one optional embodiment, the preset finite element proxy model is an alternative fitting model constructed using the finite element simulation dataset of the unit structure as training samples. By fitting the nonlinear mapping relationship between the load parameters and structural deformation through machine learning, it can balance the deformation calculation accuracy and the real-time computing efficiency of industrial scenarios.

[0097] In one optional embodiment, the binary Gaussian spatial attention mask represents a pixel-level weight matrix generated based on a two-dimensional binary Gaussian distribution function, with the physical mapping center point as the core. This matrix guides the LLM visual encoder to accurately locate the heat / excitation source. Pixel regions closer to the center point have higher weights, while the weights of distant background regions decrease progressively with each layer.

[0098] Specifically, step S2023 above includes: Step b1: Based on the real-time operating condition parameter set, displacement calculation is performed using a preset finite element proxy model to obtain the deformation offset of the three-dimensional spatial mesh of the hydropower equipment relative to the static reference at the current moment.

[0099] Step b2: Determine the camera vibration compensation amount using visual odometry or gantry vibration frequency.

[0100] Step b3: Based on the deformation offset and camera vibration compensation, the coordinates of the mapping center point are obtained through dynamic spatial mapping equation processing.

[0101] Step b4: Construct a binary Gaussian spatial attention mask based on the coordinates of the mapped center point.

[0102] In one optional embodiment, the three-dimensional spatial mesh of the generator unit represents a set of three-dimensional structured meshes discretized based on the actual geometric dimensions and assembly relationships of key components such as the turbine generator unit frame, bearing housing, top cover, and support structure, used to describe structural deformation. Each mesh node corresponds to the spatial coordinates of the equipment entity.

[0103] In an optional embodiment, the static reference represents the original three-dimensional coordinates and camera intrinsic and extrinsic parameters under standard stable operating conditions where the unit is shut down, without load, vibration, or deformation.

[0104] In an optional embodiment, deformation offset It represents the three-dimensional displacement increment (sub-millimeter to millimeter level) of the unit's flexible body under full load / variable operating conditions due to stress, and is used to characterize the stress deformation state of the equipment structure.

[0105] In an optional embodiment, camera vibration compensation amount This indicates the real-time offset of the camera's attitude, shooting angle, and installation position caused by the transmission of vibrations and slight swaying of the unit during operation to the monitoring camera.

[0106] In one optional embodiment, strain gauge feedback data deployed at key stress points of the frame are first collected as direct boundary conditions for physical deformation, and then input together with the real-time operating condition parameter set into a preset finite element (FEM) proxy model.

[0107] Furthermore, this proxy model can quickly map working parameters and strain data into structural response, calculate the dynamic displacement changes of all grid nodes in the 3D spatial grid, compare them with the original coordinates under the static datum, calculate the position difference in the 3D direction for each node, and statistically integrate the results to obtain the deformation offset of the entire equipment at the current moment. .

[0108] Furthermore, real-time data from the rack vibration sensors is acquired, and the dominant frequency and amplitude characteristics are extracted. Simultaneously, optical flow tracing of feature points in consecutive frames of the monitoring video can be performed using visual odometry to obtain the minute offset trajectories of these feature points.

[0109] Furthermore, by combining the vibration frequency domain analysis results with the pixel offset data of the visual odometry, a camera extrinsic parameter offset model was established. The minute translational and angular deflection errors of the camera mounting bracket or lens caused by the excitation of the unit operation were quantified, and finally these errors were converted into dynamic correction amounts for the camera extrinsic parameters. .

[0110] Furthermore, combining camera static extrinsic parameters and deformation offset... and dynamic correction amount Through the dynamic space mapping equation, the three-dimensional coordinates of a one-dimensional physical heat / excitation source can be represented. Projected onto the two-dimensional pixel plane, the following relationship (3) is shown: (3) In the formula: These are two-dimensional pixel plane coordinates, representing the coordinates in the plane. At any given moment, the physical source is mapped to the corresponding pixel location in the monitoring image. ; This represents the camera intrinsic parameter matrix, which is determined by the camera's own optical characteristics (such as focal length and principal point coordinates). It is used to convert three-dimensional points in the camera coordinate system into two-dimensional image coordinates. This represents the camera's static extrinsic parameter matrix, which may include the rotation matrix. Translation vector It describes the standard position and orientation of the camera relative to the device's global coordinate system in its initial installation state; This indicates the initial placement of physical sensors (such as vibration measuring points or infrared temperature measuring points) in the global three-dimensional space of the equipment.

[0111] Firstly, through Calculate the true instantaneous three-dimensional position of the sensor after the device deforms. Secondly, through... Correct for camera viewing angle deviations caused by environmental vibrations. Then, transform the corrected physical points from the world coordinate system to the camera coordinate system. Finally, use intrinsic parameters... Projected onto the image, forming the final pixel coordinates. That is, the coordinates of the center point of the mapping.

[0112] Furthermore, using the obtained mapping center point Centered on the vector, a spatial attention mask with a binary Gaussian distribution is constructed, as shown in the following relation (4): (4) In the formula: This represents the coordinates of any pixel within the mask's coverage area; Let represent the covariance matrix, and be an adaptive scaling factor whose magnitude is adaptively adjusted according to the energy intensity of the sensor signal frequency band. Specifically, when the energy measured by the sensor is very high, The corresponding increase expands the receptive field (focus area) of the visual encoder; when energy is concentrated in a specific high-frequency band, It shrinks down and achieves sub-pixel level precise locking.

[0113] Furthermore, the mask has the highest weight at the mapping point, and by continuously attenuating it towards the surrounding background area, it can form a pixel-level continuous weight matrix.

[0114] Furthermore, the mask serves to provide a precise crosshair for the visual encoder of the large language model, guiding it to eliminate irrelevant interference in the background (such as light reflections and water vapor fluctuations) and accurately locate the actual physical heat or vibration source area. At the same time, when the unit undergoes millimeter-level displacement due to load changes, the mask will drift in real time with the dynamic formula to ensure that the visual features and physical signals are always semantically aligned.

[0115] Step S2024: Determine the target visual feature set based on the binary Gaussian space attention mask and the initial visual feature set.

[0116] In an optional embodiment, a pixel-wise multiplication fusion is performed on the binary Gaussian spatial attention mask and the initial visual feature map to form a corresponding target visual feature map.

[0117] Furthermore, this weighted operation can retain and enhance the feature information of the high-weighted areas of the mask (corresponding to the physical heat / vibration source areas of the device), while suppressing the features of the low-weighted areas (such as the background of the computer room, invalid walls, water vapor noise, and other irrelevant areas), thereby eliminating irrelevant interference and achieving accurate focusing of the large model visual encoder.

[0118] Furthermore, the weighted visual features can be normalized and scaled, and the feature value distribution can be unified, which can avoid fluctuations in the feature value range caused by differences in mask weights and ensure the stability of subsequent model processing.

[0119] Step S2025: Based on the real-time operating condition parameter set, time-series signal set, and target visual feature map set, determine the multimodal perception dataset of the hydropower equipment.

[0120] In one optional embodiment, the real-time operating condition parameter set, the time-aligned temporal signal set, and the spatially optimized target visual feature map set are associated and bound together using the video exposure center time as a unified index, and integrated to form a corresponding multimodal perception dataset.

[0121] In some optional implementations, the PINN joint loss function of the neural network used to obtain the physical information of the fused rotor dynamics equations of the preset large model in step S203 above includes: Step c1: Obtain the simplified rotor dynamics differential equation with damping and gyroscopic effects, and the cross-entropy loss and mean square error loss of the pre-set large model.

[0122] Step c2: Extract latent variables representing the displacement or vibration trend of the axial system from the hidden layer of the preset large model, and use the automatic differentiation method of the deep learning framework to obtain the second and first derivatives of the latent variables with respect to time.

[0123] Step c3 involves inputting the hidden variables, second-order and first-order derivatives into the simplified rotor dynamics differential equation to obtain the physical residuals.

[0124] Step c4: Construct the joint loss function of the neural network PINN based on the cross-entropy loss, mean squared error loss, and physical residual.

[0125] In an optional embodiment, the simplified rotor dynamics differential equation is the motion equation of the turbine generator shaft system with damping and gyroscopic effect, which is used to determine whether the model output conforms to the laws of mechanical physics, as shown in the following relationship (5): (5) In the formula: This represents the mass matrix, used to characterize the rotor's mass and moment of inertia. This represents the damping matrix, used to characterize the energy dissipation characteristics of vibration. This represents the gyroscope matrix, used to characterize the rotational precession torque; This represents the stiffness matrix, used to characterize the ability to resist elastic deformation; It represents the displacement / vibration trend of the shaft system, i.e., the implicit variable; Represents velocity (first derivative); Represents acceleration (second derivative); This indicates the unbalanced force on the rotor; Indicates hydraulic excitation force; Represents physical residuals.

[0126] In one alternative embodiment, cross-entropy loss represents the fundamental loss of a large language model, used to ensure that the text / semantic symbols generated by the model are logically coherent and accurate.

[0127] In an optional embodiment, the mean squared error loss represents the health regression head loss, which measures the deviation between the model's predicted health and the true label.

[0128] In one optional embodiment, latent variables characterizing the displacement or vibration trend of the shaft system are extracted from the hidden layer of a pre-defined large model. Then, the first derivative is calculated using the deep learning framework Autograd. and second derivative .

[0129] Furthermore, the latent variables With derivative and Substitute the dynamic equation shown in the above relation (5) and calculate the physical residual. Furthermore, if If it is true, then it means that it conforms to the laws of physics; if This indicates that the prediction violates the laws of physics and generates a penalty gradient.

[0130] Furthermore, cross-entropy loss, mean squared error loss, and physical residual are integrated to construct the final PINN joint loss function for the neural network. The following relation (6) is shown: (6) In the formula: Represents cross-entropy loss; This represents the mean squared error loss; Represents the physical residual term; This represents the weighting coefficient for health loss; This represents the physical residual penalty coefficient.

[0131] Furthermore, when the predicted trend of the large model violates the basic mass-stiffness-damping characteristics of the device (i.e., When deviating from 0), This will generate a gradient penalty, driving the model learning process to converge to the physical law boundary.

[0132] In some optional implementations, step S205 above includes: Step S2051: Based on the multimodal sensing dataset, the temporal semantic symbol sequence is obtained by processing the temporal segmentation network based on the physical spectrum envelope.

[0133] In one alternative embodiment, the temporal segmentation network based on physical spectral envelope represents a 1D convolutional and FFT network that transforms one-dimensional high-frequency physical signals into LLM-understandable dynamic semantic symbols.

[0134] In one alternative embodiment, the temporal semantic symbol sequence represents a large model-compatible token sequence after the vibration / swing signal has been symbolized.

[0135] In an optional embodiment, by inputting high-frequency time-series sensing signals from multimodal sensing data into a time-series segmentation network based on physical spectrum envelope, one-dimensional continuous physical signals can be converted into a sequence of time-series semantic symbols that can be directly input into large models, thereby eliminating the semantic gap between industrial signals and LLM.

[0136] For example, after inputting the high-frequency time-series sensing signal from the multimodal sensing data into a time-series segmentation network based on the physical spectrum envelope, a one-dimensional residual network containing multi-scale dilated convolution is first used: parallel convolution kernels of different sizes (such as 1×64 and 1×256) are set to construct dual residual branches; using convolution kernels with different receptive fields, long-period mechanical friction modulation signals and high-frequency transient impact excitation signals are captured in parallel.

[0137] Then, the energy vectors of key dynamic frequency bands such as frequency conversion, harmonics, and 0.5 harmonics are extracted through the FFT fast Fourier transform operator layer.

[0138] Furthermore, after linear projection using a multilayer perceptron (MLP), these continuous numerical vectors can be discretized into a temporal semantic symbol sequence of dimension 1. .in, Indicates the length of the timing symbol; This indicates the feature dimension of the large model.

[0139] Step S2052: Based on the topological mask matrix, a multi-head cross-attention fusion network is used to process the multimodal perception dataset and the temporal semantic symbol sequence to obtain the multimodal fused semantic feature sequence.

[0140] In one alternative embodiment, the multi-head cross-attention fusion network represents a Transformer fusion module with topological mask constraints that achieves cross-modal alignment of text, vision, and temporal data.

[0141] In one optional embodiment, using a topological mask matrix as a physical constraint, a multi-head cross-attention fusion network is used to align and fuse three modalities—textual conditions, visual features, and temporal semantic symbols—in the same semantic space, thereby ultimately outputting a fused multimodal semantic feature sequence.

[0142] For example, the real-time operating text information in the multimodal perception dataset can be converted into a PromptToken, such as: "The current unit is operating in the high head, full load area, and the water guide bearing area exhibits the following spatiotemporal multimodal characteristics...".

[0143] Furthermore, the text token, visual token, and temporal token are concatenated and input into the Transformer layer inside the large model, while a topological mask matrix is ​​introduced into the multi-head cross-attention calculation. Then, through deep fusion of multiple layers of cross-attention and feedforward networks, a unified multimodal fused semantic feature sequence can be finally output.

[0144] Among them, visual tokens can be obtained by extracting features from the target region locked by the mask through the visual encoder of the large model and projecting them onto the semantic manifold space of the LLM; temporal tokens are temporal semantic symbol sequences.

[0145] In one example, a method for assessing the health of hydropower equipment by integrating a large linguistic model of physical information with multimodal feature alignment is provided. This method aims to alleviate the limitations of existing technologies through deep fusion of mechanisms and data, and is expected to achieve the following technical effects: 1. Constructing a high-precision spatiotemporal joint feature representation: A compensation algorithm considering flexible deformation and sub-pixel-level phase delay is proposed and combined with low-light image enhancement technology to significantly reduce the spatiotemporal alignment error of multimodal data under complex working conditions, providing a high-quality underlying data space for subsequent diagnosis.

[0146] 2. Achieve deep semantic mapping and joint reasoning for cross-modal features: Employ temporal symbolization technology to establish a mapping bridge between one-dimensional physical signals and high-dimensional semantic space, and provide health assessment with confidence intervals to improve the scientific accuracy of equipment degradation trajectory prediction.

[0147] 3. Effectively improve the physical consistency and interpretability of large model diagnosis: By introducing an anti-illusion reasoning module with physical constraints, the device transmission mechanism is transformed into partial differential equation (PDE) constraints and dynamic causal graphs. At the algorithm logic level, the "diagnosis illusion" of large models is effectively suppressed, making its reasoning process highly consistent with the basic laws of physics and significantly improving the engineering credibility of the diagnosis results.

[0148] Specifically, its core system consists of a bottom-layer multimodal data perception and spatiotemporal precise alignment module, a mid-layer multimodal semantic symbolization and fusion reasoning module, and a top-layer physical constraint anti-hallucination reasoning module. It includes the following core steps and module designs: 4.1 Multimodal data perception and spatiotemporal precise alignment module.

[0149] A hardware-aware and microsecond-level spatiotemporal dual compensation alignment mechanism is proposed. For large-scale mega-units (such as mixed-flow turbine generator units) under different head and load conditions, the complex electromechanical transients and hydrodynamic disturbances are addressed in this step by implementing high-fidelity data layer reconstruction to ensure the unification of spatiotemporal benchmarks for multi-source data.

[0150] 4.1.1 Edge-side clock synchronization and fractional delay filtering.

[0151] (1) Hardware clock synchronization: An industrial-grade edge computing gateway supporting the IEEE 1588 PTP v2 protocol is used to stamp nanosecond-level hardware timestamps on high-definition surveillance cameras (such as 25fps / 30fps) and high-frequency IEPE vibration / sway sensors (such as 10kHz~50kHz).

[0152] (2) Fractional delay filtering and resampling: The temporal resolution of video frames is much lower than that of high-frequency signals. The system uses the exposure center time of the video keyframe. As an absolute reference point, extract the set of discrete sensor sampling points corresponding to that moment. This is done to eliminate phase truncation errors due to sub-sampling periods. A fractional-order Sinc filter based on Hamming window truncation is constructed to resample the high-frequency sequence, as shown in the above relation (2).

[0153] Furthermore, this step precisely shifts the phase of the high-frequency signal to the same microsecond as the opening of the video shutter, eliminating "temporal artifacts" during fusion.

[0154] 4.1.2 Low-light image enhancement.

[0155] For the physical environment of a closed water chiller room with limited lighting and water vapor interference, a low-light image enhancement network based on Retinex theory and self-attention mechanism is constructed. Before alignment, the monitoring video frames are first subjected to brightness adaptive equalization and contrast stretching to highlight the surface texture of the equipment and extract high-dimensional visual features, so as to avoid the harsh lighting from masking the real thermal / deformation defect signals.

[0156] 4.1.3 Dynamic spatial mapping of strain compensation in flexible bodies.

[0157] During the operation of large rotating machinery (such as mixed-flow hydro-generator units), the unit is not in an absolutely rigid state. When the unit is under full load or experiences drastic fluctuations in operating conditions (such as changes in head and active power), the enormous fluid pressure and mechanical stress can cause sub-millimeter to millimeter-level low-frequency creep or high-frequency dynamic deformation in giant flexible bodies such as the top cover and frame. Traditional static spatial calibration models rely solely on pre-measured camera intrinsic and extrinsic parameters, failing to capture this physical displacement during operation. This results in a "semantic misalignment" between visual features in the image (such as surface oil mist and thermal infrared pixels) and the physical sensor positions (such as bearing vibration measurement points) in the spatial dimension. This solution upgrades the static mapping model to a dynamic real-time calibration model by introducing a real-time deformation compensation mechanism.

[0158] Step 1: Real-time sensing of multi-source operating conditions and structural strain. The system collects key variables affecting structural displacement in real time through an edge gateway: Operating Parameter Set These include real-time active power, guide vane opening, and head height, used to characterize the hydrodynamic load borne by the unit.

[0159] Strain data at critical locations: Real-time feedback data from strain gauges deployed at critical stress points on the frame are obtained as direct boundary conditions for physical deformation.

[0160] Step 2: Displacement calculation based on the finite element (FEM) proxy model. Since traditional finite element analysis (FEA) involves enormous computational costs and cannot meet real-time alignment requirements, this solution uses a pre-defined FEM proxy model: Computational mechanism: The surrogate model is based on a pre-trained response surface or neural network, which quickly maps the operating parameters C(t) to the structural response.

[0161] Output: Calculates the deformation offset vector of the unit's three-dimensional spatial mesh relative to the static datum at the current moment. .

[0162] Step 3: Camera dynamic disturbance compensation. Considering the impact of the industrial environment on the imaging equipment, calculate the real-time correction amount for the camera: Vibration compensation amount By using visual odometry or rack vibration frequency, the dynamic extrinsic parameter offset of the camera is corrected in real time to ensure that camera shake does not affect the accuracy of coordinate projection.

[0163] Step 4: Construct a dynamic spatial mapping equation system to map the three-dimensional physical coordinates of the physical sensors. Real-time conversion to two-dimensional pixel coordinates of the image plane As shown in the above relation (3).

[0164] Step 5: Generate adaptive spatial attention mask. After obtaining the accurate mapping center point, the system generates a mask at the corresponding position in the visual feature map to guide the large language model to focus: Binary Gaussian distribution construction: Construct a spatial attention mask M_spatial with the mapping point as the center.

[0165] Covariance adaptive scaling: The coverage area of ​​the mask (covariance matrix) is adaptively adjusted according to the energy intensity of the sensor signal frequency band, ensuring that the receptive field of the LLM visual encoder can accurately lock the real physical heat or vibration source area.

[0166] Among them, 1. The mathematical form of the mask: the binary Gaussian distribution system at the center point of the mapping At this point, a weight matrix based on a binary Gaussian distribution is constructed, as shown in the above relation (4).

[0167] 2. The core function of a mask.

[0168] Precisely pinpointing the receptive field: Through this mask, the visual encoder of the large model can eliminate irrelevant interference in the background (such as irrelevant light reflections and water vapor fluctuations) and precisely pinpoint the actual physical heat or vibration source area.

[0169] Solving the "spatiotemporal misalignment": Even if the unit undergoes millimeter-level displacement due to load changes, the mask will "drift" in real time with the dynamic formula to ensure that visual features and physical signals are always semantically aligned.

[0170] Guided depth alignment: It provides spatial priors for subsequent Transformer layers, enabling the model to logically associate "the color change of a certain pixel" with "a vibration signal of a certain frequency".

[0171] 3. Subsequent processing: Feature manifold fusion.

[0172] The generated mask is multiplied by the preprocessed visual feature map. Subsequently, these weighted visual tokens, along with the temporally symbolized (1D-Patch Embedding) physical signal tokens, are fed into the larger model. Time-series tokens contain dynamic features such as frequency conversion and frequency multiplication.

[0173] Text Token: Contains current operating status information (such as "Running at full capacity").

[0174] Visual Token: Defect image features after masking and weighting (such as blade cracks, bearing oil leakage).

[0175] 4.2 Multimodal semantic symbolization and fusion reasoning module.

[0176] We propose a cross-modal temporal symbolization and feature manifold mapping network to transform one-dimensional physical signals into high-dimensional semantic manifolds natively supported by large models (large language models with reasoning capabilities).

[0177] 4.2.1 Temporal Segmentation Network Based on Physical Spectrum Envelope (1D-Patch Embedding): Unlike 2D patches for visual images, this scheme constructs a one-dimensional residual network containing multi-scale dilated convolution. The first layer of the network uses convolution kernels of different sizes (e.g., and ), constructing parallel residual branches to capture long-period mechanical friction modulation signals and high-frequency transient impact excitation signals in parallel. Subsequently, energy vectors of key dynamic feature frequency bands such as 1 (frequency conversion), 2, and 0.5 are extracted through a Fast Fourier Transform (FFT) operator layer. After linear projection through a Multilayer Perceptron (MLP), these continuous numerical vectors are discretized into a temporal semantic symbol sequence of dimension . .

[0178] 4.2.2 Dynamic Prompt-Guided Multi-Head Cross-Attention Fusion: Real-time extracted textual operating condition information is converted into Prompt Tokens (e.g., "The current unit is operating in the high-head, full-load zone, and the water guide bearing area exhibits the following spatiotemporal multimodal features..."). Text Tokens, visual Tokens, and temporal Tokens are concatenated and input into the Transformer layer within the larger model. In Multi-head Cross-Attention, cross-modal feature alignment is achieved, enabling the model to physically correlate "pixel-level oil / water mist changes at the water guide bearing" with "high-frequency water pressure pulsation signals" in the feature manifold space.

[0179] 4.2.3 Bayesian Health Assessment and Confidence Measurement: In the regression head of the large language model, a Monte Carlo dropout (Drop rate = 0.1~0.3) mechanism is used. M forward propagations (e.g., M=30) are performed on the same set of input data, and the mean and variance of these M health output values ​​are calculated. The system ultimately outputs a 95% confidence interval to the operations and maintenance personnel. The diagnostic conclusions enhance the model's ability to anticipate risks under extreme conditions.

[0180] 4.3 Physically Constrained Anti-Hallucination Reasoning Module

[0181] We propose an anti-hallucination inference module based on DSTCG graph attention and PDE joint constraints. By embedding physical mechanisms as rigid constraints into LLM, we change the black-box characteristics of pure data fitting of large models and apply rigid industrial physical boundaries to ensure the physical consistency of inference.

[0182] 1. Dynamic Spatiotemporal Causal Graph (DSTCG) Attention Mask Constraint: Establishing the set of prior structural nodes for the unit. (Including: generator upper guide, thrust bearing, stator frame, lower guide, water guide bearing, impeller, etc.). The edge weights between nodes are dynamically calculated, integrating the static transfer function with data-driven instantaneous transfer entropy. For example, during the transient process at the initial stage of grid connection, the transfer weight of electromagnetic pull increases significantly; when the unit is in a steady state, the mechanical transfer weight dominates. The DSTCG is encoded using a Graph Attention Network (GAT) to generate a topological mask matrix of dimension . When calculating the attention weights in the LLM (multi-head cross-attention fusion process in section 4.2.2 above), this mask matrix forces the weights between nodes without direct physical connections to be zero, suppressing non-physical connections from the source.

[0183] Among them, DSTCG is based on the prior structure node set of the unit. The edge weights between nodes consist of: Prior structure node set First, based on the actual mechanical structure of the hydro-generator unit, a series of physically significant nodes were pre-designed. These nodes represent the unit's key transmission and support components.

[0184] Dynamic edge weights The "edges" between nodes are not fixed, but are calculated in real time by integrating static mechanisms and dynamic data.

[0185] Static transfer function Based on the mechanical model in the equipment design phase, define the inherent physical connections between components.

[0186] Instantaneous entropy transfer (data-driven): calculates the dynamic correlation between nodes by monitoring data in real time.

[0187] Furthermore, the ultimate goal of constructing DSTCG is to add a layer of "physical constraint" to the Large Language Model (LLM). Encoding DSTCG using a Graph Attention Network (GAT) generates a topological mask matrix. When the LLM calculates attention weights, this mask forces the weights of nodes with no direct physical connection (e.g., between the upper guide bearing and the tailrace pipe) to be set to zero. This prevents the model from generating associations that violate physical common sense (i.e., "diagnostic illusions") from the outset, ensuring that the inference path perfectly conforms to the boundaries of industrial physics.

[0188] 2. Physical Information Neural Network (PINN) Loss Function for Integrating Rotor Dynamics Equations: During the supervised fine-tuning (SFT) stage of the model, the latent variables mapped by the LLM output layer (characterizing the predicted shaft displacement or vibration trend) are extracted. A simplified rotor dynamics differential equation with damping and gyroscopic effects is introduced, as shown in the above relation (5).

[0189] Furthermore, by utilizing the automatic differentiation (Autograd) function of the deep learning framework, the second and first derivatives of the latent variables with respect to time are obtained, and the physical residuals are substituted into the above equation to calculate the final joint loss function, as shown in the above relation (6).

[0190] Furthermore, when the predicted trend of the large model violates the basic mass-stiffness-damping characteristics of the device (i.e., When deviating from 0), This will generate a gradient penalty, driving the model learning process to converge to the physical law boundary.

[0191] In one specific embodiment, taking the continuous health assessment and defect detection of a hydro-generator unit as an example, a detailed implementation process of a device health assessment method based on physical topology constraints and multimodal spatiotemporal alignment is provided. The system includes an edge computing perception layer, a feature alignment and symbolization layer, and a cloud-based LLM inference layer. The specific implementation steps are as follows: Step S1: Synchronous acquisition of multi-source heterogeneous data and nanosecond-level timestamp marking.

[0192] The system operates at the edge of the unit. High-definition monitoring cameras (sampling rate) are deployed in the turbine room and generator floor. Video streams are acquired; simultaneously, vibration and sway signals (sampling rate) are obtained using IEPE high-frequency sensors deployed on the water guide bearing, thrust bearing, and stator frame. All sensors are uniformly connected to an edge gateway that supports the IEEE 1588PTP v2 protocol, and each frame of image and each sampling point is stamped with an absolute timestamp at the nanosecond level.

[0193] Step S2: Low-light visual enhancement and multimodal defect feature preprocessing.

[0194] To address the harsh environment affecting turbine blades, guide bearings, and other areas with limited light and interference from oil mist / water vapor, the acquired monitoring video frames are first preprocessed. Based on Retinex theory, the original images... Decomposed into reflection components and light component As shown in the above relation (1).

[0195] Furthermore, by constructing a lightweight convolutional network that includes a spatial self-attention mechanism, the illumination component is dynamically balanced and smoothed. This allows for adaptive enhancement of visual defect features such as micro-cracks on the blade surface, cavitation erosion, or bearing oil seepage without amplifying background noise, and outputs an enhanced high-dimensional visual feature map.

[0196] Step S3: Microsecond-level time phase alignment based on fractional delay filtering.

[0197] Expose the center moment with enhanced video keyframes Based on this, the subsampling period phase truncation error of the high-frequency vibration signal at this moment is addressed. The system activates a fractional-order Sinc filter based on Hamming window truncation at the edge to resample the discrete sampling sequence of the sensor, as shown in the above relation (2).

[0198] This allows for precise shifting of the time-domain phase of a one-dimensional high-frequency signal to the same microsecond as the opening of the video shutter, eliminating "temporal artifacts" during cross-modal fusion.

[0199] The enhanced video keyframe is a direct product of step S2, and it is composed of reflection components. With optimized light components The result of recombination (or feature fusion).

[0200] Further, step S2 involves processing and The component transforms the "unclear" image into a "clear" one; step S3 then uses this clear image as a time scale to "correct" the phase of the high-frequency sensor, eliminating time artifacts.

[0201] Furthermore, resampling the discrete sampling sequence of the sensor specifically refers to a fractional-order delay filtering process. Its core purpose is to solve the microsecond-level temporal phase misalignment problem between high-frequency sensors (such as vibration signals of 10kHz-50kHz) and low-frequency video (such as 30fps). This is because the exposure center time of the video shutter... Often, the sampling point falls between two discrete sensor sampling points. Directly taking the nearest point will lead to "phase truncation error," while resampling uses an interpolation algorithm to accurately calculate the phase truncation error. The signal value corresponding to the given time.

[0202] Step S4: Fuse subpixel-level spatial mapping with dynamic deformation compensation.

[0203] Real-time acquisition of unit operating condition parameters (Active power, guide vane opening, head, etc.), input the preset finite element (FEM) proxy model, and calculate the three-dimensional mesh creep offset of the giant frame and top cover under multiphysics coupling. .

[0204] Furthermore, combining camera static extrinsic parameters and vibration compensation amount The three-dimensional coordinates of the one-dimensional physical heating / excitation source are projected onto the two-dimensional pixel plane through the dynamic mapping equation, as shown in the above relation (3).

[0205] Furthermore, the above method ensures that the visual encoder of the large model can accurately locate the "real physical heat / vibration source area", avoiding the failure problem of traditional static mapping under varying working conditions.

[0206] Furthermore, a two-dimensional Gaussian spatial attention mask is generated at the center point of the mapping to ensure that the LLM visual encoder accurately locks the target region and achieves sub-pixel-level spatial alignment.

[0207] Step S5: Symbolization and manifold fusion of cross-modal temporal physical signals.

[0208] The aligned high-frequency physical signals are input into a one-dimensional temporal segmentation network (1D-Patch Embedding). Through multi-scale dilated convolution and Fast Fourier Transform (FFT), key frequency band energy vectors related to water pressure pulsation and rotor rotation are extracted and mapped into a temporal semantic token sequence of dimension 1 using an MLP. Subsequently, natural language condition prompts (PromptTokens), visual defect tokens, and temporal tokens are fed into a multi-head cross-attention network of an LLM to complete the semantic joint mapping of heterogeneous data in the manifold space.

[0209] Step S6: Anti-hallucination diagnostic reasoning based on dual physical constraints.

[0210] During the LLM's inference diagnostic process, rigid physical boundary constraints are applied: 1. Topological attention constraint: Based on the shaft transfer function and transient characteristics of the generator set, a mask matrix between nodes of the spatiotemporal causal graph (DSTCG) is dynamically generated. In the attention calculation, the weights between nodes without direct physical transmission paths (such as the upper guide bearing and the tailrace pipe) are forced to zero.

[0211] 2. Differential Equation Optimization Constraints: Extract the predicted displacement vector x represented by the LLM hidden layer, substitute it into the rotor dynamics equation with damping C and gyroscope matrix G, i.e., the above relation (5), and calculate the physical residual.

[0212] Furthermore, if A value not equal to 0 indicates that the trend predicted by LLM violates the fundamental physical law of "mass-stiffness-damping". This residual will be added to the loss function as a penalty term, forcing the model to correct the "logic illusion" and make its predictions conform to the industrial physical boundaries.

[0213] Step S7: Bayesian uncertainty quantification and health assessment report generation.

[0214] Monte Carlo Dropout (MC Dropout) is enabled in the final regression head of the LLM. Thirty random forward propagation calculations are performed with the current input state, and the mean and variance of the health score are calculated. The final output is a physically interpretable diagnostic report text with a 95% confidence interval. The health metrics are fed back to the monitoring dashboard and the operation and maintenance decision-making system.

[0215] The method for assessing the health of hydropower equipment, which integrates a large physical information language model with multimodal feature alignment, provided in this example has the following beneficial effects: (1) A microsecond-level spatiotemporal dual compensation alignment mechanism for giant flexible units is proposed: Unlike traditional static calibration, this example integrates fractional delay filtering and dynamic deformation compensation based on FEM proxy model, combined with preprocessing enhancement of low-light images, to achieve microsecond-level time synchronization and sub-pixel-level spatial alignment of visual and high-frequency sensing data under complex variable working conditions, fundamentally solving the problem of physical misalignment of multimodal data and feature loss under harsh working conditions.

[0216] (2) A physical-inspired cross-modal temporal symbolization and manifold mapping method is proposed: In response to the semantic gap in LLM processing of industrial signals, a one-dimensional symbolization network containing multi-scale dilated convolution and physical frequency band attention module is designed to transform the original physical signal into a token sequence rich in dynamic semantics, thereby realizing a high-fidelity mapping between physical signals and LLM semantic space.

[0217] (3) A dual-physical-constrained anti-hallucination reasoning framework integrating DSTCG and PINN was constructed: the topological mask generated by the Dynamic Spatiotemporal Causal Graph (DSTCG) constrains the inference path of LLM at the attention level; at the same time, the rotor dynamics equation constraint is embedded at the optimization objective level through the loss function of the Physical Information Neural Network (PINN). The two work together to ensure the physical consistency and interpretability of the model inference from different dimensions.

[0218] (4) A Bayesian health assessment system with uncertainty quantification capability was established: the Monte Carlo random inactivation (MC Dropout) mechanism was integrated into the LLM regression head to improve the point estimation output into a probability distribution and output the health status with confidence interval, which significantly enhanced the engineering credibility and decision support value of the assessment conclusion under extreme and rare working conditions.

[0219] This embodiment also provides a hydroelectric equipment health assessment device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0220] This embodiment provides a device for assessing the health of hydroelectric equipment, such as... Figure 3 As shown, the device includes: The first acquisition module 301 is used to acquire the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of hydropower equipment based on a unified time reference.

[0221] The first processing module 302 is used to process the initial monitoring video stream, high-frequency sensor signal set and real-time operating condition parameter set through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain a multimodal sensing dataset of hydropower equipment.

[0222] The second acquisition module 303 is used to acquire the dynamic spatiotemporal causal graph of the hydropower equipment and the physical information of the fused rotor dynamics equation of the preset large model using a neural network PINN joint loss function.

[0223] The encoding module 304 is used to encode dynamic spatiotemporal causal graphs using a graph attention network and generate a topological mask matrix.

[0224] The second processing module 305 is used to process the multimodal perception dataset based on the topological mask matrix and using cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fused semantic feature sequence.

[0225] The third processing module 306 is used to input the multimodal fused semantic feature sequence into a preset large model and process it using the Monte Carlo random deactivation mechanism to obtain the device health assessment result containing the confidence interval that satisfies the PINN joint loss function of the neural network.

[0226] The hydroelectric equipment health assessment device provided in this embodiment of the invention can execute the hydroelectric equipment health assessment method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.

[0227] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0228] The following is a detailed reference. Figure 4 This diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from memory 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the electronic device. The processor 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0229] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0230] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a memory 408, or installed from a ROM 402. When the computer program is executed by the processor 401, it performs the functions defined in the hydroelectric equipment health assessment method of the embodiments of the present invention.

[0231] Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.

[0232] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the hydroelectric equipment health assessment method shown in the above embodiments is implemented.

[0233] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0234] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for assessing the health of a hydroelectric facility, characterized in that, The method includes: Based on a unified time reference, the initial monitoring video stream, high-frequency sensor signal set, and real-time operating condition parameter set of hydropower equipment are acquired. The initial monitoring video stream, the high-frequency sensor signal set, and the real-time operating condition parameter set are processed through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain the multimodal sensing dataset of the hydropower equipment. The PINN joint loss function is used to obtain the physical information of the dynamic spatiotemporal causal graph of the hydropower equipment and the fused rotor dynamics equation of the preset large model. The dynamic spatiotemporal causal graph is encoded using a graph attention network, and a topological mask matrix is ​​generated. Based on the topological mask matrix, the multimodal sensing dataset is processed using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fusion semantic feature sequence. The multimodal fused semantic feature sequence is input into the preset large model and processed using the Monte Carlo random deactivation mechanism to obtain the device health assessment result containing the confidence interval that satisfies the PINN joint loss function of the neural network.

2. The method of claim 1, wherein, The initial monitoring video stream, the high-frequency sensor signal set, and the real-time operating condition parameter set are processed through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain the multimodal sensing dataset of the hydropower equipment, including: A low-light image enhancement network based on Retinex theory and self-attention mechanism is used to perform brightness equalization and contrast stretching on the initial monitoring video stream to obtain an initial visual feature map set. Using the exposure center time of the video keyframe corresponding to the initial visual feature map set as a reference, the high-frequency sensor signal set is resampled using a fractional Sinc filter truncated by a Hamming window to obtain a time-series signal set. Based on the real-time operating condition parameter set, displacement calculation is performed through a preset finite element proxy model, and a binary Gaussian spatial attention mask is constructed. Based on the binary Gaussian space attention mask and the initial visual feature map set, the target visual feature map set is determined; Based on the real-time operating condition parameter set, the time-series signal set, and the target visual feature map set, the multimodal sensing dataset of the hydropower equipment is determined.

3. The method of claim 2, wherein, A low-light image enhancement network based on Retinex theory and self-attention mechanism is used to perform brightness equalization and contrast stretching on the initial monitoring video stream to obtain an initial visual feature map set, including: The video frames in the initial monitoring video stream are preprocessed to obtain the target monitoring video stream; Using the Retinex theory, the image in the target surveillance video stream is decomposed to obtain multiple initial reflection components and multiple initial illumination components; Based on a lightweight convolutional network containing a spatial self-attention mechanism, the multiple initial illumination components are dynamically balanced and smoothed to obtain multiple target illumination components. The initial visual feature map is generated based on multiple target illumination components and multiple initial reflection components.

4. The method of claim 2, wherein, Based on the real-time operating condition parameter set, displacement calculation is performed using a preset finite element proxy model, and a binary Gaussian spatial attention mask is constructed, including: Based on the real-time operating condition parameter set, displacement calculation is performed using the preset finite element proxy model to obtain the deformation offset of the three-dimensional spatial mesh of the hydropower equipment relative to the static reference at the current moment. The camera vibration compensation amount is determined by using visual odometry or frame vibration frequency. Based on the deformation offset and the camera vibration compensation, the coordinates of the mapping center point are obtained after processing by the dynamic spatial mapping equation. Based on the coordinates of the center point of the mapping, the binary Gaussian spatial attention mask is constructed.

5. The method of claim 1, wherein, Based on the aforementioned topological mask matrix, the multimodal sensing dataset is processed using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fused semantic feature sequence, including: Based on the multimodal sensing dataset, a temporal semantic symbol sequence is obtained after processing by a temporal segmentation network based on physical spectrum envelope; Based on the topological mask matrix, the multimodal perception dataset and the temporal semantic symbol sequence are processed using a multi-head cross-attention fusion network to obtain the multimodal fused semantic feature sequence.

6. The method of claim 1, wherein, The PINN joint loss function for obtaining physical information from the fused rotor dynamics equations of a pre-defined large model includes: Obtain the simplified rotor dynamics differential equation with damping and gyroscopic effects, and the cross-entropy loss and mean square error loss of the preset large model; In the hidden layer of the preset large model, latent variables representing the displacement or vibration trend of the axial system are extracted, and the second and first derivatives of the latent variables with respect to time are obtained using the automatic differentiation method of the deep learning framework. The hidden variables and the second and first derivatives are input into the simplified rotor dynamics differential equation to obtain the physical residuals. The PINN joint loss function of the neural network is constructed based on the cross-entropy loss, the mean squared error loss, and the physical residual.

7. A hydroelectric plant health assessment device, comprising: The device includes: The first acquisition module is used to acquire the initial monitoring video stream, high-frequency sensor signal set and real-time operating condition parameter set of hydropower equipment based on a unified time reference. The first processing module is used to process the initial monitoring video stream, the high-frequency sensor signal set, and the real-time operating condition parameter set through a microsecond-level spatiotemporal dual compensation alignment mechanism to obtain the multimodal sensing dataset of the hydropower equipment. The second acquisition module is used to acquire the dynamic spatiotemporal causal graph of the hydropower equipment and the PINN joint loss function of the physical information of the fused rotor dynamics equation of the preset large model. The encoding module is used to encode the dynamic spatiotemporal causal graph using a graph attention network and generate a topological mask matrix. The second processing module is used to process the multimodal perception dataset based on the topological mask matrix using a cross-modal temporal symbolization and feature manifold mapping network to obtain a multimodal fusion semantic feature sequence. The third processing module is used to input the multimodal fused semantic feature sequence into the preset large model and process it using the Monte Carlo random deactivation mechanism to obtain the device health assessment result containing the confidence interval that satisfies the PINN joint loss function of the neural network.

8. An electronic device, comprising: include: A memory and a processor are interconnected, the memory storing computer instructions, and the processor executing the computer instructions to perform the hydroelectric equipment health assessment method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the hydroelectric equipment health assessment method according to any one of claims 1 to 6.

10. A computer program product, characterised in that, Includes computer instructions for causing a computer to perform the hydroelectric equipment health assessment method according to any one of claims 1 to 6.