Ultrahigh voltage metering data intelligent acquisition method based on multi-modal fusion

By combining multimodal fusion and deep learning models, the problems of insufficient accuracy and stability in ultra-high voltage power grid metering data acquisition are solved, and highly reliable metering data acquisition in complex environments is achieved.

CN121577957APending Publication Date: 2026-02-27MAINTENANCE & TEST CENTRE CSG EHV POWER TRANSMISSION CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511755380.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies lack multi-modal data fusion in ultra-high voltage power grid metering data acquisition, resulting in insufficient accuracy and stability of metering data, which cannot meet high reliability requirements.

Method used

A multimodal fusion intelligent acquisition method is adopted, including the acquisition and preprocessing of electrical signals, acoustic signals and image data. A feature set is generated through a deep learning model, and a modified PerceiverIO model is used for multimodal fusion and physical constraints. A Transformer network is combined for temporal interactive fusion, and finally calibration is performed at the edge and center.

Benefits of technology

It improves the accuracy and stability of measurement data, enhances the robustness of the system, maintains the reliability of acquisition results in complex environments, and dynamically adapts to equipment aging or environmental changes through a collaborative calibration mechanism between the edge and the center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121577957A_ABST
    Figure CN121577957A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric power measurement, and discloses an intelligent ultrahigh voltage measurement data acquisition method based on multi-modal fusion, which comprises the following steps: acquiring on-site electrical, acoustic and image data, extracting features through a deep learning model, inputting the data into an improved general architecture PerceiverIO for core fusion, introducing a physical constraint mechanism in a fusion process, and performing multi-modal fusion on the fusion process to obtain a fusion result; a self-adaptive modal gating mechanism is applied, modal weight is dynamically adjusted according to data quality to enhance robustness, a fused high-dimensional sequence is subjected to time sequence interaction through a Transform network, comprehensive feature representation is generated, a collection result is generated through classification and regression operation, an edge node corrects errors in real time through a Kalman filter, and a high-dimensional image is obtained. And the central database carries out long-term calibration and issuing on the metering coefficient. According to the method, the problems of single information, low precision and poor long-term stability of a traditional method are solved, and high-precision and high-robustness intelligent data acquisition is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of electric power metering, in particular to a kind of superhigh voltage metering data intelligent acquisition method based on multi-modal fusion. BACKGROUND

[0002] With the rapid development of superhigh voltage power grid, the accuracy and real-time of metering data play an important role in equipment operation monitoring and state evaluation, and the prior art relies on single modal data acquisition method, and voltage, current and other electrical quantities are collected by electrical signal measurement equipment, or auxiliary diagnosis is carried out by image and acoustic monitoring device.

[0003] However, the existing method lacks the fusion of multi-modal data, cannot simultaneously consider electrical quantities, acoustic characteristics and image information, and has deficiencies in long time sequence dependent modeling, noise interference suppression and physical consistency constraints, which can easily lead to the decrease of the precision of metering data and the lack of stability, and it is difficult to meet the high reliability requirement of metering and operation monitoring in superhigh voltage scene.Therefore, the present application provides a kind of superhigh voltage metering data intelligent acquisition method based on multi-modal fusion to solve the deficiencies of prior art. SUMMARY

[0004] In view of the deficiencies of the prior art, the present application provides a kind of superhigh voltage metering data intelligent acquisition method based on multi-modal fusion, which solves the problem that electrical quantities, acoustic characteristics and image information cannot be considered simultaneously, and has deficiencies in long time sequence dependent modeling, noise interference suppression and physical consistency constraints.

[0005] To achieve the above purpose, the present application is realized by the following technical scheme: a kind of superhigh voltage metering data intelligent acquisition method based on multi-modal fusion, characterized by comprising the following steps:

[0006] S1, acquisition and preprocessing, collect multi-modal original data of superhigh voltage metering field, and preprocess the multi-modal original data to generate standardized multi-modal data set, the multi-modal original data includes electrical signal, acoustic signal and image data;

[0007] S2, feature extraction and set generation, the standardized multi-modal data set is input into deep learning model to generate feature set, the feature set includes electrical signal time sequence feature vector, acoustic feature vector and image feature vector;

[0008] S3, multi-modal fusion, the feature set is input into improved universal architecture PerceiverIO, the improved universal architecture PerceiverIO carries out physical constraint mechanism and self-adapting modal gate mechanism and passes through multi-granularity output interface, generates high-dimensional semantic feature sequence;

[0009] S4, time sequence interaction fusion, inputting the high-dimensional semantic feature sequence to a multi-head self-attention layer of a Transformer network to generate a multi-modal comprehensive feature representation;

[0010] S5, result generation, performing classification and regression operation on the multi-modal comprehensive feature representation to generate a collection result;

[0011] S6, edge correction and center calibration, transmitting the collection result to an edge computing node for real-time error correction and uploading to a central database for calibration update of a measurement coefficient.

[0012] The application provides a multi-modal fusion-based ultrahigh voltage measurement data intelligent collection method.

[0013] 1. The application fuses electrical signals, acoustic signals and image data in the ultrahigh voltage measurement field, compared with a single data source, can perceive the device state from multiple dimensions, and through an adaptive modal gating mechanism, can dynamically evaluate the quality of each modal data, automatically reduce the weight when the sensor data is disturbed or the quality is reduced, ensure the stability and reliability of the collection result in complex or harsh environments, and enhance the robustness of the system.

[0014] 2. In the multi-modal fusion process, the application introduces a physical constraint mechanism, decodes the internal voltage and current representation, calculates the theoretical power, and compares it with the reference value, so that the physical law is introduced as a loss function constraint into the training process of the deep learning model, forcing the model to learn not only the feature representation that fits the data, but also the real physical law, avoiding incorrect results that contradict physical common sense, and improving the accuracy and reliability of the electrical measurement value regression result.

[0015] 3. The application constructs an edge and center collaborative closed-loop calibration system, uses a Kalman filter in the edge computing node to fuse the collection result with the original measurement value in real time, can correct the instantaneous error, the central database compares the corrected result with the high-precision standard model to continuously calibrate and update the measurement coefficient, and the updated coefficient is sent to the edge node, the correction and long-term calibration mechanism can dynamically adapt to device aging or environmental changes, and ensures the measurement ability in long-term operation. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is the overall flowchart of the method of the application;

[0017] Figure 2 is a schematic diagram of the improved PerceiverIO internal structure of the application;

[0018] Figure 3A multi-modal data feature extraction flowchart of the present application;

[0019] Figure 4 A Transformer time sequence interaction fusion schematic diagram of the present application;

[0020] Figure 5 An edge center collaborative closed-loop calibration system schematic diagram of the present application. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0022] Referring to the accompanying drawings Figure 1 The embodiments of the present application provide a multi-modal fusion based ultrahigh voltage metering data intelligent acquisition method: acquisition and preprocessing, acquiring multi-modal original data of an ultrahigh voltage metering site, the multi-modal original data including electrical signals, acoustic signals and image data; then preprocessing the multi-modal original data to generate a standardized multi-modal data set;

[0023] Feature extraction and set generation, inputting the standardized multi-modal data set into a deep learning model, the deep learning model including a PatchTST model (a time sequence prediction model), a wav2vec model (a speech representation model based on self-supervised learning) and a BEiT model (a self-supervised visual representation model); modeling time sequence features of the electrical signals, extracting convolutional features of the image data and converting spectral features of the acoustic data through the deep learning model, respectively generating feature vectors and combining to generate a feature set;

[0024] Multi-modal fusion, inputting the feature set into an improved universal architecture PerceiverIO (a universal multi-modal learning model), performing multi-modal feature alignment fusion, physical constraints and modal dynamic adjustment through the improved universal architecture to generate a high-dimensional semantic feature sequence;

[0025] Time sequence interaction fusion, inputting the high-dimensional semantic feature sequence into a Transformer network (a deep learning network architecture), calculating the correlation in the time dimension and performing interaction fusion through the Transformer network to generate a multi-modal comprehensive feature representation;

[0026] Result generation, performing classification and regression operations on the multi-modal comprehensive feature representation, and combining the classification results and the regression results to generate an acquisition result;

[0027] The edge correction and center calibration transmit the collection results to the edge computing node, and perform real-time error correction and noise reduction processing, and then calibrate and update the measurement coefficient of the device in the center database.

[0028] The data collection operation specifically refers to obtaining various data capable of reflecting the core working conditions, mechanical and insulation states, and external and thermal states of the ultrahigh voltage device in the running process. The multi-modal original data includes electrical signals, acoustic signals and image data.

[0029] The collection of electrical signals refers to obtaining voltage signals measured by voltage transformers, current signals measured by current transformers, and temperature signals of key components of the device measured by temperature sensors.

[0030] The collection of acoustic signals refers to obtaining partial discharge sound, mechanical vibration sound and environmental noise signals generated during device operation.

[0031] The collection of image data refers to obtaining appearance images and running state video frames of the device captured by visible light cameras, and surface defect images captured by infrared thermographs and the like.

[0032] After completing the data collection operation, a preprocessing operation is performed:

[0033] Time synchronization processing configures a time source for all data collection actions.

[0034] Noise filtering processing applies a digital low-pass filter to filter high-frequency noise or applies a notch filter to eliminate specific frequency power harmonic interference for electrical signals. A band-pass filter is applied to retain specific frequency bands and filter out low-frequency environmental noise for acoustic signals. Median filtering or Gaussian smoothing algorithm is applied to reduce image sensor noise for image data.

[0035] Missing value completion processing can use linear interpolation, forward filling or backward filling method to calculate and fill the missing data points for electrical signals and acoustic signals with time sequence continuity.

[0036] Standardization processing can specifically use Z-Score standardization method to subtract the sequence mean of signal values and then divide by the sequence standard deviation, or can use Min-Max standardization method to linearly scale the signal values to a fixed interval.

[0037] Refer to the attached Figure 3For electrical signals, the standardized electrical signals are constructed into an input sequence in chronological order, and the input sequence is divided into multiple continuous and non-overlapping time segments in the time dimension. Vectorization encoding operation is performed on the time segments to convert them into fixed-dimensional feature vectors. The sequence of feature vectors of all time segments is input into the PatchTST model. The PatchTST model captures the dependent features and patterns of electrical signals at different time scales, aggregates all processed time segments, and generates the final electrical signal time sequence feature vector.

[0038] For acoustic signals, the standardized acoustic signals are preprocessed. The continuous acoustic waveform signal is converted into a spectrogram by a short-time Fourier transform method, and the spectrogram is framed and segmented. Then, feature encoding is performed on the acoustic signals, which specifically includes frequency band division of the spectrum and calculation of energy values in the frequency band, and vector representation of the energy values and phase dimensions. The vector representation sequence is input into the wav2vec model. The wav2vec model first extracts local acoustic patterns through convolution layers, then performs feature encoding on the acoustic signals through its context modeling layer to calculate the correlation between different frames, and finally aggregates the feature representations of all frames to generate the final acoustic feature vector.

[0039] For image data, the input standardized image data is processed by blocking. A single image is divided into a grid-shaped, non-overlapping image block sequence. Vectorization encoding is performed on the image blocks. The mask image modeling method is used to randomly select a portion of the image blocks according to a predetermined proportion, and the image block vector representation is replaced with a mask placeholder. The sequence of vector representations of all image blocks is input into the BEiT model. The BEiT model learns feature representations of all image blocks through a Transformer encoder and calculates spatial correlations between different image blocks using a multi-head self-attention mechanism. The BEiT model aggregates the feature representations of all image blocks to generate the final image feature vector. The electrical signal time sequence feature vector, acoustic feature vector, and image feature vector are combined to form a feature set.

[0040] Referring to the accompanying Figure 2 The multi-modal fusion receives the feature set. The first improvement of the improved PerceiverIO model is the construction of a hierarchical latent space. The hierarchical latent space includes a long-time sequence latent space and a high-frequency transient latent space.

[0041] Construction of the long-time sequence latent space: define a long-time sequence latent array; extract input features covering an extended time window from the generated feature set; input the input features into the long-time sequence latent array to extract relevant time-dependent features and generate long-time sequence latent representations.

[0042] High-frequency transient latent space construction: define a high-frequency transient latent array; from the feature set, extract input features covering a short time window; input the input features into the high-frequency transient latent array to extract rapidly changing features within a local range related to phenomena such as transient pulses and burst impacts, and generate high-frequency transient latent representations.

[0043] After generating the long-term latent representation and the high-frequency transient latent representation, the two are combined to obtain an initial latent representation, which will be used as input for subsequent feature alignment and updating operations in the improved PerceiverIO model.

[0044] Performing a feature alignment operation takes the initial latent representation as the source of the query; the generated feature set is used as the source of the key and value together. In the operation process, the initial latent representation is transformed to generate a query matrix, while the feature set is transformed to generate a key matrix and a value matrix. The dot product of the query matrix and the transpose of the key matrix is calculated to obtain a relevance score matrix. The relevance score matrix is scaled and normalized by applying the Softmax function along a specific dimension to convert it into an attention weight matrix. The sum of the weight values in the attention weight matrix is 1. The attention weight matrix and the value matrix are multiplied to output the updated latent representation. The updated latent representation maintains the same dimensions as the initial latent representation but incorporates weighted feature information from the modalities. The feature alignment operation can be repeated multiple times in the improved PerceiverIO model, i.e., as a processing block. The last generated updated latent representation serves as the query source for the next execution, enabling iterative refinement and deep fusion of features.

[0045] Refer to the attached Figure 2 The second improvement in the improved PerceiverIO model is the physical constraint mechanism:

[0046] The physical constraint mechanism is activated during the training phase of the improved PerceiverIO model. The calculation process of the physical constraint loss term is as follows:

[0047] From the features output by the improved PerceiverIO model, decode or map specific physical quantity representations related to electrical quantities. According to the basic physical formula of electrical power, multiply the voltage representation and the current representation element by element to obtain a theoretically calculated power representation. Obtain a reference power representation for comparison. Calculate the difference between the theoretically calculated power representation and the reference power representation by calculating the mean square error or other distance measurement functions to obtain the physical constraint loss term.

[0048] In each iteration of the PerceiverIO model training, the physical constraint loss term and the standard loss function are weighted and summed according to a preset weight to form a total loss function, wherein the standard loss function is a basic function for measuring the difference between the model prediction result and the true label, which is composed of two parts, namely the cross-entropy loss function for evaluating the performance of the classification task of the model and the mean square error loss function for evaluating the performance of the regression task of the model. By conducting the gradient generated by the total loss function back to the model parameters in the back propagation process, the PerceiverIO model is forced to not only make the prediction result approach the true label, but also make the generated feature representation of voltage and current meet the physical law of electric power when combined during the learning process.

[0049] Referring to the drawings Figure 3 The third improvement in the improved PerceiverIO model is an adaptive modal gating mechanism.

[0050] Perform modal quality assessment to assess the quality of the original data constituting the feature set and calculate one or more quality indicators for the modal.

[0051] Perform gating weight calculation, combine the calculated quality indicators of all modes into a quality feature vector, input the quality feature vector into the gating unit, and perform nonlinear transformation on the quality feature vector. The gating unit can be a small neural network, which outputs the gating weight through the Softmax activation function layer. The weight value in the gating weight corresponds to the mode, and the sum of all weight values is 1.

[0052] Perform feature weighting operation, multiply the electrical signal time series feature vector, acoustic feature vector and image feature vector in the feature set with the corresponding gating weight element by element, adjust the feature set after the feature weighting operation, and then enter the subsequent feature alignment and fusion process to generate a high-dimensional semantic feature sequence.

[0053] Referring to the drawings Figure 2 The fourth improvement in the improved PerceiverIO model is a multi-granularity output interface.

[0054] The coarse-grained output query array is a set of learnable vectors optimized during training to probe and extract high-level and general features from the final latent representation. These features correspond to the overall running state of the super-high voltage equipment, the statistical trend of key parameters or the macro classification of environmental conditions within a long time span.

[0055] The fine-grained output query array is another set of learnable vectors optimized during training to probe and extract low-level and detailed features from the final latent representation of the improved PerceiverIO model.

[0056] In execution, the coarse-grained output query array and the fine-grained output query array are executed as two parallel cross-attention queries, respectively, which jointly use the final latent representation generated by the improved PerceiverIO model after the main processing procedure as the key and the value.

[0057] Through cross-attention operation, the coarse-grained output query array generates a coarse-grained feature vector; the fine-grained output query array generates a fine-grained feature vector; the generated coarse-grained feature vector sequence and the fine-grained feature vector sequence are combined to form a high-dimensional semantic feature sequence, which contains the fused multi-modal information and encodes the features of different time scales and different semantic levels in the internal structure.

[0058] Referring to the accompanying drawings Figure 4 The time sequence interaction fusion receives the generated high-dimensional semantic feature sequence as input, inputs the high-dimensional semantic feature sequence into a standard Transformer network, and performs a position encoding operation on the input sequence before entering the main part of the Transformer network. The position encoding is realized by adding pre-computed sine and cosine function values related to the position to the corresponding feature vectors.

[0059] After completing the position encoding, the sequence is input into the multi-head self-attention layer of the Transformer network. In the multi-head self-attention layer, the input feature sequence is sent into multiple independent heads. In each head, according to the input feature sequence, a query matrix, a key matrix and a value matrix are generated through three independent linear transformations, respectively. Then, the dot product of the query matrix and the key matrix is calculated to obtain an attention score matrix. Next, the attention score matrix is scaled and Softmax normalized to convert it into an attention weight matrix. Finally, the attention weight matrix and the value matrix are multiplied. For each time step in the sequence, the feature representation of itself is recalculated as the weighted sum of all other time step features in the entire time sequence.

[0060] Since it is multi-head attention, the above process occurs in multiple heads in parallel, and each head uses different linear transformation parameters. All output results are spliced together and then linearly transformed to form the final output of the multi-head self-attention layer. After residual connection and layer normalization, it is sent to the feedforward neural network. The feedforward neural network performs independent nonlinear transformation on the sequence, and the output is also subjected to residual connection and layer normalization.

[0061] The complete structure of the multi-head self-attention and the feedforward network constitutes a Transformer encoder layer. Multiple such encoder layers can be stacked to form a deeper network. After processing by the entire Transformer network, the final output is a multi-modal comprehensive feature representation, which is a feature sequence of the same length as the input sequence.

[0062] Referring to the accompanying drawings Figure 1 The multi-modal comprehensive feature representation output from the Transformer network is a sequence of feature vectors. An aggregation operation is performed on the sequence of feature vectors to extract the feature vector corresponding to the last time step, which has already contained the context information of the entire sequence under the self-attention mechanism of the Transformer. Alternatively, a global average pooling operation can be performed on all feature vectors of the entire sequence to obtain an aggregated feature vector that can also represent global information. After obtaining the aggregated feature vector, the following task branches are executed:

[0063] Classification task branch: The aggregated feature vector is input into one or more independent classification heads. The classification head is composed of one or more fully connected layers, and the output dimension of the last layer is equal to the total number of classes of the classification task. For example, for a classification head used for device state evaluation, the output dimension is 3, corresponding to the original scores of the normal, warning, and fault classes, respectively. Then the original scores are input into a Softmax activation function to convert them into a probability distribution, and the class with the highest probability is selected as the final classification result of the task.

[0064] Regression task branch: The aggregated feature vector is input into one or more independent regression heads. Each regression head is also composed of one or more fully connected layers. Unlike the classification task, the last layer of the regression task branch does not use a nonlinear activation function, so it directly outputs continuous values in any range. The output dimension of the regression head is equal to the number of electrical quantity measurements (e.g., A / B / C three-phase voltage, A / B / C three-phase current, total active power, total reactive power, etc.) that need to be predicted.

[0065] Finally, all classification results output by the classification task branch and all regression results output by the regression task branch are combined to form a complete collection of information.

[0066] Referring to the accompanying drawings Figure 5 Edge correction and center calibration include real-time correction performed at the edge and periodic calibration performed at the center.

[0067] Edge real-time correction, receiving complete acquisition results and performing real-time correction Utilizing a locally deployed lightweight correction model, the output is quickly corrected, and the lightweight correction model can be a Kalman filter, which takes the acquisition results output by the central model as the predicted state, and takes the latest raw measurement values directly obtained from the local sensor as the observation value, through the fusion of the two, the corrected and more optimal state estimation value is output, which filters out the possible small lag or deviation of the model prediction, and the instantaneous noise that the original sensor readings may contain, thereby generating the corrected acquisition results, and as the final output of the acquisition method at the current time step, which can be used for local real-time monitoring and display, while being packaged and uploaded to the centralized computing resource to perform central periodic calibration.

[0068] Central periodic calibration, receiving corrected acquisition results, comparing the corrected acquisition results with high-precision standard measurement models or historical reference values stored in the central database, the high-precision standard measurement model can be a simulation model based on more complex physical principles and offline precision calibration, by comparing, the calibration deviation between the corrected acquisition results and the standard value is calculated, according to the calibration deviation, the measurement coefficient stored in the center is updated, the measurement coefficient can be the parameter of the correction model used in the edge correction part, or the gain or bias parameter applied to some link, the update algorithm can adopt the idea of gradient descent, the updated measurement coefficient is stored back to the central database, and is distributed to the edge computing node at the next calibration period or when requested by the edge node, the edge computing node will update the subsequent measurement coefficient in the subsequent edge real-time correction operation to make the correction more accurate.

[0069] The present application has been verified in a certain ultra-high voltage converter station, the environment of the field area is complex, there is strong electromagnetic interference and variable climate, the experimental results show that the present application has advantages in various performance indicators compared with the traditional method.

[0070] Table 1. Performance comparison table of the present application and traditional ultra-high voltage measurement data intelligent acquisition method

[0071] Indicator category Conventional method Inventive method Mean square error (MSE) of electrical measurement value 0.018 0.009 Measurement stability (standard deviation) 0.036 0.021 Anti-interference packet loss rate (%) 3.2 0.9 Abnormality detection rate (%) 82.5 93.4 False alarm rate (%) 7.6 3.1 Accuracy of environmental auxiliary parameter (%) 86.3 94.7 Accuracy of equipment health status evaluation (%) 84.1 92.8

[0072] Performance advantage analysis summary:

[0073] More accurate and stable measurement: the mean square error of electrical measurement value is halved, and the measurement stability is improved, thanks to the hierarchical modeling and physical constraint mechanism of the improved PerceiverIO,

[0074] Stronger anti-interference capability: the data packet loss rate is reduced from 3.2% to 0.9% in an interference environment, thanks to the real-time noise reduction processing of the edge node.

[0075] The abnormality detection rate is improved to 93.4%, while the false positive rate is reduced to 3.1%. This is due to the adaptive modal gating mechanism that enhances the ability to identify multi-source abnormal signals, and the Transformer network that eliminates irrelevant noise.

[0076] More comprehensive evaluation: Higher accuracy in environmental parameter and device long-term health status evaluation, reflecting the advantages of multi-granularity output interface in multi-level feature extraction.

Claims

1. A method for intelligent acquisition of ultra-high pressure metering data based on multimodal fusion, characterized in that, Includes the following steps: S1. Acquisition and Preprocessing: Acquire multimodal raw data from the ultra-high voltage metering site, and preprocess the multimodal raw data to generate a standardized multimodal dataset. The multimodal raw data includes electrical signals, acoustic signals, and image data. S2. Feature extraction and set generation: The standardized multimodal dataset is input into a deep learning model to generate a feature set, which includes electrical signal timing feature vectors, acoustic feature vectors, and image feature vectors. S3. Multimodal fusion: The feature set is input into the improved general architecture PerceiverIO. The improved general architecture PerceiverIO performs physical constraint mechanism and adaptive modal gating mechanism and generates high-dimensional semantic feature sequence through multi-granularity output interface. S4. Temporal interactive fusion: The high-dimensional semantic feature sequence is input into the multi-head self-attention layer of the Transformer network to generate a multimodal comprehensive feature representation; S5. Result generation: Classify and regress the multimodal integrated feature representation to generate the acquisition results; S6. Edge correction and center calibration: The collected results are transmitted to the edge computing node for real-time error correction and uploaded to the central database for calibration and update of the measurement coefficients.

2. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, The S2 step is specifically as follows: The deep learning models include: the PatchTST model corresponding to electrical signals, the wav2vec model corresponding to acoustic signals, and the BEiT model corresponding to image data; The electrical signals in the standardized multimodal dataset are input into the PatchTST model to perform time-series feature modeling and generate electrical signal time-series feature vectors. The acoustic signals from the standardized multimodal dataset are input into the wav2vec model, and spectral feature transformation is performed to generate acoustic feature vectors. The image data from the standardized multimodal dataset is input into the BEiT model to perform feature extraction and generate image feature vectors. The electrical signal timing feature vector, the acoustic feature vector, and the image feature vector are combined to form the feature set.

3. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, The S3 step also includes the initial latent representation construction: The improved general-purpose architecture PerceiverIO extracts input features covering the extended time window from the feature set to generate long-term latent representations; The improved general-purpose architecture PerceiverIO extracts input features covering short-time windows from the feature set to generate high-frequency transient latent representations; The initial latent representation is obtained by combining the long-term latent representation and the high-frequency transient latent representation.

4. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, In step S3, the physical constraint mechanism specifically includes: The voltage and current representations are decoded from the features output by the improved general-purpose architecture PerceiverIO, and multiplied to obtain the theoretically calculated power representation, and a reference power representation is obtained. The difference between the theoretical power representation and the reference power representation is calculated to obtain the physical constraint loss term; The physical constraint loss term is weighted and summed with the standard loss function to form the total loss function, and the parameters of the improved general architecture PerceiverIO are optimized based on the total loss function.

5. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, In step S3, the adaptive modal gating mechanism specifically includes: The quality of the multimodal raw data constituting the feature set is evaluated, and quality indicators are calculated for the multimodal raw data. The quality index is input into the gating unit, which performs a nonlinear transformation on the quality index and applies the Softmax activation function to calculate the gate control weight. The electrical signal timing feature vector, acoustic feature vector, and image feature vector in the feature set are multiplied element-wise with the corresponding gating weights to obtain the adjusted feature set. The adjusted feature set is input into the improved general architecture PerceiverIO for feature alignment and feature fusion.

6. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, In step S3, generating the high-dimensional semantic feature sequence specifically includes: In the physical constraint mechanism, a coarse-grained output query array and a fine-grained output query array are defined. The query array is a learnable parameter array that is preset within the improved general architecture PerceiverIO and can be optimized during the training phase. The coarse-grained output query array and the fine-grained output query array are respectively used as queries for two parallel cross-attention operations, and both use the final latent representation of the improved general architecture PerceiverIO as the key and value. The two parallel cross-attention operations generate coarse-grained feature vectors and fine-grained feature vectors, respectively. The coarse-grained feature vector and the fine-grained feature vector are combined to form the high-dimensional semantic feature sequence.

7. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, The S4 step is specifically as follows: Position encoding is performed on the high-dimensional semantic feature sequence; The multi-head self-attention layer uses its internally preset query weight matrix, key weight matrix, and value weight matrix to perform a linear transformation on the sequence that has completed position encoding, generating a query matrix, a key matrix, and a value matrix, respectively. The dot product of the query matrix, key matrix, and value matrix is ​​normalized to obtain the attention weight matrix. The attention weight matrix is ​​then multiplied by the value matrix and interactively fused in the time dimension to generate the multimodal comprehensive feature representation.

8. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, The S5 step is specifically as follows: An aggregation operation is performed on the multimodal integrated feature representation to obtain an aggregated feature vector. The aggregated feature vector is then input in parallel to a preset classification head and a regression head. The classification head outputs the classification result of the state assessment, and the regression head outputs the regression result of the electrical quantity measurement. The classification result and the regression result are combined to generate the acquisition result.

9. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 1, characterized in that, In step S6, the real-time error correction performed by the edge computing node specifically involves: The edge computing node receives measurement coefficients distributed from the central database; The acquisition results are used as the predicted state of the Kalman filter, and the original measurement values ​​obtained from the ultra-high pressure metering site are used as the observation values ​​of the Kalman filter. The predicted state and observed values ​​are fused using the Kalman filter and the distributed metric coefficients to generate the corrected acquisition result as the final output of the current time step.

10. The intelligent acquisition method for ultra-high pressure metering data based on multi-modal fusion according to claim 9, characterized in that, In step S6, the calibration update performed by the central database is specifically as follows: The system receives the corrected acquisition results and compares them with a high-precision standard metrology model stored in the central database to calculate the calibration deviation. Based on the calibration deviation, the central database updates the stored measurement coefficients, and then distributes the updated measurement coefficients to the edge computing nodes for subsequent real-time error correction.