Probability distribution restoration method based on improved Wasserstein regression
By improving the Wasserstein regression method to segment and spatially transform sensor data, the problem of probability distribution repair when wireless sensor data is lost is solved, realizing high-fidelity data reconstruction and accurate characterization of structural state. It is suitable for health monitoring systems of critical infrastructure such as large bridges and high-rise buildings.
Patent Information
- Application Number
- CN202610056370.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2046-01-16
AI Technical Summary
Existing data restoration methods cannot effectively reflect the spatial coupling relationship and overall dynamic characteristics between multiple points of a structure, and cannot guarantee that the restored data has the same probability distribution as the original data. In particular, it is difficult to achieve high-fidelity reconstruction when wireless sensor data is lost.
An improved Wasserstein regression method is adopted to divide sensor data into equal segments and construct a set of probability density functions. Through logarithmic and exponential mapping, spatial transformation is performed, and the probability distribution information of multiple intact sensors is used to predict the probability density function of missing data segments. This achieves multi-source probability distribution fusion and Wasserstein spatial transformation, improving the accuracy and statistical consistency of missing data reconstruction.
When sensor data is lost, it can effectively utilize the joint distribution information of other sensors in the same span to achieve high-fidelity repair, improve the data recovery capability and the accuracy of structural state representation in complex missing scenarios, and support intelligent operation and maintenance and real-time repair of large-scale infrastructure.
Smart Images

Figure CN121542584A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a probability distribution repair method based on improved Wasserstein regression. Background Technology
[0002] In the field of structural health monitoring, wireless sensors are prone to data loss due to harsh environments, accompanied by a lack of probability distribution information, which is crucial for structural performance assessment. Existing data restoration methods rely solely on historical information from a single sensor to repair sample points, making it difficult to reflect the spatial coupling relationships and overall dynamic characteristics between multiple points in the structure, and failing to guarantee that the restored data has the same probability distribution as the original data. Summary of the Invention
[0003] This disclosure provides a probability distribution repair method based on improved Wasserstein regression.
[0004] According to one aspect of this disclosure, a probability distribution repair method is provided, comprising: acquiring structural monitoring data from different sensors for the same structure within a target time period, the structural monitoring data including first structural monitoring data with missing data and second structural monitoring data without missing data; dividing the first structural monitoring data and the second structural monitoring data into equal segments to determine a target structural monitoring data segment containing missing data and a reference structural monitoring data segment corresponding to the target structural monitoring data segment, the reference structural monitoring data segment including second structural monitoring data within the same time window as the target structural monitoring data segment; estimating the probability density function of the reference structural monitoring data segment to determine a set of probability density functions; predicting the probability density function of the target structural monitoring data segment based on a regression model using probability density functions in the set of probability density functions, the regression model being an improved Wasserstein regression model based on spatial transformation of probability density functions using logarithmic and exponential mappings; and repairing the probability distribution of the target structural monitoring data segment based on the probability density function of the target structural monitoring data segment.
[0005] One approach to probability distribution restoration based on improved Wasserstein regression involves dividing structural monitoring data from different sensors into equal segments to obtain target structure monitoring data segments containing missing data and corresponding reference structure monitoring data segments. A probability density function set is constructed using the reference structure monitoring data segments. This set is then combined with a multi-function regression model based on functional partial least squares regression (FPLS) and regression models involving logarithmic and exponential mappings to predict the probability density function of the target structure monitoring data segments. This method achieves probability distribution restoration based on multi-source probability distribution fusion and Wasserstein spatial transformation, improving the accuracy and statistical consistency of missing data reconstruction.
[0006] According to at least one embodiment of the probability distribution repair method based on improved Wasserstein regression of this disclosure, based on a regression model, the probability density function of the target structure monitoring data segment is predicted through the probability density functions in the probability density function set, including: performing a logarithmic mapping on the probability density functions in the probability density function set to determine the representation function of the probability density functions in the probability density function set; inputting the representation function of the probability density functions in the probability density function set into the regression model to determine the representation function of the target structure monitoring data segment; and performing an exponential mapping on the representation function of the target structure monitoring data segment to determine the probability density function of the target structure monitoring data segment.
[0007] According to at least one embodiment of the probability distribution repair method based on improved Wasserstein regression of this disclosure, the representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment, including: using B-spline basis functions to perform dimensionality reduction expansion on the response variable function and covariate function in the regression model, wherein the response variable function is determined based on the tangent vector of the target structure monitoring data segment, and the covariate function is determined based on the representation function of the probability density function in the probability density function set; and using functional partial least squares regression to estimate the coefficient vector of the response variable function to determine the representation function of the target structure monitoring data segment.
[0008] According to at least one embodiment of the probability distribution repair method based on improved Wasserstein regression of this disclosure, a logarithmic mapping is performed on the probability density functions in the probability density function set to determine the characterization function of the probability density functions in the probability density function set, including: determining a reference probability distribution based on the probability density functions in the probability density function set; calculating the cumulative distribution function of the probability density functions in the probability density function set and the reference probability distribution; determining a transfer mapping from the reference probability distribution to the probability density functions in the probability density function set based on the cumulative distribution function; and determining the characterization function of the probability density functions in the probability density function set based on the difference between the transfer mapping and the identity mapping.
[0009] According to at least one embodiment of the present disclosure, a probability distribution repair method based on improved Wasserstein regression is provided, wherein the reference probability distribution is the Frescher mean of the probability density function of a probability density function set in Wasserstein space.
[0010] According to at least one embodiment of the probability distribution repair method based on improved Wasserstein regression of this disclosure, the first structure monitoring data and the second structure monitoring data are divided into equal segments to determine a target structure monitoring data segment containing missing data and a reference structure monitoring data segment corresponding to the target structure monitoring data segment. This includes: aligning the first structure monitoring data and the second structure monitoring data on a time axis; dividing the aligned first structure monitoring data and second structure monitoring data into multiple equal and continuous time segments to determine corresponding first structure monitoring data segments and second structure monitoring data segments; filtering the first structure monitoring data segments to determine a target structure monitoring data segment containing missing data; and filtering the second structure monitoring data segments based on the time window in which the target structure monitoring data segment is located to determine a reference structure monitoring data segment.
[0011] The probability distribution repair method based on improved Wasserstein regression according to at least one embodiment of this disclosure further includes: establishing a joint conditional distribution function of the sensor in the target structure monitoring data segment based on the probability density function of the target structure monitoring data segment and the probability density functions of other sensors in the same time period; further generating data points, wherein the number of data points is the same as the number of sampling points in the target structure monitoring data segment, and the sampling frequency of the data points is consistent with that of the target structure monitoring data segment; and inserting the data composed of the data points into the target structure monitoring data segment to achieve probability distribution repair of the target structure monitoring data segment.
[0012] According to at least one embodiment of the probability distribution repair method based on improved Wasserstein regression of the present disclosure, based on the probability density function of the target structure monitoring data segment, the physical laws of the structure are used as sampling constraints, and data points are generated by a conditional Monte Carlo sampling method. The sampling constraints include the rate of change of adjacent data points not exceeding the maximum acceleration response threshold of the structure, local energy conservation conditions, and / or the frequency domain power spectral density within the target bandwidth.
[0013] According to another aspect of this disclosure, an electronic device is provided, comprising: a memory storing execution instructions; and a processor executing the execution instructions stored in the memory, such that the processor performs a probability distribution repair method based on improved Wasserstein regression according to any embodiment of this disclosure.
[0014] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a probability distribution repair method based on improved Wasserstein regression according to any embodiment of this disclosure. Attached Figure Description
[0015] The accompanying drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0016] Figure 1 This is a schematic diagram of the overall process of a probability distribution repair method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0017] Figure 2 This is a flowchart illustrating the process of determining a reference structure monitoring data segment in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of this disclosure.
[0018] Figure 3 This is a flowchart illustrating the process of determining the probability density function of a target structure monitoring data segment in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of this disclosure.
[0019] Figure 4 This is a flowchart illustrating the process of determining the representation function of the probability density function in a probability density function set according to an embodiment of the present disclosure, based on an improved Wasserstein regression-based probability distribution repair method.
[0020] Figure 5This is a flowchart illustrating the process of determining the characterization function of a target structure monitoring data segment in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of this disclosure.
[0021] Figure 6 This is a schematic diagram of the probability distribution repair process in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0022] Figure 7 This is a schematic diagram illustrating the probability distribution estimation of continuous missing data from multiple sensors and corresponding data segments in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of this disclosure.
[0023] Figure 8 This is a schematic diagram of an experimental data scenario in a probability distribution repair method based on improved Wasserstein regression according to one embodiment of this disclosure.
[0024] Figure 9 This is a schematic diagram of the recovery result of the missing probability density function in case 1 of a probability distribution repair method based on improved Wasserstein regression according to an embodiment of this disclosure.
[0025] Figure 10 This is a schematic diagram of the recovery result of the missing probability density function in case 2 of a probability distribution repair method based on improved Wasserstein regression according to an embodiment of this disclosure.
[0026] Figure 11 This is a schematic diagram of the recovery result of the missing probability density function in case 3 of a probability distribution repair method based on improved Wasserstein regression according to an embodiment of this disclosure.
[0027] Figure 12 This is a schematic structural block diagram of a probability distribution repair device according to one embodiment of the present disclosure.
[0028] Figure 13 This is a schematic structural block diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation
[0029] The present disclosure will now be described in further detail with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present disclosure are shown in the accompanying drawings.
[0030] It should be noted that, where there is no conflict, the embodiments and features described in this disclosure can be combined with each other. The technical solutions of this disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0031] In large bridge structural health monitoring systems, 8 to 12 wireless accelerometers are typically deployed along the length of the bridge deck to continuously collect vibration responses under environmental excitations. However, in actual operation, due to radio interference, power supply fluctuations, or hardware aging, data loss from a single sensor often occurs for several hours or even days. In extreme cases, 2 to 3 sensors within the same span may fail simultaneously. Since the structure does not deteriorate in the short term, there is a stable spatial correlation between the sensor data. However, current technology can only use the probability distribution regression of a single intact sensor to obtain the probability distribution of the missing sensor. When multiple neighboring sensors have intact data, this data information cannot be fully utilized.
[0032] To address this, this disclosure proposes a probability distribution repair method based on improved Wasserstein regression. Under the premise of stable short-term structural performance, the method segments the monitoring data of each sensor and estimates their probability density function. Then, it uses logarithmic mapping to transform the distribution characteristics to the tangent space for functional regression modeling. This fully integrates the probability distribution information of multiple intact sensors to predict the distribution pattern of the monitoring data segment of the target structure. Finally, it restores the data through exponential mapping and generates repair data that conforms to statistical characteristics. Thus, even when one or more neighboring sensors fail, it can still effectively utilize the joint distribution information of other sensors in the same span to achieve high-fidelity repair, improving the data recovery capability and the accuracy of structural state representation in complex missing scenarios.
[0033] The probabilistic distribution repair method disclosed herein can be deployed on cloud servers for centralized intelligent operation and maintenance of large-scale infrastructure, and can also be integrated into edge computing gateways or field monitoring terminals for low-latency real-time repair. In health monitoring systems for critical infrastructure such as large bridges, high-rise buildings, dams, or wind turbine towers, this probabilistic distribution repair method can run on embedded edge devices deployed on-site. Utilizing locally acquired multi-sensor data streams, it can autonomously complete probabilistic distribution-level repair of missing data at key measuring points even in the event of communication interruptions or bandwidth limitations, ensuring the continuity and reliability of the status assessment module. Simultaneously, this probabilistic distribution repair method can also be integrated as a core algorithm into mobile inspection terminals (such as industrial tablets or smartphones), allowing engineering technicians to wirelessly access sensor fragment data on-site, perform probabilistic distribution repair and visualization analysis in real time, and improve operation and maintenance efficiency. Furthermore, in city-level structural cluster monitoring platforms for smart city construction, this probabilistic distribution repair method can be deployed on a central server to batch process historical and real-time data from massive distributed sensor networks, construct cross-structure and cross-regional collaborative repair models, and serve data consistency maintenance for urban safety early warning and digital twin systems, realizing extended applications from individual repair to group intelligent diagnosis.
[0034] Figure 1 This is a schematic diagram of the overall process of a probability distribution repair method according to one embodiment of this disclosure. Figure 1 The method M100 shown includes steps S110 to S150. This method can be executed by an electronic device such as a mobile phone or tablet.
[0035] In step S110, for the same structure within the target time period, structural monitoring data from different sensors are acquired. The structural monitoring data includes first structural monitoring data with missing data and second structural monitoring data without missing data.
[0036] For the same engineering structure within a target time period, structural monitoring data from multiple sensors are collected. The structural monitoring data includes first structural monitoring data where sampling is missing or communication is interrupted in some time periods, and second structural monitoring data with no data loss and good quality within the same time period.
[0037] For example, in the actual operation of health monitoring of large bridges, it is common for some sensors to experience intermittent data loss due to complex field environments, unstable power supply, or aging equipment. For instance, during a typhoon, multiple vibration sensors in a span may experience temporary wireless module disconnection due to strong winds, but continuous data from sensors in other unaffected areas can still be acquired. Furthermore, in the dense monitoring network of elevated sections of urban rail transit, hundreds or thousands of sensors operate simultaneously. Manually checking data integrity is inefficient. Automated scripts can be used to batch scan the data streams of each channel, automatically marking the monitoring data of the first structure and the monitoring data of the second structure based on set thresholds, thus achieving an efficient and standardized data preprocessing workflow.
[0038] Optionally, structural monitoring data consists of synchronous time-series data acquired by sensors deployed on engineering structures such as bridges, buildings, or dams. This data is used to analyze the correlation between different measuring points and serves as the fundamental input for probability density function estimation and probability distribution repair. It reflects the dynamic response of the sensors under environmental excitation and external loads, including measurements of physical quantities such as acceleration, strain, displacement, tilt, and / or temperature. These data record the vibration modes, stress distribution, and deformation characteristics of the structure during service, serving as a core basis for assessing structural health, identifying damage, and providing early warnings of potential safety hazards.
[0039] In step S120, the first structure monitoring data and the second structure monitoring data are divided into equal segments to determine the target structure monitoring data segment containing missing data and the reference structure monitoring data segment corresponding to the target structure monitoring data segment. The reference structure monitoring data segment contains the second structure monitoring data that is in the same time window as the target structure monitoring data segment.
[0040] The first and second structure monitoring data are synchronized and aligned on the time axis, and then divided into multiple equal and continuous time segments, forming a one-to-one correspondence between the first and second structure monitoring data segments. Each structure monitoring data segment corresponds to a fixed-duration observation window (e.g., 1 hour, 6 hours, or 1 day), ensuring that all sensors have the same time range within the same sequence number of the structure monitoring data segment. The first structure monitoring data segment contains at least one incomplete target structure monitoring data segment due to missing original sampling; this target structure monitoring data segment serves as the object for subsequent probability density function prediction and probability distribution repair. The second structure monitoring data segment contains a reference structure monitoring data segment within the same time window as the target structure monitoring data segment.
[0041] In step S130, the probability density function is estimated for the reference structure monitoring data segment to determine the probability density function set.
[0042] The reference structure monitoring data segments are transformed into geometrically meaningful probability distribution representations, constructing a probability density function set reflecting the cooperative response characteristics. Probability density function (PDF) estimation is performed on the reference structure monitoring data segments, and the estimation results are organized into a time series set, forming a probability density function set used to predict the distribution of the target structure monitoring data segments. Each probability density function in the set corresponds to the probability distribution of a reference sensor within a time window, constituting the input variable of the regression model. The reference sensor is the sensor that collects complete structural monitoring data, i.e., the second structural monitoring data.
[0043] Optionally, based on the time segment index, each probability density function in the probability density function set is structured and stored, wherein each probability density function in the probability density function set contains the identifier of the corresponding reference sensor.
[0044] In step S140, based on the regression model, the probability density function of the target structure monitoring data segment is predicted by the probability density function in the probability density function set. The regression model is an improved Wasserstein regression model based on the spatial transformation of probability density functions using logarithmic and exponential mappings.
[0045] Logarithmic mapping is used to spatially transform the probability density functions corresponding to each reference sensor in the probability density function set. The transformed probability density function set is then input into a regression model to learn the mapping relationship between the representation functions of multiple reference sensors and the corresponding representation function of the target sensor within the same time window, thus obtaining the representation function of the target structure monitoring data segment. An exponential mapping operation is then performed on the predicted representation function of the target structure monitoring data segment, remapping it from the tangent space back to the Wasserstein space to recover the probability density function of the target structure monitoring data segment. Here, the target sensor is the sensor whose collected structural monitoring data contains missing data.
[0046] Preferably, the regression model is an improved Wasserstein regression model based on multi-function regression with functional partial least squares regression (FPLS) and spatial transformation of probability density functions using logarithmic and exponential mappings.
[0047] Preferably, given that structural degradation or changes in the boundary conditions of the structural system typically do not occur in the short term, and the spatial correlation between the structural monitoring data recorded by sensors does not change significantly in a short period of time, it is assumed that the spatial correlation between the probability density functions of different sensors in the probability density function set at different time periods is consistent, and the probability density function set is used as a functional sample to train the regression model from multiple probability distributions to a distribution.
[0048] Preferably, the regression model is constructed based on the geometric framework of Wasserstein regression. By introducing logarithmic and exponential mapping operations of the reference distribution in the Wasserstein space, the complex distribution-to-distribution prediction problem is transformed into a functional regression problem in the tangent space, thereby achieving efficient and geometrically consistent prediction of the structural response distribution.
[0049] Preferably, a logarithmic mapping is applied to the probability density function set to map the probability distribution to the tangent space, transforming it into a common function, called the representation function, thus forming a representation function set. Within the tangent space, multi-function regression based on functional partial least squares regression (FPLS) is used to predict the representation function of the target structure monitoring data segment during the same time period using the representation functions of multiple sensors in the representation function set. An exponential mapping is then applied to the predicted representation functions of the target structure monitoring data segment during the same time period, mapping them back to the probability space, thus completing the repair of the missing probability distribution function of the target structure data segment.
[0050] In step S150, the probability distribution of the target structure monitoring data segment is repaired based on the probability density function of the target structure monitoring data segment.
[0051] The predicted probability density function is used as a mathematical representation of the statistical regularity of the target structure monitoring data segment. It is directly output and stored as the repaired probability distribution, thereby achieving high-fidelity reconstruction of the structural response amplitude distribution characteristics (such as kurtosis, skewness, and energy concentration intervals), restoring the complete distribution shape of the target structure monitoring data segment in the probability space. It can be used for subsequent health status assessment, anomaly detection, multi-source data consistency, or data repair analysis tasks based on distribution characteristics.
[0052] For multi-sensor monitoring data of the same structure within a target time period, the system distinguishes between missing first-structure monitoring data and complete second-structure monitoring data. Both first-structure and second-structure monitoring data are then equally segmented to form a first-structure monitoring dataset and a second-structure monitoring dataset containing the target structure monitoring data segment. Based on the reference structure monitoring data segment, a probability density function set is constructed. Using this set, a multi-function-function regression based on functional partial least squares regression within the Wasserstein regression framework, employing a spatial transformation mechanism based on logarithmic and exponential mapping, predicts the probability density function of the target structure monitoring data segment. Then, time-series data is generated based on the probability density function of the target structure monitoring data segment to complete the probability distribution restoration of the target structure monitoring data segment. This achieves high-fidelity restoration of structural monitoring data, with multi-source distribution fusion, geometrically consistent mapping, and distribution-level reconstruction as its core principles.
[0053] Regarding step S120, in some embodiments of this disclosure, it may include, for example... Figure 2Steps S1201 to S1204 are shown.
[0054] In step S1201, the first structural monitoring data and the second structural monitoring data are aligned on the time axis.
[0055] The system acquires timestamp information from each sensor, identifies its time reference (such as UTC, local system time, or GPS time synchronization), and unifies the structural monitoring data from all sensors onto a common time axis through interpolation, resampling, or offset correction. The first and second structural monitoring data are aligned with the same time resolution and start / end times to ensure that the structural monitoring data segments from each sensor have a strictly consistent time range in subsequent segmentation operations. This supports probability density function estimation and regression prediction based on the response of multiple measurement points under the same operating condition.
[0056] In step S1202, the aligned first and second structure monitoring data are divided into multiple time segments of equal and continuous length, and the corresponding first and second structure monitoring data segments are determined.
[0057] According to a preset time length, the first and second structure monitoring data are divided into a series of equal-length, continuous time segments with no or partial overlap. Each time segment corresponds to a fixed time interval, ensuring that all sensors cover the exact same time range within the same numbered segment. This forms a one-to-one correspondence between the first and second structure monitoring data segments, where the former contains at least one incomplete target structure monitoring data segment due to sampling interruptions. This division method guarantees the comparability of the responses of each sensor under the same operating conditions, providing a synchronous input basis for probability density function estimation and regression modeling.
[0058] In step S1203, the first structure monitoring data segment is filtered to determine the target structure monitoring data segment containing missing data.
[0059] After dividing the first structure monitoring data into multiple continuous and equal-length time segments, a data quality assessment is performed on each segment. Based on preset missing data criteria (such as a sampling point loss rate exceeding a threshold, consecutive NaN value length reaching a set proportion, or signal amplitude remaining constant), segments containing missing data are identified. Data segments that meet these criteria are marked as target structure monitoring data segments.
[0060] In step S1204, the second structure monitoring data segment is filtered based on the time window in which the target structure monitoring data segment is located, and the reference structure monitoring data segment is determined.
[0061] For the identified target structure monitoring data segment, its corresponding time window information is obtained. All second structure monitoring data segments are traversed, and segments whose time ranges completely overlap or highly overlap with theirs are selected. These second structure monitoring data segments within the same time interval are used as reference structure monitoring data segments to characterize the overall dynamic state of the structure during that period. The reference structure monitoring data segments are derived from other sensors that are spatially adjacent to or functionally related to the target structure monitoring data segment, ensuring that their response characteristics are comparable and correlated, thereby supporting multi-source fusion modeling in subsequent probability density function estimation and regression prediction tasks.
[0062] The first and second structural monitoring data are precisely aligned on the timeline based on timestamp information and unified to the same sampling resolution and time reference. The aligned first and second structural monitoring data are then divided into multiple equal-length, continuous time segments with strictly consistent time ranges, forming one-to-one corresponding structural monitoring data segments. Based on this, the integrity of the first structural monitoring data segments is assessed to identify target structural monitoring data segments containing sampling interruptions or data loss. According to the time window of this target segment, second structural monitoring data segments with complete data from other sensors are located and selected as reference structural monitoring data segments, ensuring that their response characteristics reflect the dynamic behavior under the same external operating conditions. This guarantees the comparability of each measuring point within the same time interval.
[0063] Regarding step S140, in some embodiments of this disclosure, it may include, for example... Figure 3 Steps S1401 to S1403 are shown.
[0064] In step S1401, a logarithmic mapping is performed on the probability density functions in the probability density function set to determine the characterization function of the probability density function in the probability density function set.
[0065] Logarithmic mapping transforms probability density functions, which are inherently impossible to directly add or regress, into function vectors in the tangent space, making classical statistical methods applicable. Based on optimal transport theory, logarithmic mapping ensures the transformation follows the shortest path principle, avoiding distribution ambiguity or quality leakage caused by traditional L² space averaging. Furthermore, all representation functions share the same reference coordinate system, facilitating the construction of unified multivariate regression relationships and improving the interpretability of the regression model. The representation functions typically have small amplitudes and smooth variations, making them suitable for basis function expansion and dimensionality reduction, which improves the training speed and convergence of the regression model.
[0066] In step S1402, the representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment.
[0067] The representation functions of the probability density functions corresponding to each sensor in the probability density function set are used as input variables and fed into a pre-trained or real-time constructed regression model. The regression model learns the functional mapping relationship between the representation functions of the probability density functions of each reference sensor in the probability density function set and the representation function of the target structure monitoring data segment of the target sensor based on historical complete time-period structural monitoring data. Within the current target time period, the regression model outputs the representation function of the target structure monitoring data segment based on the changing trends of the probability density functions of each reference sensor in the probability density function set. The representation function of the target structure monitoring data segment represents the optimal transmission offset direction and amplitude that the target sensor should have relative to the reference distribution under the same operating conditions, serving as the basis for subsequently recovering its complete probability density function through exponential mapping.
[0068] In step S1403, the representation function of the target structure monitoring data segment is subjected to exponential mapping to determine the probability density function of the target structure monitoring data segment.
[0069] The exponential mapping follows geodesic paths in Wasserstein space, ensuring optimal transmission between the reconstructed and reference distributions and avoiding non-physical oscillations or distortions. When the characterization function contains multiple local fluctuations, the exponential mapping can naturally evolve into complex distribution patterns such as bimodal, wide-tailed, or multimodal distributions, adapting to non-stationary responses in real-world engineering. The exponential and logarithmic mappings give the entire Wasserstein regression framework both mathematical closure and engineering practicality.
[0070] Preferably, a logarithmic mapping is performed on the probability density functions in the probability density function set to obtain unconstrained representation functions with linear spatial structures corresponding to the probability density functions in the probability density function set. Within the tangent space, the representation functions corresponding to the probability density functions in the probability density function set are input into a multi-function regression based on functional partial least squares regression established using historical data to determine the representation function of the target structure monitoring data segment; an exponential mapping is then performed on the representation function of the target structure monitoring data segment to determine the probability density function of the target structure monitoring data segment.
[0071] By performing a logarithmic mapping on the probability density functions in the probability density function set, the nonlinear distribution space is projected to a flat tangent space, transforming it into a computable representation function. This process is based on optimal transport theory and unifies all representation functions under the same geometric coordinate system, supporting subsequent multivariate functional regression modeling. These representation functions are fed as input variables into a multi-function regression model constructed using partial least squares regression, learning the distribution evolution coupling relationship of the probability density functions of multiple reference sensors and the target sensor under the same operating conditions. This allows for the prediction of the representation function of the target structure monitoring data segment, accurately representing its optimal transport offset direction and magnitude relative to the reference state. An exponential mapping operation is then performed on the representation function of the target structure monitoring data segment, reconstructing it into the probability density function of the target structure monitoring data segment. This not only ensures the mathematical validity of the output distribution but also naturally generates complex shapes such as bimodal and / or wide-tailed patterns, truly reflecting the dynamic response characteristics under non-stationary excitation. This achieves the goal of intelligently inferring the statistical behavior of missing locations from the distribution evolution of multi-source sensors, improving the robustness, physical consistency, and engineering usability of the structural health monitoring system in scenarios with incomplete data.
[0072] Regarding step S1401, in some embodiments of this disclosure, it may include, for example... Figure 4 Steps S410 to S440 are shown.
[0073] In step S410, a reference probability distribution is determined based on the probability density function in the probability density function set.
[0074] Preferably, the reference probability distribution is the Friesian mean of the probability density functions of the probability density function set in the Wasserstein space.
[0075] Specifically, based on the probability density functions corresponding to each reference sensor in the probability density function set, their Fréchet mean under the Wasserstein distance metric is calculated, and the probability distribution corresponding to the Fréchet mean is determined as the reference probability distribution. This reference probability distribution serves as the common base point for subsequent logarithmic and exponential mappings, used to project each probability density function from the Wasserstein manifold to the tangent space, or conversely, to reconstruct the prediction results. This reference distribution comprehensively reflects the common statistical characteristics of multiple sensors within the same time window, exhibiting good representativeness and stability, and supporting probability distribution comparison and fusion modeling across sensors and time periods.
[0076] In step S420, the probability density function of the probability density function set and the cumulative distribution function of the reference probability distribution are calculated.
[0077] For each probability density function in the probability density function set and the reference probability distribution, their corresponding cumulative distribution functions are calculated. The cumulative distribution function is continuous and monotonically increasing within its domain. This ensures that all distributions are expressed in a unified functional form, providing the necessary input foundation for subsequent logarithmic mapping based on optimal transport theory, and supporting cross-sensor mass element matching and deformation field calculation.
[0078] In step S430, based on the cumulative distribution function, the transfer mapping from the reference probability distribution to the probability density function in the probability density function set is determined.
[0079] Based on the cumulative distribution function of the reference probability distribution and the cumulative distribution function of the probability density functions in the probability density function set, an optimal transfer mapping is constructed from the reference distribution to the cumulative distribution function of each probability density function. The transfer mapping represents the transformation function that maps each mass point in the reference distribution to the same quantile position in the target distribution, reflecting the local shift behavior of the probability density function relative to the common benchmark. This optimal transfer mapping is strictly monotonically increasing, ensuring mass conservation and no cross-transfer, and supports subsequent extraction of its corresponding representation function through logarithmic mapping for use in multivariate regression modeling.
[0080] In step S440, the representation function of the probability density function in the probability density function set is determined based on the difference between the transport mapping and the identity mapping.
[0081] The pointwise difference between each transport map and the identity map is calculated to obtain the representation function of the corresponding probability density function. Here, the nonlinear Wasserstein manifold is locally approximated as a flat tangent space, allowing traditional statistical methods to be directly applied to probability distribution data. The original probability density function requires thousands of points to describe, while the representation function, after dimensionality reduction, can be represented by dozens of parameters, facilitating real-time processing and transmission.
[0082] Based on the probability density functions in the probability density function set, the Friesian mean of each function in the Wasserstein space is calculated and used as the reference probability distribution. This reference distribution serves as the geometric center of the entire distribution set and the common base point for subsequent mapping operations, ensuring good representativeness and stability in the modeling process. The cumulative distribution function (CDF) of each probability density function and the reference probability distribution is calculated separately, transforming the probability distribution into a continuously monotonically increasing functional form, providing a unified and differentiable mathematical foundation for optimal transport analysis. Based on the reference distribution and the CDF of each probability density function, an optimal transport mapping from the reference distribution to the target distribution is constructed. This mapping strictly maintains the quality order and accurately characterizes the nonlinear deformation path of each sensor response relative to the common benchmark. By calculating the difference between the transport mapping and the identity mapping, the representation function of each probability density function in the tangent space is obtained, transforming the distribution offset on the nonlinear manifold into an additive and regressible functional vector in a flat space. A complete Wasserstein geometric processing chain was constructed, which not only avoids the distribution ambiguity problem caused by traditional Euclidean averaging, but also enables the probability density function, which was originally not directly computable, to be multivariate statistically modeled under a unified coordinate system, thereby improving the accuracy, interpretability and computational efficiency of multi-source distributed data fusion.
[0083] Regarding step S1402, in some embodiments of this disclosure, it may include, for example... Figure 5 Steps S510 to S520 are shown.
[0084] In step S510, B-spline basis functions are used to perform dimensionality reduction expansion on the response variable function and covariate function in the regression model. The response variable function is determined based on the representation function of the target structure monitoring data segment, and the covariate function is determined based on the representation function of the probability density function in the probability density function set.
[0085] By compressing thousands of dimensions of functional data into tens of dimensions of coefficient vectors, high-dimensional regression problems are transformed into conventional multivariate regression, significantly improving training speed and convergence. B-spline basis functions possess strong smoothness and local support properties, avoiding high-frequency noise interference and ensuring a natural and oscillating reconstructed function shape. The small coefficient vectors of the response and covariate functions are suitable for edge-cloud collaborative architectures, reducing bandwidth pressure and supporting remote intelligent operation and maintenance of wide-area infrastructure clusters. The output format can be directly integrated with mature statistical models such as FPLS, FPCR, and GAM, enhancing the method's versatility and maintainability. Furthermore, the basis function expansion itself has a regularization effect, suppressing overfitting.
[0086] Preferably, the number of basis functions in the regression model is determined through cross-validation, and the number of basis functions is used to describe the complexity of the regression model. The number of basis functions is used as a core hyperparameter controlling the complexity of the regression model. During the training phase of the regression model, cross-validation is used to iterate through a set of candidate basis functions, constructing a corresponding regression model for each basis function, and calculating its prediction error on the validation set. The number of basis functions that minimizes the validation error is selected as the optimal number of basis functions for constructing the final model. This ensures that the determined regression model has sufficient flexibility to capture the spatial evolution characteristics of the representation function while effectively suppressing overfitting and improving prediction stability under unknown conditions.
[0087] In step S520, the coefficient vector of the response variable function is estimated using the functional partial least squares regression method to determine the characterization function of the target structure monitoring data segment.
[0088] Using the representation function expanded from B-spline basis functions as the input variable and the representation function of the target structure monitoring data segment as the response variable, a functional partial least squares regression method is employed to estimate the coefficient parameters of the response variable function. By iteratively extracting latent variables and maximizing the covariance between the input and output blocks, a compact multiple linear regression relationship is established. After the regression model is trained, for a new target structure monitoring data segment, the regression model can automatically predict the B-spline expansion coefficients of the representation function of the new target structure monitoring data segment based on the B-spline expansion coefficients of the representation function of the reference structure monitoring data segment acquired by a concurrent reference sensor. This allows for the reconstruction of a complete representation function, which is then used for subsequent exponential mapping to recover its probability density function.
[0089] B-spline basis functions are used to reduce the dimensionality of the response variable function and cofunction, compressing the functional variables that originally required thousands of sampling points into low-dimensional vectors with only dozens of coefficients. This significantly reduces computational load and communication overhead while preserving key morphological features of the distribution offset. Furthermore, its smoothness and local support characteristics effectively suppress noise interference, ensuring the reconstructed function is natural and oscillatory. Cross-validation is used to systematically evaluate the predictive performance of regression models with different numbers of basis functions. The optimal number of basis functions is automatically determined based on minimizing the validation set error, scientifically balancing the model's expressive power and generalization ability, avoiding underfitting or overfitting, and improving the model's robustness under diverse operating conditions. Functional partial least squares regression (FPLS) is used to establish a multiple regression relationship between the input and output representation function coefficients. By extracting the latent variable with the largest covariance, the problem of multicollinearity among multiple input sources is effectively addressed, achieving accurate inference from the distribution evolution patterns of representation functions from multiple reference sensors to the target sensor's representation function. This transforms complex functional regression problems into efficient and solvable numerical tasks, providing an accurate, stable, and physically consistent predictive foundation for reconstructing the probability density function of the target structure monitoring data segment through exponential mapping.
[0090] In some embodiments of this disclosure, it may include, for example Figure 6 Steps S610 to S620 are shown.
[0091] In step S610, based on the probability density function of the target structure monitoring data segment, and in conjunction with the probability density functions of other sensors in the same time period, a joint conditional distribution function of the sensor in the target structure monitoring data segment is established, and data points are further generated. The number of data points is the same as the number of sampling points in the target structure monitoring data segment, and the sampling frequency of the data points is consistent with that of the target structure monitoring data segment.
[0092] Preferably, based on the probability density function of the target structure monitoring data segment, and in conjunction with the probability density functions of other sensors in the same time period, a joint conditional distribution function of the sensor in the target structure monitoring data segment is established. The physical laws of the structure are used as sampling constraints, and data points are generated through the conditional Monte Carlo sampling method. The sampling constraints include the rate of change of adjacent data points not exceeding the maximum acceleration response threshold of the structure, local energy conservation conditions, and / or the frequency domain power spectral density being within the target bandwidth.
[0093] Specifically, based on the probability density function of the target structure monitoring data segment reconstructed through exponential mapping, combined with the probability density functions of other sensors in the same time period, a joint conditional distribution function for the sensor in the target structure monitoring data segment is established. Monte Carlo sampling or inverse transform sampling methods are then used to randomly select a set of numerical samples, i.e., data points, that conform to this distribution. The number of generated data points is strictly equal to the total number of sampling points that should be collected under normal operating conditions for that time segment. Simultaneously, the time interval between data points is ensured to be consistent with the sampling frequency of the target structure monitoring data segment, forming a complete and equidistant data point sequence. This data point sequence not only restores the data length and temporal structure but also accurately reflects the statistical characteristics of the structural response within that time period, and can be used to replace missing data in subsequent health assessments and data analysis.
[0094] In step S620, the data composed of data points is inserted into the target structure monitoring data segment to achieve data repair of the target structure monitoring data segment.
[0095] The data point sequence replaces or fills the target structure monitoring data segment according to its corresponding time window and time sequence. The insertion operation maintains the original database structure, timestamp format, and unit system unchanged, forming a logically continuous and uninterrupted complete structure monitoring data segment. Metadata tags can also be attached to identify the source of the repair, supporting audit traceability and credibility assessment, ensuring the availability and legitimacy of the repaired data in various downstream tasks.
[0096] Based on the probability density function of the target structure monitoring data segment reconstructed by exponential mapping, and combined with the probability density functions of other sensors in the same time period, a joint conditional distribution function for the sensor in the target structure monitoring data segment is established. Monte Carlo or inverse transform sampling methods are then used to generate a data point sequence with the same number of sampling points and consistent sampling frequency as the original sampling points, ensuring that the filled data statistically reflects the dynamic laws of the structural response in that time period. Structural physical laws are introduced as conditional sampling constraints. Conditional Monte Carlo sampling controls the rate of change of adjacent data points to not exceed the maximum acceleration response threshold of the structure, satisfies local energy conservation conditions and / or frequency domain power spectral density distribution characteristics, effectively avoiding unreasonable abrupt changes or non-physical interpretations, and improving the temporal smoothness and dynamic realism of the synthesized data. The generated data point sequence is seamlessly inserted into the target structure monitoring data segment according to the time window and temporal sequence, replacing or filling the original missing intervals, maintaining complete consistency with the database structure, timestamp format, and unit system. The source and method of the repair can be identified through additional metadata tags, supporting data traceability and credibility assessment. Not only did it restore the length and temporal continuity of the data, but it also ensured a high degree of consistency in statistical distribution, physical rationality, and engineering usability. This allows the repaired data to be directly used for routine tasks such as subsequent spectrum analysis, damage identification, or over-limit alarms, greatly improving the robustness, intelligence level, and practical operation and maintenance value of the structural health monitoring system in scenarios with missing data.
[0097] In one specific embodiment, such as Figure 7 As shown in sub-diagram (a), it is assumed that the components are installed at different locations on the structure. m There are 10 sensors, of which some sensors intermittently fail (let's say the 10th sensor). h 1 and the h Two sensors were used, and consecutive gaps occurred at different time intervals, resulting in the loss of probability distribution information for the corresponding time periods. (The sentence about the first sensor appears unrelated and likely refers to a separate topic.) h Taking a multi-probability distribution to probability distribution regression task from a single sensor as an example, segments with continuously missing data on this sensor are denoted as... (i.e., the target structure monitoring data segment), its missing data is as follows Figure 7 The sub-diagram shown in (b) is denoted as Except for the first h Other sensors besides the two main sensors serve as collaborating sensors, or reference sensors, used to regress missing data. The probability distribution. The set of indices for collaborative sensors is denoted as... .like Figure 7 As shown in subplot (b), to establish a multi-probability distribution to distribution regression model, each complete data sequence is equally divided into several data segments to construct probability distribution function samples. Figure 7As shown in subplot (c), it is assumed that data within the same data segment follows the same probability distribution, but the probability distributions differ between different data segments. The probability density function is estimated for each data segment without missing data to obtain the probability density function set of the reference structure data segment.
[0098] Given that structural degradation or changes in the boundary conditions of the structural system typically do not occur in the short term, and the spatial correlation between sensor-recorded data does not change significantly in a short period, it is assumed that the probability density functions of each data segment have consistent correlation. The probability density function of each intact data segment (i.e., the reference structure monitoring data segment) is used as a functional sample to train a multi-probability distribution-to-distribution regression model. Therefore, the task of multi-probability distribution-to-distribution regression of multi-sensor data is to identify the missing data segments. Above, the probability distribution of multiple collaborative sensors in the same data segment is utilized. , restore the h Missing data from one sensor Missing probability distribution .like Figure 7 As shown in subgraph (a), for the th h For two sensors, due to the long continuous data loss period, the missing data should first be segmented into shorter fragments, and then processed. h For each missing segment of the two sensors, perform the same multi-distribution to distribution regression task as above.
[0099] First, apply the logarithmic mapping in equation (1) to all probability density functions in the probability density function set to map them from the Wasserstein space to the tangent space, thereby obtaining the unconstrained representation function, i.e., the covariate function. and response variable function Probability density function in a set of probability density functions The logarithmic mapping can be defined as: (1) In the formula, Describing the probability density function in a set. The representation function in the tangent space, This represents an identity mapping. Is with the first i The cumulative distribution function (CDF) of the Fréchet mean corresponds to a probability distribution sequence (i.e., probability density function). Since we only consider univariate probability distributions here, we can first calculate the quantile function of the Fréchet mean. Then, by inverting the function, the corresponding cumulative distribution function can be obtained. Quantile function It can be calculated according to formula (2).
[0100] (2) Since the tangent space is a subspace of the Hilbert space, regression methods for ordinary functional data can be used in this space. To extend Wasserstein regression to regression from multiple probability distributions to probability distributions, this disclosure embeds a multi-function-to-function regression method based on functional partial least squares regression (FPLS) into the tangent space to characterize the mapped covariate function. With response variable function The relationship between them can be expressed in the following form: (3) in, Indicates the subject With response variable function The regression coefficient function between them; Represents a quadratic term With response variable function The regression coefficient function between them This represents the error term function.
[0101] Since functional data is inherently infinite-dimensional, directly estimating these coefficient functions is computationally expensive and statistically unstable, especially with limited sample sizes. To alleviate this problem, basis function expansion is typically used to reduce dimensionality and improve computational efficiency. B- Spline basis functions are widely used because they combine flexibility, computational efficiency, and local support. B - Construction of spline bases from the objective function Let a sequence of nodes be set as the starting point on the domain of the variable. Let the variable be defined in the interval... Above, its node vector is denoted as ,in d for B - The degree of the spline (i.e., the order of the polynomial). M The number of basis functions, i.e. B - Number of terms in the spline expansion. Based on this sequence of nodes, piecewise polynomial basis functions can be constructed using the Cox-de Boor recurrence relation. Thus, the objective function... It can be approximated as a linear combination of these basis functions: (4) in, express B -spline basis functions, Let T be the corresponding coefficient vector, and T denote the matrix transpose. This disclosure can be determined through cross-validation. B - The number of terms in the spline expansion is adjusted to achieve an optimal balance between flexibility and simplicity. Therefore, the response variable function, independent variable function, and error function in equation (5) can all be used.B The spline basis function expansion is as follows: (5) (6) (7) in, and They are respectively represented as and Selected B -Spline basis functions; and For the corresponding B - Number of terms in the spline expansion; and It is the corresponding expansion coefficient vector; for The expansion coefficients.
[0102] For quadratic terms Its expanded form can be expressed as: (8) in, , .
[0103] Similarly, the main effect coefficient function and quadratic coefficient function You can also do it through B- Spline expansion takes the following form: (9) (10) In the formula, and They represent the first i The function of the main effect coefficient and the ( i , j ) quadratic coefficient function of B- Spline expansion coefficient matrix. and The specific form is as follows: (11) According to the above B- During the spline expansion process, equation (3) can be rewritten in the following form: (12) Simplifying the above equation, we get: (13) In the formula , Therefore, the regression problem from multiple probability distributions to a probability distribution is transformed into solving the coefficient matrix of the multiple regression equation (13). , and The problem lies in the use of more predictors or more basis functions. and The dimensionality increases rapidly, potentially leading to ill-conditioned problems (multicollinearity). To address this, dimensionality reduction techniques such as Principal Component Regression (PCR) or Partial Least Squares Regression (PLS) can be employed. PCR reduces dimensionality by extracting principal components from the prediction matrix. However, the extracted principal components do not consider the correlation between the predictor and response variables, making it impossible to determine the impact of each regression parameter on the response variable. PLS, on the other hand, extracts principal components from both the prediction and response matrices simultaneously, maximizing the correlation between these principal components. The prediction and response matrices are then regressed on the extracted principal components, respectively. This process is iterated until satisfactory results are obtained through cross-validation. The coefficients are then obtained using PLS. , and After that, the main effect coefficient function Quadratic coefficient function and It can be obtained through equations (9), (10) and (7) respectively.
[0104] Once the regression model from multiple probability distributions to probability distributions is established as shown in formula (3), the segment containing the missing data is determined. The representation function, i.e., the response variable function, can be calculated as follows: (14) Finally, an exponential mapping is used to represent the function, i.e., the response variable function. The mapping from the tangent space back to the Wasserstein space is defined as follows: (15) In the formula express The inverse function of .
[0105] Therefore, the embodiments of this disclosure are verified through the following experimental data. The example data comes from a wireless monitoring system installed on a pedestrian overpass, which is a two-span continuous rigid frame bridge located on a university campus. Figure 8(a) subgraph and Figure 8 As shown in sub-figure (b), the length is 44 meters and the width is 3.7 meters. The signal-to-noise ratio of the observation data from sensors 4 and 8 is low; therefore, this disclosure uses only the monitoring data from accelerometers 1, 2, 3, 5, 6, and 7. A total of 17 weeks of acceleration data were recorded in the provided data; this disclosure uses only the acceleration data from the third week under full excitation.
[0106] To fully verify the effectiveness of the method proposed in this disclosure, the following missing operating conditions were considered: Condition 1: Only one sensor has data loss (assuming sensor 2 has data loss), and the data from the other 5 sensors is used to repair the data. Condition 2: Two sensors have data loss at the same time (assuming that sensors 2 and 6 have data loss), and the data from the other four sensors is used to repair the data. Condition 3: Three sensors have data loss at the same time (assuming that sensors 1, 3 and 6 have data loss), and the data from the remaining three sensors is used to repair the data.
[0107] like Figures 9 to 11 As shown, the improved Wasserstein regression (IWR) and the original Wasserstein regression (WR) are compared under three missing data scenarios to restore the probability density functions. In each figure, the solid line represents the result of the method of this disclosure, and the dashed line represents the result of the original Wasserstein regression. Furthermore, the legend a to b indicates that the original Wasserstein regression recovers the probability distribution of sensor b based on the data of sensor a.
[0108] The table below compares the average IAE values of the three-segment recovery probability density functions under different scenarios. Under three different operating conditions, when a specific faulty sensor fails, the probability density function of the faulty sensor is recovered using other different sensors. In operating condition 1, the average integral absolute error (IAE) of the improved Wasserstein regression is only 0.0260, far lower than the minimum error of any single-sensor recovery, i.e., 0.0700. In more complex scenarios with simultaneous failure of multiple sensors (such as operating condition 3), the improved Wasserstein regression still maintains stable low errors (i.e., 0.0553, 0.0430, 0.0353), while the original Wasserstein regression method suffers from large error fluctuations and poor accuracy due to its reliance on a single sensor. This fully demonstrates that the improved Wasserstein regression, by fusing multi-sensor information and preserving the distribution geometry, can more accurately repair missing probability distributions and effectively improve the integrity of structural health monitoring data.
[0109] Based on any of the above embodiments, this disclosure also provides a probability distribution repair device.
[0110] Figure 12 This is a schematic block diagram of the probability distribution repair device according to one embodiment of the present disclosure.
[0111] like Figure 12 As shown, the probability distribution repair device includes: The data acquisition module 1202 acquires structural monitoring data from different sensors for the same structure within a target time period. The structural monitoring data includes first structural monitoring data with missing data and second structural monitoring data without missing data. The data segmentation module 1204 divides the first structure monitoring data and the second structure monitoring data into equal segments to determine the target structure monitoring data segment containing missing data and the reference structure monitoring data segment corresponding to the target structure monitoring data segment. The estimation module 1206 estimates the probability density function of the reference structure monitoring data segment and determines the set of probability density functions that contain probability density functions in the same time window; Prediction module 1208, based on a regression model, uses the probability density functions in the general probability density function set to predict the probability density function of the target structure monitoring data segment; The probability distribution repair module 1210 repairs the probability distribution of the target structure monitoring data segment based on the probability density function of the target structure monitoring data segment.
[0112] The aforementioned probability distribution repair device can be in the form of computer software, and each module of the aforementioned probability distribution repair device can be implemented through computer software modules.
[0113] The specific implementation process of the functions and roles of each module in the above probability distribution repair device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0114] This disclosure also provides an electronic device. Figure 13 A schematic diagram of the hardware implementation using the processing system is shown.
[0115] The hardware structure of electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnect buses and bridges, depending on the specific application and overall design constraints of the hardware. Bus 1100 connects various circuits including one or more processors 1200, memory 1300, and / or hardware modules. Bus 1100 can also connect various other circuits 1400 such as peripherals, voltage regulators, power management circuits, external antennas, etc. Bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one connection line is used in this figure, but this does not indicate that there is only one bus or one type of bus.
[0116] For ease of explanation, certain steps of the above method are described in relation to modules. It should be understood that the corresponding module performing one or more steps of the above method may be one or more hardware modules specifically configured to perform the corresponding step, or implemented by a processor configured to perform the corresponding step, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination thereof.
[0117] This disclosure also provides a readable storage medium storing a computer program that, when executed by a processor, is used to implement the methods described above. A "readable storage medium" can be any means capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples of a readable storage medium include: an electrical connection with one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable read-only memory (CDROM), etc.
[0118] This disclosure also provides a computer program product, the methods of which can be implemented wholly or partially through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, all or part of the processes or functions of this disclosure are performed.
[0119] Computer programs or instructions can be stored in a readable storage medium or transferred from one readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The readable storage medium can be any available medium capable of access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video optical disc; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or it can include both volatile and non-volatile types of storage media.
[0120] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0121] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0122] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0123] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0124] In the description of this specification, the references to terms such as "one embodiment / mode," "some embodiments / modes," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, or characteristic described in connection with that embodiment / mode or example is included in at least one embodiment / mode or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / mode or example. Moreover, the specific features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments / modes or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments / modes or examples described in this specification, as well as the features of different embodiments / modes or examples.
[0125] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0126] Those skilled in the art should understand that the above embodiments are merely for illustrating the present disclosure and are not intended to limit the scope of the disclosure. Those skilled in the art can make other changes or modifications based on the above disclosure, and these changes or modifications still fall within the scope of the present disclosure.
Claims
1. A probability distribution repair method based on improved Wasserstein regression, characterized in that, include: For the same structure within a target time period, structural monitoring data from different sensors are acquired. The structural monitoring data includes first structural monitoring data with missing data and second structural monitoring data without missing data. The first structural monitoring data and the second structural monitoring data are divided into equal segments to determine the target structural monitoring data segment containing missing data and the reference structural monitoring data segment corresponding to the target structural monitoring data segment. The reference structural monitoring data segment contains the second structural monitoring data that is in the same time window as the target structural monitoring data segment. The probability density function is estimated for the monitoring data segment of the reference structure to determine the set of probability density functions that contain probability density functions in the same time window; Based on a regression model, the probability density function of the target structure monitoring data segment is predicted using the probability density functions in the probability density function set. The regression model is an improved Wasserstein regression model based on spatial transformation of the probability density function using logarithmic and exponential mappings. Based on the probability density function of the target structure monitoring data segment, the probability distribution of the target structure monitoring data segment is repaired.
2. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, Based on a regression model, the probability density function of the target structure monitoring data segment is predicted using probability density functions from the set of probability density functions, including: Logarithmic mapping is performed on the probability density functions in the probability density function set to obtain the characterization function of the probability density function in the probability density function set; The representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment; An exponential mapping is performed on the characterization function of the target structure monitoring data segment to determine the probability density function of the target structure monitoring data segment.
3. The probability distribution repair method based on improved Wasserstein regression as described in claim 2, characterized in that, The representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment, including: The response variable function and covariate function in the regression model are expanded by B-spline basis functions to reduce dimensionality. The response variable function is determined based on the characterization function of the target structure monitoring data segment, and the covariate function is determined based on the characterization function of the probability density function in the probability density function set. The coefficient vector of the response variable function is estimated using functional partial least squares regression, and the characterization function of the target structure monitoring data segment is determined.
4. The probability distribution repair method based on improved Wasserstein regression as described in claim 2, characterized in that, To determine the representation function of the probability density function in the set of probability density functions by performing a logarithmic mapping on the probability density functions, the following steps are included: Based on the probability density function set, a reference probability distribution is determined; Calculate the probability density function in the probability density function set and the cumulative distribution function of the reference probability distribution; Based on the cumulative distribution function, a transfer mapping from the reference probability distribution to the probability density functions in the probability density function set is determined; Based on the difference between the transport map and the identity map, the characterization function of the probability density function in the probability density function set is determined.
5. The probability distribution repair method based on improved Wasserstein regression as described in claim 4, characterized in that, The reference probability distribution is the Frieser mean of the probability density function of the probability density function set in Wasserstein space.
6. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, The first structural monitoring data and the second structural monitoring data are each divided into equal segments to determine the target structural monitoring data segment containing missing data and the reference structural monitoring data segment corresponding to the target structural monitoring data segment, including: Align the first structure monitoring data and the second structure monitoring data on the time axis; The aligned first and second structure monitoring data are divided into multiple consecutive time segments of equal length, and the corresponding first and second structure monitoring data segments are determined. The first structure monitoring data segment is filtered to determine the target structure monitoring data segment containing missing data; Based on the time window in which the target structure monitoring data segment is located, the second structure monitoring data segment is filtered to determine the reference structure monitoring data segment.
7. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, Also includes: Based on the probability density function of the target structure monitoring data segment, and in conjunction with the probability density functions of other sensors in the same time period, a joint conditional distribution function of the sensor in the target structure monitoring data segment is established, and data points are further generated. The number of data points is the same as the number of sampling points in the target structure monitoring data segment, and the sampling frequency of the data points is consistent with that of the target structure monitoring data segment. The data formed by combining the data points is inserted into the target structure monitoring data segment to achieve data repair of the target structure monitoring data segment.
8. The probability distribution repair method based on improved Wasserstein regression as described in claim 7, characterized in that, Based on the probability density function of the target structure monitoring data segment, the physical laws of the structure are used as sampling constraints. Data points are generated by the conditional Monte Carlo sampling method. The sampling constraints include the rate of change of adjacent data points not exceeding the maximum acceleration response threshold of the structure, local energy conservation conditions, and / or the frequency domain power spectral density being within the target bandwidth.
9. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes the execution instructions stored in the memory, causing the processor to perform the probability distribution repair method based on improved Wasserstein regression as described in any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the probability distribution repair method based on improved Wasserstein regression as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Bayesian model updating method and system for structural damage identification
CN114511088A
Structural health monitoring missing data reconstruction method based on WGANGP-Unet
CN116502060A
Power grid measurement data checking method and system based on artificial intelligence algorithm
CN118260541A
Data prediction method based on kernel regression primary function Bayesian dynamic linear model
CN119622655A
Distributed photovoltaic grid-connected regional power grid real-time monitoring and coordination control system
CN121238704A