Probabilistic distribution repairing method based on improved wasserstein regression
By improving the Wasserstein regression method and utilizing regression models with logarithmic and exponential mappings, the problem of data reconstruction when wireless sensor data is lost is solved. This achieves multi-source probability distribution fusion and high-fidelity repair, improving the accuracy and consistency of data recovery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-10
AI Technical Summary
Existing data restoration methods cannot effectively reflect the spatial coupling relationship and overall dynamic characteristics between multiple points of a structure, and cannot guarantee that the restored data has the same probability distribution as the original data. In particular, it is difficult to achieve high-fidelity reconstruction when wireless sensor data is lost.
An improved Wasserstein regression method is adopted. By dividing the sensor data into equal segments, a set of probability density functions is constructed. Then, by using regression models with logarithmic and exponential mappings, the probability density function of the target structure monitoring data segment is predicted, realizing the fusion and spatial transformation of multi-source probability distributions and reconstructing missing data.
It improves the accuracy and statistical consistency of missing data reconstruction, ensuring that the repaired data has the same probability distribution as the original data, and is suitable for data recovery and structural state representation in complex missing data scenarios.
Smart Images

Figure CN121542584B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a probability distribution repairing method based on improved Wasserstein regression. BACKGROUND
[0002] In the field of structural health monitoring, wireless sensors are prone to data loss due to harsh environment, and are accompanied by missing probability distribution information, while the probability distribution is crucial for structural performance evaluation. The existing data repairing method only repairs sample points through historical information of a single sensor, which is difficult to reflect the spatial coupling relationship between multiple points of the structure and the overall dynamics characteristics, and cannot guarantee that the repaired data and the original data have the same probability distribution. SUMMARY
[0003] The present disclosure provides a probability distribution repairing method based on improved Wasserstein regression.
[0004] According to one aspect of the present disclosure, a probability distribution repairing method is provided, comprising: acquiring structural monitoring data of different sensors for the same structure in a target time period, wherein the structural monitoring data comprises first structural monitoring data with missing data and second structural monitoring data without missing data; equally segmenting the first structural monitoring data and the second structural monitoring data respectively, to determine a target structural monitoring data segment containing missing data and a reference structural monitoring data segment corresponding to the target structural monitoring data segment, wherein the reference structural monitoring data segment contains second structural monitoring data in the same time window as the target structural monitoring data segment; performing probability density function estimation on the reference structural monitoring data segment to determine a probability density function set; predicting the probability density function of the target structural monitoring data segment based on the probability density functions in the probability density function set through a regression model, wherein the regression model is an improved Wasserstein regression model based on logarithmic mapping and exponential mapping for spatial conversion of the probability density function; and repairing the probability distribution of the target structural monitoring data segment based on the probability density function of the target structural monitoring data segment.
[0005] According to the probability distribution repair method based on improved Wasserstein regression according to an aspect, target structural monitoring data segments containing missing data and corresponding reference structural monitoring data segments are obtained by equally dividing and segmenting structural monitoring data of different sensors; a set of probability density functions is constructed using the reference structural monitoring data segments; the probability density functions of the target structural monitoring data segments are predicted by combining a multi-function-function regression based on functional partial least squares regression (FPLS) and a regression model of logarithmic mapping and exponential mapping, thereby realizing probability distribution repair based on multi-source probability distribution fusion and Wasserstein space conversion, and improving the accuracy and statistical consistency of missing data reconstruction.
[0006] According to the probability distribution repair method based on improved Wasserstein regression according to at least one embodiment of the present disclosure, the probability density functions of the target structural monitoring data segments are predicted by the probability density functions in the set of probability density functions based on a regression model, including: performing logarithmic mapping on the probability density functions in the set of probability density functions to determine the representation functions of the probability density functions in the set of probability density functions; inputting the representation functions of the probability density functions in the set of probability density functions into the regression model to determine the representation function of the target structural monitoring data segments; and performing exponential mapping on the representation function of the target structural monitoring data segments to determine the probability density function of the target structural monitoring data segments.
[0007] According to the probability distribution repair method based on improved Wasserstein regression according to at least one embodiment of the present disclosure, inputting the representation functions of the probability density functions in the set of probability density functions into the regression model to determine the representation function of the target structural monitoring data segments includes: using B-spline basis functions to perform dimension reduction expansion on a response variable function and a covariate function in the regression model, the response variable function being determined based on tangent vectors of the target structural monitoring data segments, and the covariate function being determined based on the representation functions of the probability density functions in the set of probability density functions; and estimating a coefficient vector of the response variable function using a functional partial least squares regression method to determine the representation function of the target structural monitoring data segments.
[0008] According to the probability distribution repairing method based on improved Wasserstein regression, the probability density functions in the set are logarithmically mapped, a representation function of the probability density functions in the set is determined, including: determining a reference probability distribution based on the probability density functions in the set; calculating a cumulative distribution function of the probability density functions in the set and the reference probability distribution; determining a transport mapping from the reference probability distribution to the probability density functions in the set based on the cumulative distribution function; and determining the representation function of the probability density functions in the set based on a difference between the transport mapping and an identity mapping.
[0009] According to the probability distribution repairing method based on improved Wasserstein regression, the reference probability distribution is a Frechet mean of the probability density functions in the set in a Wasserstein space.
[0010] According to the probability distribution repairing method based on improved Wasserstein regression, the first structure monitoring data and the second structure monitoring data are respectively equally divided into segments, a target structure monitoring data segment containing missing data and a reference structure monitoring data segment corresponding to the target structure monitoring data segment are determined, including: aligning the first structure monitoring data and the second structure monitoring data on a time axis; dividing the aligned first structure monitoring data and second structure monitoring data into a plurality of time segments of equal length and continuity, and determining corresponding first structure monitoring data segments and second structure monitoring data segments; screening the first structure monitoring data segments to determine the target structure monitoring data segment containing missing data; and screening the second structure monitoring data segments based on a time window in which the target structure monitoring data segment is located to determine the reference structure monitoring data segment.
[0011] According to the probability distribution repairing method based on improved Wasserstein regression, further comprising: based on the probability density function of the target structure monitoring data segment, jointly other sensors in the same period of the probability density function, establishing a joint conditional distribution function of the sensor in the target structure monitoring data segment, further generating data points, the number of data points is the same as the number of sampling points of the target structure monitoring data segment, and the sampling frequency of the data points and the target structure monitoring data segment is consistent; inserting the data combined from the data points into the target structure monitoring data segment to realize the probability distribution repairing of the target structure monitoring data segment.
[0012] According to the probability distribution repairing method based on improved Wasserstein regression, the probability density function of the target structure monitoring data segment is determined, and a structure physical law is taken as a sampling constraint to generate data points by a conditional Monte Carlo sampling method, the sampling constraint including that a variation rate of adjacent data points does not exceed a maximum acceleration response threshold of the structure, a local energy conservation condition and / or a frequency domain power spectral density within a target bandwidth range.
[0013] According to another aspect of the present disclosure, an electronic device is provided, including a memory storing execution instructions, and a processor executing the execution instructions stored in the memory, so that the processor executes the probability distribution repairing method based on improved Wasserstein regression of any embodiment of the present disclosure.
[0014] According to still another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the probability distribution repairing method based on improved Wasserstein regression of any embodiment of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description, explain the principles of the present disclosure, in which the drawings are included to provide further understanding of the present disclosure and constitute a part of the description.
[0016] Figure 1 is a schematic diagram of the overall flow of the probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0017] Figure 2 is a schematic diagram of determining the reference structure monitoring data segment in the probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0018] Figure 3 is a schematic diagram of determining the probability density function of the target structure monitoring data segment in the probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0019] Figure 4 is a schematic diagram of determining the representation function of the probability density function in the probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0020] Figure 5is a flowchart of determining a representation function of a target structure monitoring data segment in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0021] Figure 6 is a flowchart of probability distribution repairing in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0022] Figure 7 is a schematic diagram of multi-sensor data continuous missing and corresponding data segment probability distribution estimation in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0023] Figure 8 is a schematic diagram of an experimental data scenario in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0024] Figure 9 is a schematic diagram of working condition 1 missing probability density function recovery result in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0025] Figure 10 is a schematic diagram of working condition 2 missing probability density function recovery result in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0026] Figure 11 is a schematic diagram of working condition 3 missing probability density function recovery result in a probability distribution repairing method based on improved Wasserstein regression according to one embodiment of the present disclosure.
[0027] Figure 12 is a schematic structural block diagram of a probability distribution repairing apparatus according to one embodiment of the present disclosure.
[0028] Figure 13 is a schematic structural block diagram of an electronic device according to one embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The present disclosure will be further described below in conjunction with the accompanying drawings and examples. It can be understood that the specific examples described herein are only for the purpose of explanation and are not limitations of the present disclosure. In addition, it should be noted that only parts related to the present disclosure are shown in the accompanying drawings for the convenience of description.
[0030] It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The technical solutions of the present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0031] In a large bridge structure health monitoring system, 8 to 12 wireless acceleration sensors are usually deployed along the length direction of the bridge deck to continuously collect vibration responses under environmental excitations. However, in actual operation, due to radio interference, power fluctuations or hardware aging, data loss of a single sensor for several hours or even several days often occurs, and in extreme cases, 2 to 3 sensors in the same span will fail at the same time. Since the structure does not deteriorate in the short term, there is a stable spatial correlation between the data of each sensor, but the existing technology can only use the probability distribution regression of a single intact sensor to obtain the probability distribution of the target missing sensor, and cannot fully utilize the data information when multiple adjacent sensors are intact.
[0032] Therefore, the present disclosure proposes a probability distribution repair method based on improved Wasserstein regression. Under the premise that the structure is stable in the short term, by segmenting the monitoring data of each sensor and estimating its probability density function, the distribution characteristics are converted to the tangent space for function type regression modeling through logarithmic mapping, the probability distribution information of multiple intact sensors is fully integrated to predict the distribution form of the target structure monitoring data segment, and then the exponential mapping is restored and the repaired data conforming to the statistical characteristics are generated, so that when a single or multiple adjacent sensors fail, the joint distribution information of other sensors in the same span can still be effectively utilized to realize high-fidelity repair, and the data recovery capability in complex missing scenarios and the accuracy of structure state representation are improved.
[0033] The probability distribution repair method of the present disclosure can not only be deployed in cloud servers to realize centralized intelligent operation and maintenance of large-scale infrastructure, but also be integrated into edge computing gateways or on-site monitoring terminals to realize low-latency real-time repair. In the health monitoring system of key infrastructure such as large bridges, high-rise buildings, dams or wind turbine towers, the probability distribution repair method can be run on embedded edge devices deployed on site, using locally collected multi-sensor data streams to autonomously complete the probability distribution level repair of missing data of key measurement points in the case of communication interruption or limited bandwidth, ensuring the continuity and reliability of the state evaluation module. At the same time, the probability distribution repair method can also be integrated as a core algorithm in mobile inspection terminals (such as industrial tablets or smartphones) for engineering technicians to retrieve sensor fragment data through wireless connection on site, instantly perform probability distribution repair and visual analysis, and improve operation and maintenance efficiency. In addition, in the city-level structure group monitoring platform of smart city construction, the probability distribution repair method can be deployed in the central server to batch process historical and real-time data of a large number of distributed sensor networks, build a collaborative repair model across structures and regions, serve the data consistency maintenance of the city safety early warning and digital twin system, and realize the expansion application from single repair to group intelligent diagnosis.
[0034] Figure 1 is a schematic diagram of the overall process of the probability distribution repair method according to an embodiment of the present disclosure. As shown in the method M100 includes steps S110 to S150. The method can be executed by an electronic device such as a mobile phone or a tablet computer. Figure 1
[0035] In step S110, for the same structure in the target time period, the structure monitoring data of different sensors is obtained, including the first structure monitoring data with missing data and the second structure monitoring data without missing data.
[0036] For the same engineering structure in the target time period, the structure monitoring data of multiple sensors is collected, including the first structure monitoring data with sampling missing or communication interruption in part of the time period, and the second structure monitoring data without data loss and with good quality in the same time period.
[0037] Exemplarily, in the actual operation scene of large bridge health monitoring, due to the complex field environment, unstable power supply or equipment aging, it is a common phenomenon that some sensors have periodic data loss. For example, during the passage of a typhoon, multiple vibration sensors in a certain span may lose contact with the wireless module due to strong winds, at which time continuous data of sensors in other unaffected areas can still be obtained. In addition, in the dense monitoring network of the elevated section of urban rail transit, hundreds of sensors work synchronously, and manual checking of data integrity is low in efficiency. Batch scanning of each channel data stream can be realized through automatic script, and the first structure monitoring data and the second structure monitoring data are automatically marked based on the set threshold, so as to realize an efficient and standardized data preprocessing process.
[0038] Optionally, the structure monitoring data is synchronous time series data obtained by sensors deployed on engineering structures such as bridges, buildings or dams, used to analyze the correlation between different measuring points, and as the basis input for probability density function estimation and probability distribution repair, reflecting the state of dynamic response of sensors under environmental excitation and external load, including measurement values of physical quantities such as acceleration, strain, displacement, inclination and / or temperature. These data record the vibration mode, stress distribution and deformation characteristics of the structure during service, and are the core basis for evaluating the health state of the structure, identifying damage and warning potential safety hazards.
[0039] In step S120, the first structure monitoring data and the second structure monitoring data are respectively divided into equal segments, and a target structure monitoring data segment containing missing data and a reference structure monitoring data segment corresponding to the target structure monitoring data segment are determined. The reference structure monitoring data segment contains the second structure monitoring data in the same time window as the target structure monitoring data segment.
[0040] The first structure monitoring data and the second structure monitoring data are synchronized and aligned on the time axis, and then divided into multiple time segments with equal length and continuity, respectively, to form one-to-one corresponding first structure monitoring data segments and second structure monitoring data segments. Each structure monitoring data segment corresponds to a fixed length observation window (such as 1 hour, 6 hours or 1 day), and all sensors have the same time range in the same serial number of structure monitoring data segments. The first structure monitoring data segment contains at least one target structure monitoring data segment which is incomplete due to original sampling loss, and the target structure monitoring data segment is the object of subsequent probability density function prediction and probability distribution repair. The second structure monitoring data segment contains a reference structure monitoring data segment in the same time window as the target structure monitoring data segment.
[0041] In step S130, the reference structure monitoring data segment is subjected to probability density function estimation to determine a set of probability density functions.
[0042] The reference structure monitoring data segment is converted into a probability distribution representation with geometric meaning, and a probability density function set reflecting the cooperative response characteristics is constructed. The probability density function (PDF) of the reference structure monitoring data segment is estimated, and the estimation results of the reference structure monitoring data segment are organized into a time series set to form a probability density function set for predicting the distribution of the target structure monitoring data segment. Each probability density function in the probability density function set corresponds to the probability distribution of a reference sensor in a time window, and constitutes the input variable of the regression model. The reference sensor is a sensor that collects complete structure monitoring data, i.e., the second structure monitoring data.
[0043] Optionally, based on the time segment index, each probability density function in the probability density function set is structured, organized and stored, wherein each probability density function in the probability density function set contains the identification of the corresponding reference sensor.
[0044] In step S140, based on the regression model, the probability density function of the target structure monitoring data segment is predicted by the probability density functions in the probability density function set. The regression model is an improved Wasserstein regression model based on the space conversion of the probability density function by logarithmic mapping and exponential mapping.
[0045] The probability density functions corresponding to each reference sensor in the probability density function set are converted in space by logarithmic mapping. The converted probability density function set is input into the regression model to learn the mapping relationship between the representation functions of multiple reference sensors and the corresponding representation functions of the target sensor in the same time window, and to obtain the representation function of the target structure monitoring data segment. The representation function of the target structure monitoring data segment obtained by prediction is re-mapped from the tangent space back to the Wasserstein space by performing an exponential mapping operation, and is restored to the probability density function of the target structure monitoring data segment. The target sensor is a sensor that collects structure monitoring data with missing data.
[0046] Preferably, the regression model is an improved Wasserstein regression model based on the space conversion of the probability density function by logarithmic mapping and exponential mapping, and the multiple-function-function regression of the functional partial least squares regression (FPLS).
[0047] Preferably, since structural deterioration or changes in structural system boundary conditions usually do not occur in the short term, the spatial correlation between the structure monitoring data recorded by the sensors also does not change significantly in the short term. Therefore, it is assumed that the spatial correlation between the probability density functions of different sensors in the probability density function set of different time periods is consistent, and the probability density function set is used as a functional sample to train the multiple probability distribution to distribution regression model.
[0048] Preferably, the regression model is constructed based on the geometric framework of Wasserstein regression, which transforms the complex distribution-to-distribution prediction problem into a functional regression problem in the tangent space by introducing the logarithmic mapping and exponential mapping operations of the reference distribution in the Wasserstein space, and realizes efficient and geometric consistent prediction of the structural response distribution.
[0049] Preferably, the set of probability density functions is mapped to the tangent space by logarithmic mapping to transform into ordinary functions, called characteristic functions, forming a set of characteristic functions. In the tangent space, a multi-function-function regression based on functional partial least squares (FPLS) is used to predict the characteristic function of the target structural monitoring data segment at the same time period by the characteristic functions of multiple sensors in the set of characteristic functions at the same time period. The predicted characteristic function of the target structural monitoring data segment at the same time period is mapped back to the probability space by exponential mapping to complete the repair of the probability distribution function of the target structural data segment.
[0050] In step S150, the probability distribution of the target structural monitoring data segment is repaired based on the probability density function of the target structural monitoring data segment.
[0051] The predicted probability density function is directly output and stored as the repaired probability distribution as a mathematical representation of the statistical law of the target structural monitoring data segment, thereby realizing high-fidelity reconstruction of the amplitude distribution characteristics (such as kurtosis, skewness, and energy concentration interval) of the structural response and restoring the complete distribution pattern of the target structural monitoring data segment in the probability space. It can be used for subsequent health state assessment, anomaly detection, multi-source data consistency or data repair analysis tasks based on distribution characteristics.
[0052] For the multi-sensor monitoring data of the same structure in the target time period, the first structural monitoring data with missing data and the second structural monitoring data with complete data are distinguished, and the first structural monitoring data and the second structural monitoring data are equally divided into segments to form a first structural monitoring data set and a second structural monitoring data set containing the target structural monitoring data segment. Based on the reference structural monitoring data segment, a set of probability density functions is constructed, and the set is used to predict the probability density function of the target structural monitoring data segment by a space conversion mechanism based on logarithmic mapping and exponential mapping under the framework of Wasserstein regression based on multi-function-function regression of functional partial least squares. Based on the probability density function of the target structural monitoring data segment, time series data is generated to complete the probability distribution repair of the target structural monitoring data segment, and high-fidelity repair of the structural monitoring data is realized based on multi-source distribution fusion, geometric consistent mapping and distribution level reconstruction.
[0053] Regarding step S120, in some embodiments of the present disclosure, it can include, for example, Figure 2The steps S1201 to S1204 are shown.
[0054] In step S1201, the first structural monitoring data and the second structural monitoring data are aligned on a time axis.
[0055] The timestamp information collected by each sensor is obtained, the time reference (such as UTC, local system time or GPS time) is identified, and all the structural monitoring data of the sensors are unified to a common time axis through interpolation, resampling or offset correction. The first structural monitoring data and the second structural monitoring data are aligned in the same time resolution and start and end time, so as to ensure that the structural monitoring data segments contained by each sensor have strictly consistent time ranges in the subsequent equal division segmentation operation, thereby supporting the estimation and regression prediction of the probability density function based on the responses of multiple measuring points under the same working condition.
[0056] In step S1202, the aligned first structural monitoring data and second structural monitoring data are divided into a plurality of time segments of equal length and continuity, and corresponding first structural monitoring data segments and second structural monitoring data segments are determined.
[0057] According to the preset time length, the first structural monitoring data and the second structural monitoring data are divided into a series of time segments of equal length, continuity and no overlap or partial overlap. Each time segment corresponds to a fixed time interval, and ensures that all sensors cover completely consistent time ranges in the same serial number segment. Thereby, a one-to-one corresponding first structural monitoring data segment and second structural monitoring data segment are formed, wherein the former contains at least one incomplete target structural monitoring data segment caused by sampling interruption. The division method ensures the response comparability of each sensor under the same working condition, and provides a synchronous input basis for the estimation and regression modeling of the probability density function.
[0058] In step S1203, the first structural monitoring data segment is screened to determine the target structural monitoring data segment containing missing data.
[0059] After the first structural monitoring data is divided into a plurality of continuous and equal length time segments, data quality evaluation is performed on each first structural monitoring data segment; according to the preset missing determination condition (such as the missing point loss rate exceeding the threshold, the length of continuous NaN value reaching the set proportion or the signal amplitude being constant), the segment in which the missing data exists is identified. The data segment satisfying the condition is marked as the target structural monitoring data segment.
[0060] In step S1204, based on the time window in which the target structural monitoring data segment is located, the second structural monitoring data segment is screened to determine the reference structural monitoring data segment.
[0061] For the determined target structural monitoring data segment, the corresponding time window information is obtained. All second structural monitoring data segments of the second structural monitoring data are traversed, and segments that completely coincide or highly overlap with the time range are screened out. These second structural monitoring data segments located in the same time interval are taken as reference structural monitoring data segments for representing the overall dynamic state of the structure in this period. The reference structural monitoring data segments come from other sensors that are spatially adjacent or functionally related to the target structural monitoring data segment, ensuring that their response characteristics are comparable and relevant, thereby supporting the subsequent multi-source fusion modeling in the probability density function estimation and regression prediction tasks.
[0062] The first structural monitoring data and the second structural monitoring data are accurately aligned on the time axis based on timestamp information and unified to the same sampling resolution and time reference. The aligned first structural monitoring data and the second structural monitoring data are divided into multiple time segments of equal length, continuity and strictly consistent time range, forming one-to-one corresponding structural monitoring data segments. On this basis, the target structural monitoring data segment containing sampling interruption or data loss is screened out by integrity evaluation of the first structural monitoring data segment. According to the time window where the target segment is located, the second structural monitoring data segment with complete data in the same period is located and screened out from other sensors as the reference structural monitoring data segment, ensuring that its response characteristics reflect the dynamic behavior under the same external working condition. The comparability of each measuring point in the same time interval is ensured.
[0063] Regarding step S140, in some embodiments of the present disclosure, steps S1401 to S1403 as shown in the following table can be included. Figure 3
[0064] In step S1401, the probability density functions in the probability density function set are logarithmically mapped to determine the representation functions of the probability density functions in the probability density function set.
[0065] The probability density functions that cannot be directly added or regressed are converted into function vectors in the tangent space through logarithmic mapping, so that the classical statistical method can be applied. The logarithmic mapping is based on the optimal transport theory, ensuring that the transformation process follows the shortest path principle and avoids the distribution ambiguity or quality leakage caused by traditional L² space averaging. Moreover, all representation functions share the same reference coordinate system, which facilitates the construction of a unified multiple regression relationship and improves the interpretability of the regression model. Among them, the representation functions usually have small amplitude and smooth changes, which are suitable for basis function expansion and dimensionality reduction, and are beneficial to improve the training speed and convergence of the regression model.
[0066] In step S1402, the representation functions of the probability density functions in the probability density function set are input into the regression model to determine the representation functions of the target structural monitoring data segment.
[0067] The representation function of the probability density function corresponding to each sensor in the probability density function set is input into the regression model trained in advance or constructed in real time. The regression model learns the function mapping relationship between the representation function of the probability density function of each reference sensor in the probability density function set and the representation function of the target structural monitoring data segment of the target sensor based on the historical complete period structural monitoring data. In the current target time period, the regression model outputs the representation function of the target structural monitoring data segment according to the change trend of the probability density function of each reference sensor in the probability density function set. The representation function of the target structural monitoring data segment represents the optimal transmission offset direction and amplitude of the target sensor relative to the reference distribution under the same working condition, and serves as the basis for subsequent recovery of the complete probability density function through exponential mapping.
[0068] In step S1403, the representation function of the target structural monitoring data segment is subjected to exponential mapping to determine the probability density function of the target structural monitoring data segment.
[0069] The exponential mapping follows the geodesic path in the Wasserstein space, ensures that the reconstructed distribution and the reference distribution satisfy the optimal transport principle, and avoids non-physical oscillation or distortion. When the representation function contains multiple local fluctuations, the exponential mapping can naturally evolve into complex distribution patterns such as bimodal, wide-tailed or multimodal, and adapt to the non-stationary response in real engineering. The exponential mapping and the logarithmic mapping make the entire Wasserstein regression framework have mathematical closure and engineering practicability.
[0070] Preferably, the probability density functions in the probability density function set are subjected to logarithmic mapping to obtain the representation functions of the probability density functions in the probability density function set, which are unconstrained and have a linear space structure. The representation functions of the probability density functions in the probability density function set are input into the multi-function-function regression based on functional partial least squares regression established by using historical data to determine the representation function of the target structural monitoring data segment; and the representation function of the target structural monitoring data segment is subjected to exponential mapping to determine the probability density function of the target structural monitoring data segment.
[0071] The log mapping is performed on the probability density functions in the set of probability density functions, which projects them from the nonlinear distribution space to the flat tangent space, and converts them into the operable representation of the characteristic function. This process is based on the optimal transport theory, and unifies all the characteristic functions in the same geometric coordinate system, supporting subsequent multivariate functional regression modeling. These characteristic functions are sent as input variables into the regression model constructed by the multi-function-function regression of the functional partial least squares regression, learning the distribution evolution coupling relationship of the probability density functions of multiple reference sensors and the target sensor under the same working condition, and then predicting the characteristic function of the target structure monitoring data segment, accurately representing the optimal transport offset direction and amplitude of the target structure monitoring data segment relative to the reference state. The exponential mapping operation is performed on the characteristic function of the target structure monitoring data segment, which reconstructs the probability density function of the target structure monitoring data segment. Not only does it ensure the mathematical legitimacy of the output distribution, but it also naturally generates complex morphologies such as bimodal and / or wide tail, and truly reflects the dynamic response characteristics under non-stationary excitation. The target of intelligently inferring the missing position statistical behavior from the distribution evolution law of multiple source sensors is achieved, and the robustness, physical consistency and engineering usability of the structural health monitoring system in the scene of incomplete data are improved.
[0072] Regarding step S1401, in some embodiments of the present disclosure, steps S410 to S440 as shown in FIG. 4 can be included. Figure 4
[0073] In step S410, a reference probability distribution is determined based on the probability density functions in the set of probability density functions.
[0074] Preferably, the reference probability distribution is the Fréchet mean of the probability density functions in the set of probability density functions in the Wasserstein space.
[0075] Specifically, the Fréchet mean of each reference sensor corresponding to the probability density function in the set of probability density functions is calculated in the Wasserstein distance metric, and the probability distribution corresponding to the Fréchet mean is determined as the reference probability distribution. The reference probability distribution serves as the common base point for subsequent log mapping and exponential mapping, and is used to project each probability density function from the Wasserstein manifold to the tangent space, or vice versa to reconstruct the prediction result. The reference distribution embodies the common statistical characteristics of multiple sensors in the same time window, has good representativeness and stability, and supports cross-sensor and cross-period probability distribution comparison and fusion modeling.
[0076] In step S420, the cumulative distribution function of the probability density functions in the set of probability density functions and the reference probability distribution is calculated.
[0077] The cumulative distribution function corresponding to each of the probability density functions in the set of probability density functions and the reference probability distribution is calculated. The cumulative distribution function is continuous and monotonically increasing in the domain. Ensuring all distributions are expressed in a unified function form provides the necessary input basis for subsequent optimal transport theory-based log mapping, supporting cross-sensor quality element matching and deformation field calculation.
[0078] In step S430, a transport mapping from the reference probability distribution to the probability density functions in the set of probability density functions is determined based on the cumulative distribution functions.
[0079] Based on the cumulative distribution function of the reference probability distribution and the cumulative distribution function of the probability density functions in the set of probability density functions, an optimal transport mapping from the reference distribution to the cumulative distribution function of each probability density function is constructed. The transport mapping represents a transformation function that maps each quality point in the reference distribution to the same quantile position in the target distribution, reflecting the local shift behavior of the probability density function relative to the common reference. The optimal transport mapping has strict monotonicity, ensuring mass conservation and non-crossing transport, supporting subsequent extraction of its corresponding representation function through log mapping for multivariate regression modeling.
[0080] In step S440, a representation function of the probability density functions in the set of probability density functions is determined based on the difference between the transport mapping and the identity mapping.
[0081] The point-by-point difference between each transport mapping and the identity mapping is calculated to obtain the representation function of the corresponding probability density function. The non-linear Wasserstein manifold is locally approximated as a flat tangent space, allowing traditional statistical methods to be directly applied to probability distribution data. The original probability density function requires thousands of points for description, while the representation function can be represented by tens of parameters after dimension reduction, facilitating real-time processing and transport.
[0082] Based on the probability density functions in the set, the Frechet mean in the Wasserstein space is calculated, and it is determined as the reference probability distribution, which is the geometric center of the entire distribution set and the common base point of the subsequent mapping operation, ensuring that the modeling process has good representativeness and stability. The cumulative distribution functions (CDFs) of each probability density function and the reference probability distribution are calculated respectively, and the probability distribution is converted into a continuous and monotonically increasing function form, providing a unified and differentiable mathematical basis for optimal transport analysis. Based on the cumulative distribution functions of the reference distribution and each probability density function, the optimal transport mapping from the reference distribution to the target distribution is constructed, which strictly maintains the quality order and accurately describes the nonlinear deformation path of each sensor response relative to the common reference. By calculating the difference between the transport mapping and the identity mapping, the representation function of each probability density function in the tangent space is obtained, realizing the conversion of the distribution offset on the nonlinear manifold into a function type vector that is additive and regressive in the flat space. The Wasserstein geometric processing chain is completely constructed, not only avoiding the distribution ambiguity problem caused by traditional Euclidean average, but also enabling the originally non-directly-operable probability density functions to be subjected to multivariate statistical modeling in a unified coordinate system, improving the accuracy, interpretability and computational efficiency of multi-source distribution data fusion.
[0083] Regarding step S1402, in some embodiments of the present disclosure, steps S510 to S520 as shown can be included. Figure 5
[0084] In step S510, the B-spline basis function is used to reduce the dimension of the response variable function and the covariate function in the regression model. The response variable function is determined based on the representation function of the target structure monitoring data segment, and the covariate function is determined based on the representation function of the probability density function in the set of probability density functions.
[0085] The function type data of thousands of dimensions is compressed into a tens of dimension coefficient vector, which changes the high-dimensional regression problem into a regular multivariate regression, greatly improving the training speed and convergence. The B-spline basis function has strong smoothing and local support characteristics, avoiding high-frequency noise interference, and ensuring that the reconstructed function form is natural and without oscillation. The coefficient vector of the response variable function and the covariate function is small in volume, suitable for edge-cloud collaborative architecture, reduces the bandwidth pressure, and supports remote intelligent operation and maintenance of wide-area infrastructure groups. The output form can be directly connected to mature statistical models such as FPLS, FPCR, GAM, etc., enhancing the method universality and maintainability. Moreover, the basis function expansion itself has a regularization effect, which suppresses overfitting.
[0086] Preferably, the number of basis functions in the regression model is determined by cross-validation, which is used to express the complexity of the regression model. The number of basis functions is used as a core hyperparameter to control the complexity of the regression model. During the training phase of the regression model, a set of candidate numbers of basis functions is traversed using the cross-validation method. The corresponding regression model is constructed for each number of basis functions, and the prediction error on the validation set is calculated. The number of basis functions that minimizes the validation error is selected as the optimal number of basis functions for the construction of the final model. The determined regression model has both sufficient flexibility to capture the spatial evolution characteristics of the characteristic function and effectively suppresses overfitting, thereby improving the prediction stability under unknown working conditions.
[0087] In step S520, the coefficient vector of the response variable function is estimated using the functional partial least squares regression method to determine the characteristic function of the target structure monitoring data segment.
[0088] The B-spline basis function expanded characteristic function is used as the input variable, and the characteristic function of the target structure monitoring data segment is used as the response variable. The coefficient parameters of the response variable function are estimated using the functional partial least squares regression method. The latent variables are extracted by iteration to maximize the covariance between the input and output blocks, and a compact multiple linear regression relationship is established. After the regression model is trained, for a new target structure monitoring data segment, the regression model can automatically predict the B-spline expansion coefficients of the characteristic function of the new target structure monitoring data segment according to the B-spline expansion coefficients of the characteristic function of the reference structure monitoring data segment obtained by the reference sensor at the same time, and then reconstruct the complete characteristic function for subsequent index mapping to restore the probability density function.
[0089] The B-spline basis function is used to reduce the dimension of the response variable function and the correlation function. The function type variable represented by thousands of sampling points is compressed into a low-dimensional vector containing only dozens of coefficients. While significantly reducing the computational load and communication overhead, the key morphological characteristics of the distribution offset are retained. Due to its smoothness and local support characteristics, noise interference is effectively suppressed, ensuring that the reconstructed function is natural and non-oscillatory. The prediction performance of the regression model under different basis function quantities is systematically evaluated by cross-validation method. The minimum validation set error is used as the criterion to automatically determine the optimal number of basis functions, scientifically balancing the model expression ability and generalization ability, avoiding underfitting or overfitting, and improving the robustness of the model under diversified working conditions. The functional partial least squares regression method (FPLS) is used to establish the multivariate regression relationship between the input and output characteristic function coefficients. By extracting the maximum covariance latent variable, the problem of multicollinearity among multiple sources of input is effectively addressed, and the accurate inference from the distribution evolution pattern of the characteristic function of the reference sensor to the characteristic function of the target sensor is realized. The complex functional regression problem is transformed into an efficient solvable numerical task, providing an accurate, stable and physically consistent prediction basis for subsequent reconstruction of the probability density function of the target structural monitoring data segment through exponential mapping.
[0090] In some embodiments of the present disclosure, steps S610 to S620 as shown in Figure 6 may be included.
[0091] In step S610, based on the probability density function of the target structural monitoring data segment, the joint conditional distribution function of the sensor in the target structural monitoring data segment is established in combination with the probability density functions of other sensors in the same period, and data points are further generated. The number of data points is the same as the number of sampling points of the target structural monitoring data segment, and the sampling frequency of the data points is consistent with that of the target structural monitoring data segment.
[0092] Preferably, based on the probability density function of the target structural monitoring data segment, the joint conditional distribution function of the sensor in the target structural monitoring data segment is established in combination with the probability density functions of other sensors in the same period, and the data points are generated by the conditional Monte Carlo sampling method with the structure physical law as the sampling constraint. The sampling constraints include that the variation rate of adjacent data points does not exceed the maximum acceleration response threshold of the structure, the local energy conservation condition and / or the frequency domain power spectral density within the target bandwidth range.
[0093] Specifically, the probability density function of the reconstructed target structure monitoring data segment is exponentially mapped, the probability density functions of other sensors in the same time segment are combined, and a joint conditional distribution function of the sensor in the target structure monitoring data segment is established. Based on this, a set of numerical samples or data points that satisfy the distribution are randomly extracted from the Monte Carlo sampling or inverse transform sampling method. The number of generated data points is strictly equal to the total number of sampling points that should be collected in the normal working state of the time segment. At the same time, the time interval of the data points is consistent with the sampling frequency of the target structure monitoring data segment, forming a complete and equidistant data point sequence. The data point sequence not only restores the data length and time sequence structure, but also truly reflects the statistical characteristics of the structure response in the time segment, and can be used to replace the missing data to participate in subsequent health assessment and data analysis.
[0094] In step S620, the data combined by the data points is inserted into the target structure monitoring data segment to realize data repair of the target structure monitoring data segment.
[0095] The data point sequence is replaced or filled in the target structure monitoring data segment according to its corresponding time window and time sequence order. The insertion operation keeps the original database structure, time stamp format and unit system unchanged, forming a complete structure monitoring data segment that is logically continuous and uninterrupted. At the same time, metadata tags can be attached to identify the repair source, support audit traceability and credibility evaluation, and ensure the availability and legality of the repaired data in various downstream tasks.
[0096] Based on the probability density function of the target structure monitoring data segment reconstructed by exponential mapping, and combined with the probability density functions of other sensors in the same time period, a joint conditional distribution function for the sensor in the target structure monitoring data segment is established. Monte Carlo or inverse transform sampling methods are then used to generate a data point sequence with the same number of sampling points and consistent sampling frequency as the original sampling points, ensuring that the filled data statistically reflects the dynamic laws of the structural response in that time period. Structural physical laws are introduced as conditional sampling constraints. Conditional Monte Carlo sampling controls the rate of change of adjacent data points to not exceed the maximum acceleration response threshold of the structure, satisfies local energy conservation conditions and / or frequency domain power spectral density distribution characteristics, effectively avoiding unreasonable abrupt changes or non-physical interpretations, and improving the temporal smoothness and dynamic realism of the synthesized data. The generated data point sequence is seamlessly inserted into the target structure monitoring data segment according to time windows and temporal sequence, replacing or filling the original missing intervals, maintaining complete consistency with the database structure, timestamp format, and unit system. The source and method of the repair can be identified through additional metadata tags, supporting data traceability and credibility assessment. Not only did it restore the length and temporal continuity of the data, but it also ensured a high degree of consistency in statistical distribution, physical rationality, and engineering usability. This allows the repaired data to be directly used for routine tasks such as subsequent spectrum analysis, damage identification, or over-limit alarms, greatly improving the robustness, intelligence level, and practical operation and maintenance value of the structural health monitoring system in scenarios with missing data.
[0097] In one specific embodiment, such as Figure 7 As shown in sub-diagram (a), it is assumed that the components are installed at different locations on the structure. m There are 10 sensors, of which some sensors intermittently fail (let's say the 10th sensor). h 1 and the h Two sensors were used, and consecutive gaps occurred at different time intervals, resulting in the loss of probability distribution information for the corresponding time periods. (The sentence about the first sensor appears unrelated and likely refers to a separate topic.) h Taking a multi-probability distribution to probability distribution regression task from a single sensor as an example, segments with continuously missing data on this sensor are denoted as... (i.e., the target structure monitoring data segment), its missing data is as follows Figure 7 The sub-diagram shown in (b) is denoted as Except for the first h Other sensors besides the two main sensors serve as collaborating sensors, or reference sensors, used to regress missing data. The probability distribution. The set of indices for collaborative sensors is denoted as... .like Figure 7 As shown in subplot (b), to establish a multi-probability distribution to distribution regression model, each complete data sequence is equally divided into several data segments to construct probability distribution function samples. Figure 7As shown in subplot (c), it is assumed that data within the same data segment follows the same probability distribution, but the probability distributions differ between different data segments. The probability density function is estimated for each data segment without missing data to obtain the probability density function set of the reference structure data segment.
[0098] Given that structural degradation or changes in the boundary conditions of the structural system typically do not occur in the short term, and the spatial correlation between sensor-recorded data does not change significantly in a short period, it is assumed that the probability density functions of each data segment have consistent correlation. The probability density function of each intact data segment (i.e., the reference structure monitoring data segment) is used as a functional sample to train a multi-probability distribution-to-distribution regression model. Therefore, the task of multi-probability distribution-to-distribution regression of multi-sensor data is to identify the missing data segments. Above, the probability distribution of multiple collaborative sensors in the same data segment is utilized. , restore the h Missing data from one sensor Missing probability distribution .like Figure 8 As shown in subgraph (a), for the th h For two sensors, due to the long continuous data loss period, the missing data should first be segmented into shorter fragments, and then processed. h For each missing segment of the two sensors, perform the same multi-distribution to distribution regression task as above.
[0099] First, apply the logarithmic mapping in equation (1) to all probability density functions in the probability density function set to map them from the Wasserstein space to the tangent space, thereby obtaining the unconstrained representation function, i.e., the covariate function. and response variable function Probability density function in a set of probability density functions The logarithmic mapping can be defined as:
[0100] (1)
[0101] In the formula, Describing the probability density function in a set. The representation function in the tangent space, This represents an identity mapping. Is with the first i The cumulative distribution function (CDF) of the Fréchet mean corresponds to a probability distribution sequence (i.e., probability density function). Since we only consider univariate probability distributions here, we can first calculate the quantile function of the Fréchet mean. Then, by inverting the function, the corresponding cumulative distribution function can be obtained. Quantile function It can be calculated according to formula (2).
[0102] (2)
[0103] Since the tangent space is a subspace of the Hilbert space, regression methods for ordinary functional data can be used in this space. To extend Wasserstein regression to regression from multiple probability distributions to probability distributions, this disclosure embeds a multi-function-to-function regression method based on functional partial least squares regression (FPLS) into the tangent space to characterize the mapped covariate function. With response variable function The relationship between them can be expressed in the following form:
[0104] (3)
[0105] in, Indicates the subject With response variable function The regression coefficient function between them; Represents a quadratic term With response variable function The regression coefficient function between them This represents the error term function.
[0106] Since functional data is inherently infinite-dimensional, directly estimating these coefficient functions is computationally expensive and statistically unstable, especially with limited sample sizes. To alleviate this problem, basis function expansion is typically used to reduce dimensionality and improve computational efficiency. B- Spline basis functions are widely used because they combine flexibility, computational efficiency, and local support. B - Construction of spline bases from the objective function Let a sequence of nodes be set as the starting point on the domain of the variable. Let the variable be defined in the interval... Above, its node vector is denoted as ,in d for B - The degree of the spline (i.e., the order of the polynomial). M The number of basis functions, i.e. B - Number of terms in the spline expansion. Based on this sequence of nodes, piecewise polynomial basis functions can be constructed using the Cox-de Boor recurrence relation. Thus, the objective function... It can be approximated as a linear combination of these basis functions:
[0107] (4)
[0108] in, express B -spline basis functions, Let T be the corresponding coefficient vector, and T denote the matrix transpose. This disclosure can be determined through cross-validation. B - The number of terms in the spline expansion is adjusted to achieve an optimal balance between flexibility and simplicity. Therefore, the response variable function, independent variable function, and error function in equation (5) can all be used. B The spline basis function expansion is as follows:
[0109] (5)
[0110] (6)
[0111] (7)
[0112] in, and They are respectively represented as and Selected B -Spline basis functions; and For the corresponding B - Number of terms in the spline expansion; and It is the corresponding expansion coefficient vector; for The expansion coefficients.
[0113] For quadratic terms Its expanded form can be expressed as:
[0114] (8)
[0115] in, , .
[0116] Similarly, the main effect coefficient function and quadratic coefficient function You can also do it through B- Spline expansion takes the following form:
[0117] (9)
[0118]
[0119] (10)
[0120] In the formula, and They represent the first i The function of the main effect coefficient and the ( i , j ) quadratic coefficient function of B- the spline expansion coefficient matrix. and The specific form of the above is as follows:
[0121]
[0122] (11)
[0123] According to the above B- spline expansion process, equation (3) can be rewritten as follows:
[0124]
[0125] (12)
[0126] Simplifying the above equation gives:
[0127] (13)
[0128] where , . Thus, the regression problem from multiple probability distributions to a probability distribution is converted to the problem of solving the coefficient matrix , and of the multiple regression equation (13). However, when using more predictors or using more basis functions, the dimensions of the matrices and will rapidly increase, which may lead to the emergence of ill-conditioned problems (multicollinearity). In order to solve this problem, dimension reduction techniques such as principal component regression (PCR) or partial least squares regression (PLS) can be used. Principal component regression reduces the dimension by extracting principal components from the predictor matrix. However, the extracted principal components do not take into account the correlation between the predictor variables and the response variable, resulting in the inability to determine the effect of each regression parameter on the response variable. While partial least squares regression extracts principal components from both the predictor matrix and the response matrix, maximizing the correlation between these principal components. Then, the predictor matrix and the response matrix are respectively regressed on the extracted principal components. This process is iterated until satisfactory results are obtained through cross-validation. After obtaining the coefficients , and through partial least squares regression, the main effect coefficient function , the quadratic term coefficient function and can be obtained through equations (9), (10) and (7) respectively.
[0129] When the regression model of multi-probability distribution to probability distribution is established as formula (3), the representation function of the segment where the missing data is located, i.e. the response variable function, can be calculated as follows:
[0130] (14)
[0131] Finally, the representation function, i.e. the response variable function, is mapped back to the Wasserstein space from the tangent space by using the exponential mapping, which is defined as follows:
[0132] (15)
[0133] In the formula, φ -1 represents the inverse function of φ. Therefore, the embodiment of the present disclosure is verified by the following experimental data, and the example data is from the wireless monitoring system installed on the pedestrian overpass, and the bridge is a two-span continuous rigid frame bridge located in the campus of a certain university, as shown in the (a) subgraph of FIG. 1 and the (b) subgraph of FIG. 2, the length is 44 meters, and the width is 3.7 meters. The signal-to-noise ratio of the observation data of sensors No. 4 and No. 8 is low, and the present disclosure only uses the monitoring data of accelerometers No. 1, 2, 3, 5, 6 and 7. In the provided data, a total of 17 weeks of acceleration data are recorded, and the present disclosure only uses the acceleration data under full excitation in the third week.
[0134] In order to comprehensively verify the effectiveness of the method proposed in the present disclosure, the following missing working conditions are considered: Figure 8 Figure 9 to Figure 11 Working condition 1: only one sensor has data loss (assuming that sensor 2 has data loss), and the data of the remaining 5 sensors are used for repair;
[0135] Working condition 2: two sensors have data loss at the same time (assuming that sensors 2 and 6 have data loss), and the data of the remaining 4 sensors are used for repair;
[0136] Working condition 3: three sensors have data loss at the same time (assuming that sensors 1, 3 and 6 have data loss), and the data of the remaining 3 sensors are used for repair.
[0137] As shown in FIG. 3, the acceleration data of the first week is used as the training set, and the acceleration data of the second week is used as the test set.
[0138] As shown in FIG. 4, the acceleration data of the first week is used as the training set, and the acceleration data of the second week is used as the test set.
[0139] As shown in FIG. 5, the acceleration data of the first week is used as the training set, and the acceleration data of the second week is used as the test set. Figure 12 As shown, the restored probability density functions of the improved Wasserstein regression (IWR) and the original Wasserstein regression (WR) are compared under three missing data scenarios. In each figure, the solid line represents the result of the method of the present disclosure, and the dashed line represents the result of the original Wasserstein regression. In addition, a to b in the legend means that the original Wasserstein regression restores the probability distribution of sensor b based on the data of sensor a.
[0140] The average IAE values of the three-segment restored probability density functions under different scenarios are shown in the following table. Under three different working conditions, when a specific faulty sensor fails, the probability density function of the faulty sensor is restored by the remaining different sensors. Among them, in working condition 1, the average integral absolute error (IAE) of the improved Wasserstein regression is only 0.0260, which is much lower than the minimum error of any single-sensor restoration, i.e. 0.0700; in the more complex multi-sensor simultaneous failure scenario (such as working condition 3), the improved Wasserstein regression still maintains a stable low error (i.e. 0.0553, 0.0430, 0.0353), while the original Wasserstein regression method has large error fluctuations and poor accuracy due to its dependence on a single sensor. This fully proves that the improved Wasserstein regression can more accurately restore the missing probability distribution by fusing multi-sensor information and preserving the distribution geometric structure, effectively improving the data integrity of structural health monitoring.
[0141]
[0142] Based on any one of the above embodiments, the present disclosure also provides a probability distribution restoration device.
[0143] Figure 12 is a structural schematic diagram of a probability distribution restoration device according to an embodiment of the present disclosure.
[0144] As Figure 13 shown, the probability distribution restoration device comprises:
[0145] The data acquisition module 1202 acquires structural monitoring data of different sensors for the same structure within a target time period, and the structural monitoring data includes first structural monitoring data with missing data and second structural monitoring data without missing data;
[0146] The data segmentation module 1204 equally segments the first structural monitoring data and the second structural monitoring data respectively, determines a target structural monitoring data segment containing missing data and a reference structural monitoring data segment corresponding to the target structural monitoring data segment;
[0147] The estimation module 1206 estimates the probability density function of the reference structure monitoring data segment to determine a probability density function set containing probability density functions in the same time window;
[0148] The prediction module 1208 predicts the probability density function of the target structure monitoring data segment based on the regression model and the probability density functions in the probability density function set;
[0149] The probability distribution repairing module 1210 repairs the probability distribution of the target structure monitoring data segment based on the probability density function of the target structure monitoring data segment.
[0150] The above-mentioned probability distribution repairing apparatus can be in the form of computer software, and each module of the above-mentioned probability distribution repairing apparatus can be implemented by a computer software module.
[0151] The implementation process of the functions and roles of each module in the above-mentioned probability distribution repairing apparatus is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be described here.
[0152] The present disclosure also provides an electronic device. Figure 1 A schematic diagram showing a hardware implementation of a processing system is shown.
[0153] The hardware structure of the electronic device 1000 can be implemented by using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the particular application of the hardware and overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, memories 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc. The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component bus (PCI), or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one connection line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0154] For ease of illustration, some steps of the above-mentioned method are described in correspondence with modules. It should be understood that the corresponding module performing one or more steps of the above-mentioned method can be one or more hardware modules specially configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer readable medium for implementation by a processor, or implemented by some combination.
[0155] The present disclosure also provides a readable storage medium, and the readable storage medium stores a computer program. The computer program is executed by a processor to implement the method described above. The readable storage medium can be any device that can contain, store, communicate, propagate or transport a program for use by or in connection with an instruction execution system, apparatus or device. More specific examples of the readable storage medium include the following: an electrical connection having one or more wires (electronic devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CD-ROM), and the like.
[0156] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. When implemented by software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are wholly or partially executed.
[0157] The computer program or instructions can be stored in a readable storage medium or transmitted from one readable storage medium to another readable storage medium, for example, the computer program or instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired or wireless means. The readable storage medium can be any available medium that can be accessed or a data storage device such as a server, data center, etc. that integrates one or more available media. The available media can be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; an optical medium, such as a digital video disc; and a semiconductor medium, such as a solid state disk. The computer readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile storage media.
[0158] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.
[0159] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more computer-readable media. Figure 1 one or more computer-readable media.
[0160] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more computer-readable media. Figure 1 one or more computer-readable media.
[0161] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more computer-readable media. one or more computer-readable media.
[0162] In the description of the specification, the description of the terms "one embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples" and the like means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment / way or example. Also, the specific features, structures, or characteristics described can be combined in any appropriate manner in one or more embodiments / ways or examples. In addition, the person skilled in the art can combine and combine the different embodiments / ways or examples described in the specification and the features of the different embodiments / ways or examples without contradiction with each other.
[0163] Furthermore, the terms "first", "second", etc. are used herein for descriptive purposes only and should not be construed as indicating or implying relative importance or an ordered ranking such that the technical features so designated should possess. Thus, features having a "first", "second", etc. designation can implicitly or explicitly include at least one of the features. In the description of the disclosure, the meaning of "a plurality" is at least two, for example, two, three, etc., unless otherwise specifically defined.
[0164] Those skilled in the art will understand that the above-described embodiments are merely for the purpose of clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. Other changes or modifications can be made by those skilled in the art based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A probability distribution repair method based on improved Wasserstein regression, characterized in that, include: For the same structure within a target time period, structural monitoring data from different sensors are acquired. The structural monitoring data includes first structural monitoring data with missing data and second structural monitoring data without missing data. The first structural monitoring data and the second structural monitoring data are divided into equal segments to determine the target structural monitoring data segment containing missing data and the reference structural monitoring data segment corresponding to the target structural monitoring data segment. The reference structural monitoring data segment contains the second structural monitoring data that is in the same time window as the target structural monitoring data segment. The probability density function is estimated for the monitoring data segment of the reference structure to determine the set of probability density functions that contain probability density functions in the same time window; Based on a regression model, the probability density function of the target structure monitoring data segment is predicted using the probability density functions in the probability density function set. The regression model is an improved Wasserstein regression model based on spatial transformation of the probability density function using logarithmic and exponential mappings. Based on the probability density function of the target structure monitoring data segment, the probability distribution of the target structure monitoring data segment is repaired; Specifically, based on a regression model, the probability density function of the target structure monitoring data segment is predicted using probability density functions from the set of probability density functions, including: Logarithmic mapping is performed on the probability density functions in the probability density function set to obtain the characterization function of the probability density function in the probability density function set; The representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment; The probability density function of the target structure monitoring data segment is determined by performing an exponential mapping on the characterization function of the target structure monitoring data segment. It also includes: dividing the missing data with a long continuous missing time in the first structure monitoring data into shorter target structure monitoring data segments, and performing probability distribution repair based on the shorter target structure monitoring data segments.
2. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, The representation function of the probability density function in the probability density function set is input into the regression model to determine the representation function of the target structure monitoring data segment, including: The response variable function and covariate function in the regression model are expanded by B-spline basis functions to reduce dimensionality. The response variable function is determined based on the characterization function of the target structure monitoring data segment, and the covariate function is determined based on the characterization function of the probability density function in the probability density function set. The coefficient vector of the response variable function is estimated using functional partial least squares regression, and the characterization function of the target structure monitoring data segment is determined.
3. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, To determine the representation function of the probability density function in the set of probability density functions by performing a logarithmic mapping on the probability density functions, the following steps are included: Based on the probability density function set, a reference probability distribution is determined; Calculate the probability density function in the probability density function set and the cumulative distribution function of the reference probability distribution; Based on the cumulative distribution function, a transfer mapping from the reference probability distribution to the probability density functions in the probability density function set is determined; Based on the difference between the transport map and the identity map, the characterization function of the probability density function in the probability density function set is determined.
4. The probability distribution repair method based on improved Wasserstein regression as described in claim 3, characterized in that, The reference probability distribution is the Frieser mean of the probability density function of the probability density function set in Wasserstein space.
5. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, The first structural monitoring data and the second structural monitoring data are each divided into equal segments to determine the target structural monitoring data segment containing missing data and the reference structural monitoring data segment corresponding to the target structural monitoring data segment, including: Align the first structure monitoring data and the second structure monitoring data on the time axis; The aligned first and second structure monitoring data are divided into multiple consecutive time segments of equal length, and the corresponding first and second structure monitoring data segments are determined. The first structure monitoring data segment is filtered to determine the target structure monitoring data segment containing missing data; Based on the time window in which the target structure monitoring data segment is located, the second structure monitoring data segment is filtered to determine the reference structure monitoring data segment.
6. The probability distribution repair method based on improved Wasserstein regression as described in claim 1, characterized in that, Also includes: Based on the probability density function of the target structure monitoring data segment, and in conjunction with the probability density functions of other sensors in the same time period, a joint conditional distribution function of the sensor in the target structure monitoring data segment is established, and data points are further generated. The number of data points is the same as the number of sampling points in the target structure monitoring data segment, and the sampling frequency of the data points is consistent with that of the target structure monitoring data segment. The data formed by combining the data points is inserted into the target structure monitoring data segment to achieve data repair of the target structure monitoring data segment.
7. The probability distribution repair method based on improved Wasserstein regression as described in claim 6, characterized in that, Based on the probability density function of the target structure monitoring data segment, the physical laws of the structure are used as sampling constraints. Data points are generated by the conditional Monte Carlo sampling method. The sampling constraints include the rate of change of adjacent data points not exceeding the maximum acceleration response threshold of the structure, local energy conservation conditions, and / or the frequency domain power spectral density being within the target bandwidth.
8. An electronic device, characterized in that, include: The memory stores execution instructions; as well as A processor that executes the execution instructions stored in the memory, causing the processor to perform the probability distribution repair method based on improved Wasserstein regression as described in any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the probability distribution repair method based on improved Wasserstein regression as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Structural health monitoring missing data reconstruction method based on WGANGP-Unet
CN116502060A
Power grid measurement data checking method and system based on artificial intelligence algorithm
CN118260541A