Historic building structure disease modeling analysis method and system based on multivariate data
By calculating the correlation coefficients between feature sets and building a correlation coefficient matrix, the high correlation feature sets are screened out, which solves the problem of lack of scientificity in feature set screening in the existing technology, and improves the model performance and accuracy of prediction results of ancient building structure disease analysis.
Patent Information
- Application Number
- CN202510457306.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing technology lacks scientific feature set screening indicators in the analysis of ancient building structure diseases, resulting in the neglect of correlation between feature sets, making it easy to introduce redundant information, and reduces the accuracy of model performance and prediction results.
By calculating the correlation coefficients of each pair of feature sets, a correlation coefficient matrix is constructed, and the proportion of occurrences is calculated based on the high correlation threshold, the threshold of occurrences is calculated. Based on this, the feature set is filtered, the high correlation feature set is retained, and the redundant feature set is eliminated.
Effectively eliminate redundant and irrelevant feature sets, reduce redundant information of the data set, improve the quality of the data set, improve the performance of the prediction model and the accuracy of the prediction results, and ensure the scientificity and accuracy of the final strategy.
Smart Images

Figure CN119989205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ancient building structural disease analysis, and in particular to an ancient building structural disease modeling analysis method and system based on multivariate data. Background Art
[0002] With the advancement of science and technology, the demand for the protection and restoration of ancient buildings is growing. How to effectively use modern technical means to accurately evaluate ancient buildings has become a research hotspot. Traditionally, the analysis of structural defects of ancient buildings mainly relies on the experience and judgment of experts and limited historical documents. However, this method has the disadvantages of strong subjectivity and incomplete information. Therefore, the existing technology aims to analyze the structural defects of ancient buildings by collecting multivariate data from multiple data sources and combining them with prediction models. The following are the implementation steps of the existing method: First, we collect multivariate data from various data sources, such as historical archives and documents, physical testing, and environmental monitoring. In historical archives and documents, we can not only collect geometric information on ancient building structures, but also material properties such as strength, durability, and mechanical properties such as stress and strain. In physical testing, we use non-destructive testing technologies such as ultrasonic flaw detection and infrared thermal imaging to detect internal defects in ancient building materials and record data in multiple dimensions, such as crack width, depth, expansion speed, surface roughness, etc. In environmental monitoring, we deploy a sensor network to monitor and collect various surrounding environmental data in real time, such as temperature, humidity, wind speed, and rainfall. Based on the collected multivariate data, data preprocessing is then performed, including cleaning, normalizing and standardizing the collected multivariate data to ensure data consistency and comparability, as well as removing outliers, filling missing data, and converting different types of data into a form suitable for machine learning algorithms. Next, feature extraction is performed based on the multivariate data in each data source after preprocessing, mainly to extract key features that reflect the health status of ancient buildings. For example, in historical archives and documents, multidimensional information such as design parameters, construction records, and previous repair documents are extracted; in physical inspections, the location, width, and depth of cracks are extracted. In the process of environmental monitoring, the multidimensional features such as temperature, humidity, wind speed, rainfall, etc. are extracted, and then the features extracted from each data source are integrated to form a feature set of each data source. Finally, the model training stage is the first step. In this process, the model to be trained is selected, usually a prediction model, and a deep learning algorithm such as a convolutional neural network model is used. Then, the multivariate data collected from each data source and the feature set extracted from each data source are integrated into a data set, and the integrated data set is divided into a training set and a validation set, usually in a certain ratio, such as 80% training set and 20% validation set. The prediction model is trained using a training set, and the performance of the trained prediction model is evaluated using a validation set. Common evaluation indicators include accuracy, precision, recall, F1 score, etc., and the generalization ability and prediction accuracy of the prediction model are improved by combining cross-validation and hyperparameter tuning. Finally, a trained prediction model is obtained, and then the multivariate data to be predicted is input into the trained prediction model to obtain the prediction results, including the prediction of the future health status of the ancient building structure and the development trend of potential diseases. Finally, a scientific and reasonable ancient building structure disease maintenance strategy is formulated based on the prediction results. Although the implementation based on the existing technology can effectively solve the problem of structural defect analysis of existing ancient buildings, the following defects still exist: the existing technology lacks specific scientific screening indicators for the screening of feature sets, that is, the correlation between each feature set is often ignored. Specifically, for each extracted feature set, in the absence of corresponding correlation analysis, that is, the lack of scientific screening indicators, it is very easy to cause low-correlation or even irrelevant feature sets to be included in the data set for model training, which greatly increases the redundant information of the data set, which is obviously easy to lead to poor accuracy and precision of the integrated data set for model training, thereby easily affecting the performance of the model, and easily leading to a decrease in the accuracy of the prediction results ultimately output by the model, and at the same time, it is easy to affect the accuracy of the strategy ultimately formulated based on the prediction results.
[0003] Therefore, the prior art urgently needs a method and system technical solution for modeling and analyzing ancient building structural defects based on multivariate data. Summary of the invention
[0004] In order to solve the above technical problems, the present invention provides a method for modeling and analyzing ancient building structural defects based on multivariate data, which specifically includes the following steps: Step S1, collecting multivariate data from each data source through at least two data sources, and performing preprocessing; Step S2, extracting at least two features from the preprocessed multivariate data of each data source, and integrating them to form a feature set of each data source; Step S3, analyzing the correlation between each feature set, and obtaining a retained feature set based on the analysis result; Step S3a, calculating the correlation coefficient of each pair of feature sets; Step S3a1, obtaining the observed value of each feature in each feature set, and constructing a matrix based on the observed value of each feature in each feature set; The matrix expression is:
[0005] In the formula, X represents the matrix; represents the observed value of the mth feature in the nth feature set; n represents the number of feature sets; Step S3a2: Calculate the mean of each feature set based on the matrix, and calculate the standard deviation of each feature set based on the mean; Among them, the calculation formula group of the mean and standard deviation of each feature set is:
[0006]
[0007] In the formula, Represents the mean of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Step S3a3: Calculate the covariance of each pair of feature sets based on the mean of each feature set; Among them, the calculation formula for the covariance of each pair of feature sets is:
[0008] In the formula, represents the covariance between the nth feature set and the kth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Represents the mean of the nth feature set; Represents the observed value of the mth feature in the kth feature set; represents the mean of the kth feature set; Step S3a4: Calculate the correlation coefficient of each pair of feature sets based on the standard deviation of each feature set and the covariance of each pair of feature sets; The calculation formula for the correlation coefficient of each pair of feature sets is:
[0009] In the formula, Represents the correlation coefficient between the nth feature set and the kth feature set; Represents the covariance between the nth feature set and the kth feature set; and Represent the standard deviation of the nth feature set and the standard deviation of the kth feature set respectively; Step S3b, constructing a correlation coefficient matrix based on the correlation coefficients of each pair of feature sets; Among them, the expression of the correlation coefficient matrix is:
[0010] In the formula, R represents the correlation coefficient matrix; Represents the correlation coefficient between the nth feature set and the kth feature set; Step S3c, calculating the occurrence number threshold according to the correlation coefficient matrix; Step S3c1, setting a high correlation threshold for all correlation coefficients in the correlation coefficient matrix; Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined to be a high correlation coefficient; Step S3c3, counting the high correlation coefficients between each feature set and other feature sets and the number of occurrences of all correlation coefficients; Step S3c4, according to the high correlation coefficients between each feature set and other feature sets and the occurrence times of all correlation coefficients, calculate the occurrence times ratio of the high correlation coefficients between each feature set and other feature sets; Step S3c5, summing up the occurrence ratios of high correlation coefficients between each feature set and other feature sets, and calculating the mean to obtain an occurrence threshold; Step S3d, using the occurrence number threshold to judge each feature set to obtain a retained feature set; If the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is greater than or equal to the occurrence threshold, the current feature set is retained, and all the retained feature sets are integrated into a feature set to obtain a retained feature set; if the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is less than the occurrence threshold, the current feature set is eliminated; Step S4, construct a prediction model, and integrate the retained feature set and the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate a maintenance strategy for ancient building structural defects based on the prediction result.
[0011] A system for modeling and analyzing structural defects of ancient buildings based on multivariate data, which implements the above-mentioned method for modeling and analyzing structural defects of ancient buildings based on multivariate data, includes the following modules: Data collection and preprocessing module: used to collect multivariate data from each data source through at least two data sources and perform preprocessing; Feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multi-data of each data source after preprocessing, and integrate them to form a feature set of each data source; Correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain the retained feature set based on the analysis results; Model training and strategy formulation module: connected to the correlation analysis module, used to build a prediction model, and integrate the retained feature set with the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain the prediction results, and formulate the ancient building structure disease maintenance strategy based on the prediction results.
[0012] The embodiments of the present invention have the following technical effects: The present invention calculates the occurrence threshold for screening feature sets in a data-driven manner to improve the scientificity and rationality of feature set screening. Specifically, by calculating the correlation coefficient of each pair of feature sets and constructing a correlation coefficient matrix based on this, and after determining the standard of the high correlation coefficient, the occurrence ratio of the high correlation coefficient in the correlation coefficient matrix is counted and calculated, and the average is calculated to obtain the occurrence threshold, and then the feature set is selected based on the occurrence threshold, that is, the feature set with a high occurrence ratio of the high correlation coefficient is retained. In this way, redundant and irrelevant feature sets can be effectively eliminated, and redundant information in the data set can be reduced. Therefore, on the premise of being able to effectively improve the quality of the data set, the accuracy of the prediction results output by the prediction model is effectively improved, and at the same time, the scientificity and accuracy of the final strategy formulation can be effectively ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0014] Figure 1 It is a flow chart of a method for modeling and analyzing ancient building structural defects based on multivariate data provided by an embodiment of the present invention; Figure 2 It is a framework diagram of an ancient building structural disease modeling and analysis system based on multivariate data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.
[0016] Embodiment 1: Figure 1 As shown, the present invention provides a method for modeling and analyzing ancient building structural defects based on multivariate data, comprising the following steps: Step S1, collecting multivariate data from each data source through at least two data sources, and performing preprocessing; It is worth noting that the data sources include but are not limited to historical archives and documents, physical testing, environmental monitoring and other data sources, and the multivariate data collected in each data source include: in the data source of historical archives and documents, not only the geometric information of the ancient building structure can be collected, but also material properties such as strength, durability, mechanical properties such as stress, strain and other multivariate data; in the data source of physical testing, non-destructive testing technologies such as ultrasonic flaw detection and infrared thermal imaging are used to detect the internal defects of ancient building materials, and record multivariate data in multiple dimensions such as crack width, depth, expansion speed, surface roughness, etc.; in the data source of environmental monitoring, sensor networks are deployed to monitor and collect various surrounding environmental data in real time, such as temperature, humidity, wind speed, rainfall and other multivariate data; preprocessing usually includes cleaning, normalization and standardization to ensure the consistency and comparability of the data, as well as removing outliers, filling in missing data, and converting different types of data into a form suitable for machine learning algorithms, that is, a form suitable for model learning.
[0017] Step S2, extracting at least two features from the preprocessed multivariate data of each data source, and integrating them to form a feature set of each data source; It is worth noting that the features extracted from each data source are not consistent. For example, in historical archives and documents, multidimensional information features such as design parameters, construction records, and previous repair documents are extracted; in physical inspections, multidimensional features such as crack location, width, depth, and expansion speed are extracted; and in environmental monitoring, multidimensional features such as temperature, humidity, wind speed, and rainfall are extracted. The features extracted from each data source can then be integrated to form a feature set for each data source.
[0018] Step S3, analyzing the correlation between each feature set, and obtaining a retained feature set based on the analysis result; Step S3a, calculating the correlation coefficient of each pair of feature sets; It is worth noting that by calculating the correlation coefficient of each pair of feature sets, the relationship between feature sets can be quantified, redundant features can be identified, and the interpretability of the model can be enhanced.
[0019] Step S3a1, obtaining the observed value of each feature in each feature set, and constructing a matrix based on the observed value of each feature in each feature set; The matrix expression is:
[0020] In the formula, X represents the matrix; represents the observed value of the mth feature in the nth feature set; n represents the number of feature sets; It is worth noting that in order to explain in detail how to obtain the observed values of each feature in each feature set and construct a matrix based on these observed values, we will explain the specific steps of data collection, data organization and matrix construction. First, the original data of each multivariate data needs to be collected from multiple data sources, including but not limited to historical archives and documents, physical testing and environmental monitoring, etc. The following are the specific data sources and collection methods: The multivariate data collected from the data sources of historical archives and documents include geometric information, material properties and mechanical properties, etc., where the geometric information includes but is not limited to the data such as size and shape in the design drawings and construction records of ancient buildings, the material properties include but are not limited to data such as strength and durability, and the mechanical properties include but are not limited to data such as stress and strain; The multidimensional data collected in the data source of physical testing include multidimensional data collected by nondestructive testing technology, such as crack width, depth, propagation speed, surface roughness, etc.; The multivariate data collected in the data source of environmental monitoring includes real-time monitoring data, such as temperature, humidity, wind speed, rainfall, etc.; After collecting the raw data, it is necessary to organize the data to ensure that each feature in each feature set has a corresponding observation value. The specific steps are as follows: First, record the data, that is, each feature set corresponds to a data table or file to record the data of each observation. For example, for feature sets in historical archives and documents The observed values of each feature in can be recorded as follows in Table 1: Table 1 shows the feature set in historical archives and documents A table of observation values for each feature in
[0021] As above, similarly, for the physical detection feature set and environmental monitoring feature sets , and the corresponding observations can also be recorded; Based on the above, after obtaining the observed values of each feature in each feature set in the form of a record table, the observed values of each feature in each feature set should be normalized and standardized to ensure the consistency and comparability of the data; After completing the data sorting, the observation values of each feature in each feature set can be constructed into a matrix. Please refer to the above for the detailed and complete matrix expression, which will not be repeated here.
[0022] Step S3a2: Calculate the mean of each feature set based on the matrix, and calculate the standard deviation of each feature set based on the mean; Among them, the calculation formula group of the mean and standard deviation of each feature set is:
[0023]
[0024] In the formula, Represents the mean of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Step S3a3: Calculate the covariance of each pair of feature sets based on the mean of each feature set; Among them, the calculation formula for the covariance of each pair of feature sets is:
[0025] In the formula, represents the covariance between the nth feature set and the kth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Represents the mean of the nth feature set; Represents the observed value of the mth feature in the kth feature set; represents the mean of the kth feature set; Step S3a4: Calculate the correlation coefficient of each pair of feature sets based on the standard deviation of each feature set and the covariance of each pair of feature sets; The calculation formula for the correlation coefficient of each pair of feature sets is:
[0026] In the formula, Represents the correlation coefficient between the nth feature set and the kth feature set; Represents the covariance between the nth feature set and the kth feature set; and Represent the standard deviation of the nth feature set and the standard deviation of the kth feature set respectively; Step S3b, constructing a correlation coefficient matrix based on the correlation coefficients of each pair of feature sets; It is worth noting that by constructing a correlation coefficient matrix, the relationship between feature sets can be systematically displayed, which is convenient for subsequent statistics and analysis and improves data processing efficiency.
[0027] Among them, the expression of the correlation coefficient matrix is:
[0028] In the formula, R represents the correlation coefficient matrix; Represents the correlation coefficient between the nth feature set and the kth feature set; Step S3c, calculating the occurrence number threshold according to the correlation coefficient matrix; It is worth noting that by calculating the occurrence threshold, scientific screening criteria can be set to ensure that the retained feature set has high independence and representativeness.
[0029] Step S3c1, setting a high correlation threshold for all correlation coefficients in the correlation coefficient matrix; It is worth noting that the high correlation threshold is usually defined as the absolute value of the correlation coefficient being greater than 0.8 or 0.9, so as to judge the high correlation between feature sets.
[0030] Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined to be a high correlation coefficient; Step S3c3, counting the high correlation coefficients between each feature set and other feature sets and the number of occurrences of all correlation coefficients; Step S3c4, according to the high correlation coefficients between each feature set and other feature sets and the occurrence times of all correlation coefficients, calculate the occurrence times ratio of the high correlation coefficients between each feature set and other feature sets; Step S3c5, summing up the occurrence ratios of high correlation coefficients between each feature set and other feature sets, and calculating the mean to obtain an occurrence threshold; Step S3d, using the occurrence number threshold to judge each feature set to obtain a retained feature set; If the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is greater than or equal to the occurrence threshold, the current feature set is retained, and all the retained feature sets are integrated into a feature set to obtain a retained feature set; if the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is less than the occurrence threshold, the current feature set is eliminated; It is worth noting that by using the occurrence threshold to determine the feature set, redundant and irrelevant feature sets can be effectively eliminated, the quality of the data set can be improved, and the model performance and accuracy of the prediction results can be improved.
[0031] Step S4, constructing a prediction model, and integrating the retained feature set and the preprocessed multivariate data in each data source into a data set, using the data set to train the prediction model to obtain a trained prediction model, inputting the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulating a maintenance strategy for ancient building structural defects based on the prediction result; It is worth noting that the training process of the prediction model generally includes data set preparation, model selection, model training and evaluation, so as to obtain a trained prediction model. The training process of the prediction model is an existing conventional technical means and is a technology that is well-known and will not be described in detail here. It is worth further explaining that when the prediction result is the future health status of the ancient building structure or the development trend of potential diseases, how to formulate a scientific and reasonable ancient building structure disease maintenance strategy, the prediction results should be classified and interpreted first. Usually, the prediction results can be divided into the following categories: good health status, that is, the ancient building structure is currently in good condition and has no obvious diseases; minor diseases, that is, there are some minor problems with the ancient building structure, but these problems have not yet caused a major impact on the overall structure; moderate diseases, that is, the ancient building structure has obvious diseases, which need to be dealt with in time to prevent further deterioration; serious diseases, that is, the ancient building structure has serious diseases, which may have threatened the safety of the structure and need to be taken immediately; Based on the different categories of the above prediction results, corresponding maintenance strategies for ancient building structural defects can be formulated. For example, the maintenance strategy for when the structure is in good health includes regular inspections, that is, although the structure is currently in good condition, it still needs to be inspected regularly to ensure that no new problems arise. It is recommended to conduct a comprehensive inspection every 6 months to 1 year; preventive maintenance, that is, implement some preventive measures, such as cleaning the drainage system, repairing small-scale surface damage, etc., to prevent possible problems in the future; environmental monitoring, that is, continuously monitor changes in the surrounding environment such as temperature, humidity, wind speed, rainfall, etc., in order to promptly discover factors that may cause structural changes.
[0032] Embodiment 2: Figure 2 As shown, the present invention also proposes a modeling and analysis system for ancient building structure defects based on multivariate data, which executes the above-mentioned modeling and analysis method for ancient building structure defects based on multivariate data, including the following modules: Data collection and preprocessing module: used to collect multivariate data from each data source through at least two data sources and perform preprocessing; Feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multi-data of each data source after preprocessing, and integrate them to form a feature set of each data source; Correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain the retained feature set based on the analysis results; Model training and strategy formulation module: connected to the correlation analysis module, used to build a prediction model, and integrate the retained feature set with the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain the prediction results, and formulate the ancient building structure disease maintenance strategy based on the prediction results.
[0033] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.
[0034] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0035] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A method for modeling and analyzing ancient building structural defects based on multivariate data, characterized in that: The following steps are involved: Step S1, collecting multivariate data from each data source through at least two data sources, and performing preprocessing; Step S2, extracting at least two features from the preprocessed multivariate data of each data source, and integrating them to form a feature set of each data source; Step S3, analyzing the correlation between each feature set, and obtaining a retained feature set based on the analysis result; Step S4, construct a prediction model, and integrate the retained feature set and the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate a maintenance strategy for ancient building structural defects based on the prediction result.
2. The ancient building structural disease modeling and analysis method based on multivariate data according to claim 1 is characterized in that: The analysis of the correlation between each feature set and obtaining the retained feature set according to the analysis result include: Step S3a, calculating the correlation coefficient of each pair of feature sets; Step S3b, constructing a correlation coefficient matrix based on the correlation coefficients of each pair of feature sets; Step S3c, calculating the occurrence number threshold according to the correlation coefficient matrix; Step S3d: Use the occurrence count threshold to judge each feature set to obtain a retained feature set.
3. The ancient building structural disease modeling and analysis method based on multivariate data according to claim 2 is characterized in that: The calculation of the correlation coefficient of each pair of feature sets includes: Step S3a1, obtaining the observed value of each feature in each feature set, and constructing a matrix based on the observed value of each feature in each feature set; The matrix expression is: In the formula, X represents the matrix; represents the observed value of the mth feature in the nth feature set; n represents the number of feature sets; Step S3a2: Calculate the mean of each feature set based on the matrix, and calculate the standard deviation of each feature set based on the mean; Among them, the calculation formula group of the mean and standard deviation of each feature set is: In the formula, Represents the mean of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Step S3a3: Calculate the covariance of each pair of feature sets based on the mean of each feature set; Among them, the calculation formula for the covariance of each pair of feature sets is: In the formula, represents the covariance between the nth feature set and the kth feature set; M represents the number of observations of the feature; Represents the observed value of the mth feature in the nth feature set; Represents the mean of the nth feature set; Represents the observed value of the mth feature in the kth feature set; represents the mean of the kth feature set; Step S3a4: Calculate the correlation coefficient of each pair of feature sets based on the standard deviation of each feature set and the covariance of each pair of feature sets; The calculation formula for the correlation coefficient of each pair of feature sets is: In the formula, Represents the correlation coefficient between the nth feature set and the kth feature set; Represents the covariance between the nth feature set and the kth feature set; and They represent the standard deviation of the nth feature set and the standard deviation of the kth feature set respectively.
4. The ancient building structural disease modeling and analysis method based on multivariate data according to claim 3 is characterized in that: The expression of the correlation coefficient matrix is: In the formula, R represents the correlation coefficient matrix; Represents the correlation coefficient between the nth feature set and the kth feature set.
5. The ancient building structural disease modeling and analysis method based on multivariate data according to claim 4 is characterized in that: The calculation of the occurrence threshold according to the correlation coefficient matrix includes: Step S3c1, setting a high correlation threshold for all correlation coefficients in the correlation coefficient matrix; Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined to be a high correlation coefficient; Step S3c3, counting the high correlation coefficients between each feature set and other feature sets and the number of occurrences of all correlation coefficients; Step S3c4, according to the high correlation coefficients between each feature set and other feature sets and the occurrence times of all correlation coefficients, calculate the occurrence times ratio of the high correlation coefficients between each feature set and other feature sets; Step S3c5: sum up the occurrence ratios of high correlation coefficients between each feature set and other feature sets, and calculate the mean to obtain an occurrence threshold.
6. The ancient building structural disease modeling and analysis method based on multivariate data according to claim 5 is characterized in that: The method of using the occurrence number threshold to judge each feature set to obtain a retained feature set includes: If the proportion of occurrences of high correlation coefficients between the current feature set and other feature sets is greater than or equal to the occurrence threshold, the current feature set is retained, and all retained feature sets are integrated into a feature set to obtain a retained feature set; if the proportion of occurrences of high correlation coefficients between the current feature set and other feature sets is less than the occurrence threshold, the current feature set is eliminated.
7. A system for modeling and analyzing structural defects of ancient buildings based on multivariate data, applied to a method for modeling and analyzing structural defects of ancient buildings based on multivariate data as claimed in any one of claims 1 to 6, characterized in that: The system comprises: Data collection and preprocessing module: used to collect multivariate data from each data source through at least two data sources and perform preprocessing; Feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multi-data of each data source after preprocessing, and integrate them to form a feature set of each data source; Correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain the retained feature set based on the analysis results; Model training and strategy formulation module: connected to the correlation analysis module, used to build a prediction model, and integrate the retained feature set with the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain the prediction results, and formulate the ancient building structure disease maintenance strategy based on the prediction results.
Citation Information
Patent Citations
Historic building health monitoring system, method and equipment
CN115146230A
Intelligent building quality risk prediction method
CN118735319A
Ancient building risk prediction management and control method and system based on large model
CN119624136A
Intelligent identification method and device for historical building diseases
CN119625435A
Feature selection method based on Pearson correlation and colinearity
CN119785880A
Cited By
Cable bridge intelligent inspection method and system based on unmanned aerial vehicle multi-mode perception
CN121187313A