An ancient building structure disease modeling and analysis method and system based on multi-source data

By calculating the correlation coefficient of feature sets in the disease analysis of ancient building structures, building a correlation coefficient matrix and setting a threshold, and eliminating the redundant feature set, the problem of data set redundancy in the existing technology is solved, and the accuracy of the prediction model and the scientificity of the strategy are improved.

CN119989205BActive Publication Date: 2025-07-18HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457306.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing technology lacks scientific feature set screening indicators in the analysis of ancient building structure diseases, resulting in redundancy of data sets, affecting the accuracy of model performance and prediction results.

Method used

By calculating the correlation coefficients of each pair of feature sets, a correlation coefficient matrix is constructed, a high correlation threshold is set, and the proportion of occurrences of high correlation coefficients is counted, the occurrence threshold is obtained, redundant and irrelevant feature sets are eliminated, a prediction model is constructed and a disease maintenance strategy is formulated.

Benefits of technology

The quality of the data set is improved, the accuracy of the prediction results of the prediction model and the scientificity and accuracy of the final strategy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989205B_ABST
    Figure CN119989205B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for modeling and analyzing ancient building structure diseases based on multi-source data, which relates to the technical field of ancient building structure disease analysis. The present invention calculates the occurrence frequency threshold for screening the feature set in a data-driven manner to improve the scientificity and rationality of feature set screening. Specifically, by calculating the correlation coefficient of each pair of feature sets and constructing a correlation coefficient matrix based on this, and after determining the standard of high correlation coefficient, the proportion of the occurrence frequency of high correlation coefficients in the correlation coefficient matrix is statistically calculated and averaged to obtain the occurrence frequency threshold. Subsequently, the feature set is selected based on the occurrence frequency threshold, that is, the feature set with a high proportion of the occurrence frequency of high correlation coefficients is retained, which can effectively eliminate redundant and irrelevant feature sets, reduce the redundant information in the dataset, thereby effectively improving the accuracy of the prediction results output by the prediction model, and at the same time effectively ensuring the scientificity and accuracy of the final strategy formulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of ancient building structure disease analysis, and particularly to a method and system for modeling and analyzing ancient building structure diseases based on multi-source data. Background Art

[0002] With the progress of technology, the demand for the protection and restoration of ancient buildings is increasing day by day. How to effectively use modern technical means to accurately evaluate ancient buildings has become a research hotspot. Traditionally, the analysis of ancient building structure diseases mainly relies on the experience judgment of experts and limited historical literature. However, this method has disadvantages such as strong subjectivity and incomplete information. Therefore, the existing technology aims to analyze the ancient building structure diseases through multiple data sources, collect multi-source data in each data source, and then combine with a prediction model. The following are the implementation steps of the existing method:

[0003] First, we collect multivariate data from various data sources, such as historical archives and documents, physical testing, and environmental monitoring. In historical archives and documents, we can not only collect geometric information on ancient building structures, but also material properties such as strength, durability, and mechanical properties such as stress and strain. In physical testing, we use non-destructive testing technologies such as ultrasonic flaw detection and infrared thermal imaging to detect internal defects in ancient building materials and record data in multiple dimensions, such as crack width, depth, expansion speed, surface roughness, etc. In environmental monitoring, we deploy a sensor network to monitor and collect various surrounding environmental data in real time, such as temperature, humidity, wind speed, and rainfall. Based on the collected multivariate data, data preprocessing is then performed, including cleaning, normalizing and standardizing the collected multivariate data to ensure data consistency and comparability, as well as removing outliers, filling missing data, and converting different types of data into a form suitable for machine learning algorithms. Next, feature extraction is performed based on the multivariate data in each data source after preprocessing, mainly to extract key features that reflect the health status of ancient buildings. For example, in historical archives and documents, multidimensional information such as design parameters, construction records, and previous repair documents are extracted; in physical inspections, the location, width, and depth of cracks are extracted. In the process of environmental monitoring, the multidimensional features such as temperature, humidity, wind speed, rainfall, etc. are extracted, and then the features extracted from each data source are integrated to form a feature set of each data source. Finally, the model training stage is the first step. In this process, the model to be trained is selected, usually a prediction model, and a deep learning algorithm such as a convolutional neural network model is used. Then, the multivariate data collected from each data source and the feature set extracted from each data source are integrated into a data set, and the integrated data set is divided into a training set and a validation set, usually in a certain ratio, such as 80% training set and 20% validation set. The prediction model is trained using a training set, and the performance of the trained prediction model is evaluated using a validation set. Common evaluation indicators include accuracy, precision, recall, F1 score, etc., and the generalization ability and prediction accuracy of the prediction model are improved by combining cross-validation and hyperparameter tuning. Finally, a trained prediction model is obtained, and then the multivariate data to be predicted is input into the trained prediction model to obtain the prediction results, including the prediction of the future health status of the ancient building structure and the development trend of potential diseases. Finally, a scientific and reasonable ancient building structure disease maintenance strategy is formulated based on the prediction results.

[0004] Although the implementation based on the existing technology can effectively solve the problem of analyzing the structural diseases of ancient buildings, there are still the following defects: In the existing technology, there is a lack of specific and scientific screening indicators for the selection of feature sets, that is, the correlation between each feature set is often ignored. Specifically, for each extracted feature set, without corresponding correlation analysis, that is, without scientific screening indicators, it is extremely easy to include feature sets with low correlation or even no correlation in the dataset for model training. This greatly increases the redundant information in the dataset, which obviously easily leads to poor accuracy and precision of the integrated dataset for model training. Thus, on the premise of easily affecting the performance of the model, it is easy to cause a decline in the accuracy of the prediction results finally output by the model, and at the same time, it is easy to affect the accuracy of the strategy finally formulated based on the prediction results.

[0005] Therefore, there is an urgent need for a technical solution for a method and system for modeling and analyzing the structural diseases of ancient buildings based on multi-source data in the existing technology. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a method for modeling and analyzing the structural diseases of ancient buildings based on multi-source data, specifically including the following steps:

[0007] Step S1: Collect multi-source data from at least two data sources, and perform preprocessing;

[0008] Step S2: Extract at least two features from the multi-source data of each preprocessed data source, and integrate them to form a feature set for each data source;

[0009] Step S3: Analyze the correlation between each feature set, and obtain the retained feature set according to the analysis result;

[0010] Step S3a: Calculate the correlation coefficient of each pair of feature sets;

[0011] Step S3a1: Obtain the observed values of each feature in each feature set, and construct a matrix based on the observed values of each feature in each feature set;

[0012] Among them, the expression of the matrix is:

[0013]

[0014] In the formula, X represents the matrix; represents the observed value of the m-th feature in the n-th feature set; n represents the number of feature sets;

[0015] Step S3a2: According to the matrix, calculate the mean value of each feature set, and calculate the standard deviation of each feature set according to the mean value;

[0016] Among them, the calculation formula groups for the mean and standard deviation of each feature set are as follows:

[0017]

[0018]

[0019] In the formula, represents the mean of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observed values of the features; represents the observed value of the mth feature in the nth feature set;

[0020] Step S3a3: Calculate the covariance of each pair of feature sets based on the mean of each feature set;

[0021] Among them, the calculation formula for the covariance of each pair of feature sets is:

[0022]

[0023] In the formula, represents the covariance between the nth feature set and the kth feature set; M represents the number of observed values of the features; represents the observed value of the mth feature in the nth feature set; represents the mean of the nth feature set; represents the observed value of the mth feature in the kth feature set; represents the mean of the kth feature set;

[0024] Step S3a4: Calculate the correlation coefficient of each pair of feature sets based on the standard deviation of each feature set and the covariance of each pair of feature sets;

[0025] Among them, the calculation formula for the correlation coefficient of each pair of feature sets is:

[0026]

[0027] In the formula, represents the correlation coefficient between the nth feature set and the kth feature set; represents the covariance between the nth feature set and the kth feature set; and represent the standard deviation of the nth feature set and the standard deviation of the kth feature set respectively;

[0028] Step S3b: Construct a correlation coefficient matrix based on the correlation coefficient of each pair of feature sets;

[0029] Among them, the expression of the correlation coefficient matrix is:

[0030]

[0031] In the formula, R represents the correlation coefficient matrix; Represents the correlation coefficient between the nth feature set and the kth feature set;

[0032] Step S3c, calculating the occurrence number threshold according to the correlation coefficient matrix;

[0033] Step S3c1, setting a high correlation threshold for all correlation coefficients in the correlation coefficient matrix;

[0034] Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined to be a high correlation coefficient;

[0035] Step S3c3, counting the high correlation coefficients between each feature set and other feature sets and the number of occurrences of all correlation coefficients;

[0036] Step S3c4, according to the high correlation coefficients between each feature set and other feature sets and the occurrence times of all correlation coefficients, calculate the occurrence times ratio of the high correlation coefficients between each feature set and other feature sets;

[0037] Step S3c5, summing up the occurrence ratios of high correlation coefficients between each feature set and other feature sets, and calculating the mean to obtain an occurrence threshold;

[0038] Step S3d, using the occurrence number threshold to judge each feature set to obtain a retained feature set;

[0039] If the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is greater than or equal to the occurrence threshold, the current feature set is retained, and all the retained feature sets are integrated into a feature set to obtain a retained feature set; if the occurrence ratio of the high correlation coefficient between the current feature set and other feature sets is less than the occurrence threshold, the current feature set is eliminated;

[0040] Step S4, construct a prediction model, and integrate the retained feature set and the preprocessed multivariate data in each data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate a maintenance strategy for ancient building structural defects based on the prediction result.

[0041] A system for modeling and analyzing structural defects of ancient buildings based on multivariate data, which implements the above-mentioned method for modeling and analyzing structural defects of ancient buildings based on multivariate data, includes the following modules:

[0042] Data acquisition and preprocessing module: used to collect multivariate data in each data source through at least two data sources and perform preprocessing;

[0043] Feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multivariate data of each preprocessed data source and integrate them to form a feature set for each data source;

[0044] Correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain a retained feature set based on the analysis results;

[0045] Model training and strategy formulation module: connected to the correlation analysis module, used to construct a prediction model, integrate the retained feature set with the multivariate data in each preprocessed data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate an ancient building structure disease maintenance strategy based on the prediction result.

[0046] The embodiments of the present invention have the following technical effects:

[0047] In the present invention, the occurrence frequency threshold for screening feature sets is calculated in a data-driven manner to improve the scientificity and rationality of feature set screening. Specifically, the correlation coefficient of each pair of feature sets is calculated and a correlation coefficient matrix is constructed based on this. After determining the standard for high correlation coefficients, the proportion of the occurrence times of high correlation coefficients in the correlation coefficient matrix is statistically calculated and the average value is obtained as the occurrence frequency threshold. Subsequently, feature sets are selected based on the occurrence frequency threshold, that is, feature sets with a high proportion of the occurrence times of high correlation coefficients are retained. This can effectively eliminate redundant and irrelevant feature sets, reduce redundant information in the data set, and thus effectively improve the quality of the data set while effectively improving the accuracy of the prediction results output by the prediction model. At the same time, it can effectively ensure the scientificity and accuracy of the final strategy formulation. Description of the Drawings

[0048] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1 It is a flowchart of a method for modeling and analyzing ancient building structure diseases based on multivariate data provided by an embodiment of the present invention;

[0050] Figure 2 It is a framework diagram of an ancient building structure disease modeling and analysis system based on multi-source data provided by an embodiment of the present invention. Specific implementation manners

[0051] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0052] Embodiment 1: As Figure 1 shown, the present invention provides an ancient building structure disease modeling and analysis method based on multi-source data, including the following steps:

[0053] Step S1: Collect multi-source data in each data source through at least two data sources and perform preprocessing;

[0054] It should be noted that the data source, that is, the data source, includes but is not limited to various data sources such as historical archives and documents, physical detection, and environmental monitoring. And the multi-source data collected in each data source includes: in the data source of historical archives and documents, not only the geometric information of the ancient building structure can be collected, but also material properties such as strength, durability, and mechanical properties such as stress and strain can be included. In the data source of physical detection, non-destructive testing technologies such as ultrasonic flaw detection and infrared thermal imaging are used to detect internal defects of ancient building materials, and multi-source data in multiple dimensions such as crack width, depth, expansion speed, and surface roughness are recorded. In the data source of environmental monitoring, a sensor network is deployed to monitor and collect various surrounding environmental data such as temperature, humidity, wind speed, and rainfall in real time. For preprocessing, it usually includes cleaning, normalization, and standardization processing to ensure data consistency and comparability, and includes removing outliers, filling in missing data, and converting different types of data into a form suitable for machine learning algorithms, that is, a form suitable for model learning.

[0055] Step S2: Extract at least two features from the multi-source data of each data source after preprocessing and integrate them to form a feature set for each data source;

[0056] It should be noted that the features extracted from each data source are not the same. For example, in historical archives and documents, multi-dimensional information features such as design parameters, construction records, and previous renovation documents are extracted; in physical inspections, multi-dimensional features such as crack location, width, depth, and propagation speed are extracted; and in environmental monitoring, multi-dimensional features such as temperature, humidity, wind speed, and rainfall are extracted. Subsequently, the features extracted from each data source can be integrated to form a feature set for each data source.

[0057] Step S3: Analyze the correlation between each feature set, and obtain the retained feature set based on the analysis results;

[0058] Step S3a: Calculate the correlation coefficient for each pair of feature sets;

[0059] It should be noted that by calculating the correlation coefficient for each pair of feature sets, the relationship between feature sets can be quantified, redundant features can be identified, and the interpretability of the model can be enhanced.

[0060] Step S3a1: Obtain the observed values of each feature in each feature set, and construct a matrix based on the observed values of each feature in each feature set;

[0061] Among them, the expression of the matrix is:

[0062]

[0063] In the formula, X represents the matrix; represents the observed value of the m-th feature in the n-th feature set; n represents the number of feature sets;

[0064] It should be noted that in order to elaborate in detail on how to obtain the observed values of each feature in each feature set and construct a matrix based on these observed values, we will elaborate on the specific steps of data collection, data collation, and matrix construction. First, it is necessary to collect the original data of each multivariate data from multiple data sources, and these data sources include but are not limited to historical archives and documents, physical inspections, and environmental monitoring, etc. The following are the specific data sources and collection methods:

[0065] The multivariate data collected from the data source of historical archives and documents includes geometric information, material properties, and mechanical properties, etc. Among them, the geometric information includes but is not limited to data such as design drawings of ancient buildings, dimensions, and shapes in construction records; the material properties include but are not limited to data such as strength and durability; and the mechanical properties include but are not limited to data such as stress and strain;

[0066] The multivariate data collected from the data source of physical inspections includes multi-dimensional data collected using non-destructive testing techniques, such as crack width, depth, propagation speed, surface roughness, etc.;

[0067] The multivariate data collected from the data sources of environmental monitoring include real-time monitoring data, such as temperature, humidity, wind speed, rainfall and other data;

[0068] After the original data is collected, these data need to be sorted out to ensure that each feature in each feature set has a corresponding observed value. The specific steps are as follows: First, conduct data recording, that is, each feature set corresponds to a data table or file to record the data of each observation. For example, for the feature set in historical archives and documents The observed values of each feature in it can be recorded as Table 1 below:

[0069] Table 1 is the observation value record table of the feature set in historical archives and documents for each feature in it

[0070]

[0071] As above, similarly, for the physical detection feature set and the environmental monitoring feature set , the corresponding observed values can also be recorded;

[0072] Based on the above, after obtaining the observed values of each feature in each feature set in the form of a record table, the observed values of each feature in each feature set should be normalized and standardized to ensure the consistency and comparability of the data;

[0073] Subsequently, after completing the data sorting, the observed values of each feature in each feature set can be constructed into a matrix. For the specific detailed and complete matrix expression, please refer to the above, and it will not be elaborated here one by one.

[0074] Step S3a2: According to the matrix, calculate the mean value of each feature set, and calculate the standard deviation of each feature set based on the mean value;

[0075] Among them, the calculation formula group for the mean value and standard deviation of each feature set is:

[0076]

[0077]

[0078] In the formula, represents the mean value of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observed values of the feature; represents the observed value of the mth feature in the nth feature set;

[0079] Step S3a3: Calculate the covariance of each pair of feature sets based on the mean value of each feature set;

[0080] Among them, the calculation formula for the covariance of each pair of feature sets is as follows:

[0081]

[0082] In the formula, represents the covariance between the nth feature set and the kth feature set; M represents the number of observed values of the features; represents the observed value of the mth feature in the nth feature set; represents the mean value of the nth feature set; represents the observed value of the mth feature in the kth feature set; represents the mean value of the kth feature set;

[0083] Step S3a4: Calculate the correlation coefficient of each pair of feature sets based on the standard deviation of each feature set and the covariance of each pair of feature sets;

[0084] Among them, the calculation formula for the correlation coefficient of each pair of feature sets is as follows:

[0085]

[0086] In the formula, represents the correlation coefficient between the nth feature set and the kth feature set; represents the covariance between the nth feature set and the kth feature set; and represent the standard deviation of the nth feature set and the standard deviation of the kth feature set respectively;

[0087] Step S3b: Construct a correlation coefficient matrix based on the correlation coefficients of each pair of feature sets;

[0088] It should be noted that by constructing a correlation coefficient matrix, the relationship between feature sets can be systematically displayed, facilitating subsequent statistics and analysis and improving data processing efficiency.

[0089] Among them, the expression of the correlation coefficient matrix is as follows:

[0090]

[0091] In the formula, R represents the correlation coefficient matrix; represents the correlation coefficient between the nth feature set and the kth feature set;

[0092] Step S3c: Calculate the occurrence threshold based on the correlation coefficient matrix;

[0093] It should be noted that by calculating the occurrence threshold, a scientific screening criterion can be set to ensure that the retained feature sets have high independence and representativeness.

[0094] Step S3c1: Set a high correlation threshold for all correlation coefficients in the correlation coefficient matrix;

[0095] It should be noted that the high correlation threshold is usually defined by the absolute value of the correlation coefficient being greater than 0.8 or 0.9, which is used to judge the high correlation between feature sets.

[0096] Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined as a high correlation coefficient;

[0097] Step S3c3: Count the occurrence times of high correlation coefficients and all correlation coefficients between each feature set and other feature sets;

[0098] Step S3c4: Calculate the proportion of the occurrence times of high correlation coefficients between each feature set and other feature sets based on the occurrence times of high correlation coefficients and all correlation coefficients between each feature set and other feature sets;

[0099] Step S3c5: Sum up the proportions of the occurrence times of high correlation coefficients between each feature set and other feature sets, and calculate the mean value to obtain the occurrence times threshold;

[0100] Step S3d: Use the occurrence times threshold to judge each feature set to obtain the retained feature set;

[0101] If the proportion of the occurrence times of high correlation coefficients between the current feature set and other feature sets is greater than or equal to the occurrence times threshold, retain the current feature set, and integrate all the retained feature sets into a feature set to obtain the retained feature set; if the proportion of the occurrence times of high correlation coefficients between the current feature set and other feature sets is less than the occurrence times threshold, eliminate the current feature set;

[0102] It should be noted that by using the occurrence times threshold to judge feature sets, redundant and irrelevant feature sets can be effectively eliminated, the quality of the data set can be improved, and the performance of the model and the accuracy of the prediction results can be enhanced.

[0103] Step S4: Construct a prediction model, integrate the retained feature set and the multivariate data in each preprocessed data source into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate an ancient building structure disease maintenance strategy based on the prediction result;

[0104] It should be noted that the training process of the prediction model usually includes data set preparation, model selection, model training and evaluation, so as to obtain a trained prediction model. The training process of the prediction model is an existing conventional technical means and a technology that is currently proficiently mastered, and will not be elaborated one by one here;

[0105] It is further worth noting that when predicting the future health status or the development trend of potential diseases of the ancient building structure, regarding how to formulate a scientific and reasonable maintenance strategy for the diseases of the ancient building structure, the prediction results should first be classified and explained. Generally, the prediction results can be divided into the following categories: good health status, that is, the ancient building structure is currently in a good state without obvious diseases; minor diseases, that is, there are some minor problems in the ancient building structure, but these problems have not had a significant impact on the overall structure; moderate diseases, that is, there are obvious diseases in the ancient building structure, and timely treatment is required to prevent further deterioration; serious diseases, that is, there are serious diseases in the ancient building structure, which may have threatened the safety of the structure, and immediate measures need to be taken.

[0106] Based on the different categories of the above prediction results, corresponding maintenance strategies for the diseases of the ancient building structure can be formulated. For example, the maintenance strategies for the good health status include regular inspections. That is, although the structure is currently in a good state, regular inspections are still required to ensure that no new problems occur. It is recommended to conduct a comprehensive inspection every six months to one year; preventive maintenance, that is, implementing some preventive measures, such as cleaning the drainage system, repairing small-scale surface damages, etc., to prevent possible future problems; environmental monitoring, that is, continuously monitoring the changes in the surrounding environment such as temperature, humidity, wind speed, rainfall, etc., in order to timely detect factors that may cause structural changes.

[0107] Example 2: As Figure 2 shown, the present invention also proposes an analysis system for modeling diseases of ancient building structures based on multi-source data, which executes an analysis method for modeling diseases of ancient building structures based on multi-source data as described above, and includes the following modules:

[0108] Data acquisition and preprocessing module: used to collect multi-source data in each data source through at least two data sources and perform preprocessing;

[0109] Feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multi-source data of each preprocessed data source and integrate them to form a feature set for each data source;

[0110] Correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain a retained feature set based on the analysis results;

[0111] Model training and strategy formulation module: Connected to the relevance analysis module, it is used to construct a prediction model, integrate the retained feature set and the multivariate data in each preprocessed data source into a data set, train the prediction model using the data set to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate an ancient building structure disease maintenance strategy based on the prediction result.

[0112] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" do not specifically refer to the singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method or device including the said element.

[0113] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0114] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of the present invention and do not limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.

Claims

1. A method for modeling and analyzing ancient building structure diseases based on multi-source data, characterized in that, Including the following steps: Step S1: Collect multi-source data from at least two data sources, and perform preprocessing; Step S2: Extract at least two features from the multi-source data of each preprocessed data source, and integrate them to form a feature set for each data source; Step S3: Analyze the correlation between each feature set, and obtain the retained feature set according to the analysis result; Step S3a: Calculate the correlation coefficient of each pair of feature sets; Step S3a1: Obtain the observed values of each feature in each feature set, and construct a matrix based on the observed values of each feature in each feature set; Among them, the expression of the matrix is: Wherein, X represents a matrix; represents the observed value of the m-th feature in the n-th feature set; n represents the number of feature sets; Step S3a2: According to the matrix, calculate the mean of each feature set, and calculate the standard deviation of each feature set according to the mean; Among them, the calculation formula group of the mean and standard deviation of each feature set is: wherein, represents the mean of the nth feature set; represents the standard deviation of the nth feature set; M represents the number of observed values of the features; Step S3a3: Calculate the covariance of each pair of feature sets according to the mean of each feature set; Among them, the calculation formula of the covariance of each pair of feature sets is: Wherein, represents the covariance between the nth feature set and the kth feature set; represents the mean of the nth feature set; represents the observed value of the mth feature in the kth feature set; represents the mean of the kth feature set; Step S3a4: Calculate the correlation coefficient of each pair of feature sets according to the standard deviation of each feature set and the covariance of each pair of feature sets; Among them, the calculation formula of the correlation coefficient of each pair of feature sets is: In the formula, represents the correlation coefficient between the n-th feature set and the k-th feature set; represents the covariance between the n-th feature set and the k-th feature set; and represent the standard deviation of the n-th feature set and the standard deviation of the k-th feature set, respectively; Step S3b: Construct a correlation coefficient matrix based on the correlation coefficient of each pair of feature sets; Step S3c: Calculate the occurrence times threshold according to the correlation coefficient matrix; Step S3d: Use the occurrence times threshold to judge each feature set, and obtain the retained feature set; Step S4: Construct a prediction model, integrate the retained feature set and the multi-source data in each preprocessed data source into a data set, use the data set to train the prediction model, obtain the trained prediction model, input the multi-source data to be predicted into the trained prediction model, obtain the prediction result, and formulate a maintenance strategy for the ancient building structure diseases based on the prediction result.

2. The method for modeling and analyzing ancient building structure diseases based on multi-source data according to claim 1, wherein The expression of the said correlation coefficient matrix is: Wherein, R represents a correlation coefficient matrix; represents the correlation coefficient between the nth feature set and the kth feature set.

3. The method for modeling and analyzing ancient building structure diseases based on multi-source data according to claim 2, wherein, The calculating the occurrence times threshold according to the correlation coefficient matrix includes: Step S3c1: Set a high correlation threshold for all correlation coefficients in the correlation coefficient matrix; Step S3c2: If the absolute value of the current correlation coefficient is greater than or equal to the high correlation threshold, it is determined as a high correlation coefficient; Step S3c3: Count the occurrence times of the high correlation coefficients and all correlation coefficients between each feature set and other feature sets; Step S3c4: Calculate the proportion of the occurrence times of the high correlation coefficients between each feature set and other feature sets according to the occurrence times of the high correlation coefficients and all correlation coefficients between each feature set and other feature sets; Step S3c5: Sum up the proportion of the occurrence times of the high correlation coefficients between each feature set and other feature sets, and calculate the mean value to obtain the occurrence times threshold.

4. A method for modeling and analyzing ancient building structure diseases based on multi-source data according to claim 3, characterized in that, The using the occurrence times threshold to judge each feature set and obtain the retained feature set includes: If the proportion of the occurrence times of the high correlation coefficients between the current feature set and other feature sets is greater than or equal to the occurrence times threshold, the current feature set is retained, and all the retained feature sets are integrated into a feature set to obtain the retained feature set; if the proportion of the occurrence times of the high correlation coefficients between the current feature set and other feature sets is less than the occurrence times threshold, the current feature set is removed.

5. A system for modeling and analyzing ancient building structure diseases based on multi-source data, which is applied to a method for modeling and analyzing ancient building structure diseases based on multi-source data according to any one of claims 1-4, and is characterized in that, The system includes: A data acquisition and preprocessing module: used to acquire multivariate data in each data source through at least two data sources and perform preprocessing; A feature set acquisition module: connected to the data acquisition and preprocessing module, used to extract at least two features from the multivariate data of each data source after preprocessing and integrate them to form a feature set for each data source; A correlation analysis module: connected to the feature set acquisition module, used to analyze the correlation between each feature set and obtain the retained feature set according to the analysis result; A model training and strategy formulation module: connected to the correlation analysis module, used to construct a prediction model, integrate the retained feature set and the multivariate data in each data source after preprocessing into a data set, use the data set to train the prediction model to obtain a trained prediction model, input the multivariate data to be predicted into the trained prediction model to obtain a prediction result, and formulate an ancient building structure disease maintenance strategy based on the prediction result.

Citation Information

Patent Citations

  • Historic building health monitoring system, method and equipment

    CN115146230A

  • Ancient building risk prediction management and control method and system based on large model

    CN119624136A