A multi-source heterogeneous data fusion method for cultural relic protection
Patent Information
- Application Number
- CN202611199508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-08
AI Technical Summary
[0005]本发明的目的在于克服上述现有技术的不足,提供一种面向文物保护的多源异构数据融合方法,利用数据替换的方式,将数据筛分过程集成到数据融合过程,以解决现有技术中,对于数据的清洗与评估过程存在将有效数据误删的可能,存在对不同来源数据不加可信区分地等同采信,无法在持续筛除低可信来源所引入的失真数据的同时避免真实异常被误滤的问题
[0017] Compared with existing technologies, the beneficial effect of this invention lies in that it assigns a credibility level to each analytical item according to the data source of its data stream, and distinguishes and accepts the fusion results accordingly. Since the credibility level objectively characterizes the differences in reliability of data sources in terms of acquisition methods, review processes, and traceability, the distorted information carried by low-reliability sources is no longer amplified in the fusion analysis as it is with high-reliability data, but is instead identified as an object requiring further verification. Therefore, this application can effectively avoid the problem in related technologies where distorted information is amplified due to indiscriminate acceptance without credibility distinction.
Smart Images

Figure CN122712401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital preservation technology for cultural relics, and specifically to a method for fusing multi-source heterogeneous data for the preservation of cultural relics. Background Technology
[0002] In cultural relic conservation, data related to the same cultural relic typically comes from multiple sources, including cultural relic survey data, cultural relic record archive data, digital mapping data of the cultural relic itself, damage survey data, data on the location and extent of the cultural relic, vector boundary data of the protection area and construction control zone, conservation planning data, project business data, administrative management data, as well as IoT sensing data, open-source internet data, basic geographic spatiotemporal data, and thematic data such as meteorological, seismic, and natural resource data. This multi-source data is usually associated with the corresponding cultural relic by its unique code and, after fusion, used for comprehensive anomaly analysis of the cultural relic's preservation status, risk factors, monitoring status, and the suitability of appropriate treatment measures.
[0003] For example, Chinese Patent Publication No. CN115391314A discloses a method for heterogeneous multi-source data fusion processing in cultural relic protection, including: constructing a heterogeneous multi-source data fusion processing model for cultural relic protection; using the constructed heterogeneous multi-source data fusion processing model to perform structuring and data standardization processing on the original data to obtain a first dataset; using the constructed heterogeneous multi-source data fusion processing model to perform data cleaning and quality assessment on the first dataset to obtain a second dataset; and using the constructed heterogeneous multi-source data fusion processing model to perform data fusion and data extraction on the second dataset to obtain a final dataset.
[0004] However, its cleaning and evaluation process may inadvertently delete valid data, and may indiscriminately accept data from different sources without any credibility distinction. It is unable to continuously filter out distorted data introduced by low-credibility sources while preventing real anomalies from being mistakenly filtered out. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-source heterogeneous data fusion method for cultural relic protection. By using data replacement, the data screening process is integrated into the data fusion process. This solves the problems in the prior art, such as the possibility of accidentally deleting valid data during the data cleaning and evaluation process, the indiscriminate acceptance of data from different sources without any reliable distinction, and the inability to continuously screen out distorted data introduced by low-reliability sources while avoiding the misfiltering of real anomalies.
[0006] To achieve the above objectives, this invention provides a multi-source heterogeneous data fusion method for cultural relic protection, comprising the following steps:
[0007] The multi-source heterogeneous data streams corresponding to the cultural relics are converted into analysis items for cultural relic anomaly analysis. Each analysis item is marked with the confidence level of its source data stream, which is determined based on the data source. Perform a first fusion analysis on each analysis item to determine if any outliers exist; When an anomaly exists, the decision on whether to execute a replacement verification strategy is based on the credibility level of the source data stream upon which the anomaly is based. The replacement verification strategy includes: selecting, from the source data streams of the analysis items on which the anomaly is based, the data stream that corresponds to the same cultural relic object and the same analysis item type as the low-confidence data stream and has a higher confidence level than the low-confidence data stream, as the replacement data stream; Replace the analysis item corresponding to the low-confidence data stream with the analysis item corresponding to the replacement data stream, perform a second fusion analysis, and determine whether the anomaly item still holds true. The analysis items corresponding to the low-confidence data stream are retained, and the analysis items corresponding to the replacement data stream are added together for a third fusion analysis to determine whether the anomaly is still valid. Based on the results of the second and third fusion analyses, it is determined whether to filter out the low-confidence data streams and whether to retain the anomalies.
[0008] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the step of determining whether to filter out the low-confidence data stream and whether to retain the outliers based on the results of the second fusion analysis and the third fusion analysis includes: If the anomaly still holds in the second fusion analysis, the anomaly is retained, and the low-confidence data stream is retained. If the anomaly does not hold true in the second fusion analysis but still holds true in the third fusion analysis, the low-confidence data stream is filtered out, and the anomaly is also filtered out. If the anomaly is not valid in the second fusion analysis and is not valid in the third fusion analysis, the anomaly is retained and identified as an anomaly to be reviewed, and the low-confidence data stream is not filtered out.
[0009] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the trust level includes a first trust level, a second trust level, and a third trust level that decrease sequentially. The data streams corresponding to on-site measured data, archival data, and officially published data are marked as the first trust level; the data streams corresponding to business management data and administrative management data are marked as the second trust level; and the data streams corresponding to publicly crawled data from the Internet are marked as the third trust level; the preset level is the second trust level.
[0010] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the analysis items include at least one of preservation status items, risk factor items, monitoring status items, and disposal-related items; The original fields of each data stream are converted into analysis items with unified object identifiers, time identifiers, analysis item types, and analysis item values, and the analysis items are bound to the identifiers and trust levels of their source data streams.
[0011] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the first fusion analysis, the second fusion analysis, and the third fusion analysis are fusion analyses with different analysis items but the same judgment method. The fusion analysis includes: Align the analysis items involved in the analysis according to the object identifier and analysis item type to obtain the abnormal characterization data of the cultural relic object; Anomaly analysis indicators are determined based on the anomaly characterization data. When the anomaly analysis indicators meet the preset anomaly judgment conditions, the existence of the anomaly item is determined, and the analysis item on which the anomaly item is generated and its source data stream are recorded.
[0012] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the selected replacement data stream includes: Prioritize selecting the data stream that corresponds to the same cultural relic object, the same analysis item type, has a higher confidence level than the low-confidence data stream, and whose time interval between the time marker and the low-confidence data stream is no greater than a preset time window, as the replacement data stream; When no data stream meets the aforementioned conditions, a data stream that corresponds to the same analysis item type as the low-confidence data stream, has a higher confidence level than the low-confidence data stream, and corresponds to other cultural relics objects belonging to the same category as the cultural relics object is selected as the replacement data stream; When multiple data streams meet the conditions, the one with the earlier time stamp is selected as the replacement data stream, in descending order of trust level and in ascending order of time stamp.
[0013] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the decision on whether to execute a replacement verification strategy is based on the credibility level of the source data streams of the analysis items upon which the anomalies are based, including: If the confidence level of the source data stream of the analysis item on which the anomaly is based is not lower than the preset level, the existence of the anomaly is confirmed and the replacement verification strategy is not executed. If there is a data stream with a confidence level lower than the preset level in the source data stream of the analysis item on which the anomaly is based, the data stream is identified as the low-confidence data stream, and the replacement verification strategy is executed.
[0014] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, when the replacement data stream still does not exist after selection, the abnormal item is identified as an abnormal item to be reviewed, and a supplementary collection request is generated for the low-confidence data stream; the supplementary collection result corresponding to the supplementary collection request is converted into an analysis item and used as the input for the next round of fusion analysis.
[0015] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the number of confirmed anomalies and the number of filtered anomalies generated by the corresponding analysis items of each data stream in several rounds of fusion analysis are statistically analyzed. When the ratio of the number of confirmed anomalies to the number of filtered anomalies is greater than a first threshold, the credibility level of the data stream is increased. When the ratio is less than a second threshold, the trust level of the data stream is reduced; wherein the first threshold is greater than the second threshold.
[0016] As a preferred technical solution for multi-source heterogeneous data fusion methods for cultural relic protection, the method further includes the following steps: generating an anomaly analysis report based on the confirmed anomalies and / or the anomalies to be reviewed, wherein the anomaly analysis report includes the confirmed anomalies and / or the anomalies to be reviewed, the corresponding data streams, and the confidence level of their corresponding data streams.
[0017] Compared with existing technologies, the beneficial effect of this invention lies in that it assigns a credibility level to each analytical item according to the data source of its data stream, and distinguishes and accepts the fusion results accordingly. Since the credibility level objectively characterizes the differences in reliability of data sources in terms of acquisition methods, review processes, and traceability, the distorted information carried by low-reliability sources is no longer amplified in the fusion analysis as it is with high-reliability data, but is instead identified as an object requiring further verification. Therefore, this application can effectively avoid the problem in related technologies where distorted information is amplified due to indiscriminate acceptance without credibility distinction.
[0018] Furthermore, after identifying an anomaly, the decision on whether to execute a replacement verification strategy is made based on the credibility level of the data stream upon which the anomaly is based. If the anomaly is entirely supported by data streams of a credibility level or higher than a preset threshold, its validity is unaffected by low-credibility data and can be confirmed without verification. Only when low-credibility data streams are involved in supporting the anomaly may it become distorted and therefore require verification. Thus, limiting verification calculations to truly necessary anomalies reduces unnecessary verification computations while ensuring reliability.
[0019] Furthermore, during verification, a data stream with the same cultural relic object, the same analysis item type, and a higher confidence level as the low-confidence data stream is selected as the replacement data stream. Since this replacement data stream is a real and more reliable record of the same type existing in the database, the verification can be based on objective existing data, rather than relying on constructed or hypothetical data. This avoids circular reasoning caused by verifying the object to be verified itself or its source data, thus ensuring that the verification conclusion has an objective basis.
[0020] Furthermore, low-confidence data streams and outliers are only filtered out when both conditions are met simultaneously: the second fusion analysis fails and the third fusion analysis still succeeds. The principle is as follows: the failure of the second fusion analysis indicates that the anomaly no longer exists after replacing it with data of a higher confidence level; the success of the third fusion analysis indicates that the low-confidence data still satisfies the anomaly determination condition even with the participation of data of a higher confidence level; only when both conditions are met can it be determined that the anomaly is caused by the distorted value of the low-confidence data. For cases where only one condition is met or neither is met, the data is identified as either a confirmed anomaly or an anomaly pending review, respectively. Therefore, this invention determines whether to filter out data based on whether the results of two fusion analyses simultaneously meet the above conditions. This reduces the possibility of mistakenly filtering out genuine anomalies while filtering out low-confidence distorted data, thus solving the problem of related technologies struggling to filter out distorted data without mistakenly filtering out genuine anomalies.
[0021] In summary, this invention enables the identification of anomalies in cultural relics to be differentiated based on the reliability of the data source. In cultural relic protection scenarios with multiple data sources and varying degrees of reliability, it can screen out low-reliability distorted data, filter out anomalies caused by distorted data, and reduce the situation where genuine anomalies are mistakenly filtered out, thereby improving the reliability of the results of cultural relic anomaly analysis. Attached Figure Description
[0022] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.
[0023] Figure 1 A flowchart illustrating the overall process of a multi-source heterogeneous data fusion method for cultural relic protection provided in this embodiment of the invention; Figure 2 A flowchart for determining anomalies using the first fusion analysis provided in an embodiment of the present invention; Figure 3 A flowchart for the selection of trust level gating and replacement data streams provided in an embodiment of the present invention; Figure 4 The flowcharts for the second fusion analysis, the third fusion analysis, and the anomaly handling provided in the embodiments of the present invention are as follows: Figure 5This is a flowchart for dynamically maintaining the data flow trust level, as provided in an embodiment of the present invention. Detailed Implementation
[0024] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0025] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0026] This application's embodiment is applicable to situations where a comprehensive information management platform for cultural relic protection is used to fuse multi-source heterogeneous data associated with the same cultural relic and identify anomalies. This platform typically aggregates multiple databases, including cultural relic resource databases, cultural relic spatial databases, cultural relic thematic data databases, and cultural relic-related information databases. Its data sources cover cultural relic survey results, cultural relic records and archives, digital mapping results, electronic archives of damage investigations, administrative and project management data, IoT sensing terminal data, and publicly available internet data. It is understandable that the objective reliability of data from these different sources varies significantly: on-site measurements and authoritative archival data are collected in a standardized manner, are traceable, and have high reliability; while publicly available internet data is diverse in source and contains a mix of true and false information, resulting in low reliability. If all data from different sources are accepted equally without distinction during fusion analysis, the distorted information carried by low-reliability data is easily amplified during fusion, leading to the identification of anomalies that do not actually exist, or causing real anomalies to be drowned out by noise. This embodiment aims to filter out low-reliability distorted data while retaining verified real anomalies.
[0027] like Figure 1 As shown, this embodiment provides a multi-source heterogeneous data fusion method for cultural relic protection, including steps S110 to S170.
[0028] Step S110: Convert the multi-source heterogeneous data streams corresponding to the cultural relics into analysis items for cultural relic anomaly analysis. Each analysis item is marked with the confidence level of its source data stream determined based on the data source.
[0029] In this context, an analytical item can be understood as the smallest semantic unit that can be directly used for fusion analysis after the original data stream has been structurally extracted. For example, the transformation process may include: using ETL (Extract-Transform-Load) tools to extract fields related to the preservation status, risk factors, monitoring status, or treatment measures of cultural relics from the original fields of each data stream; converting the extracted fields into analytical items with unified object identifiers, time identifiers, analytical item types, and analytical item values according to their data format; and binding the transformed analytical items to the identifiers and trust levels of their source data streams to establish a correspondence between analytical items and source data streams. The aforementioned object identifier can use a unique identification code of the cultural relic protection unit to ensure that data on the same cultural relic object across databases and sources can be accurately associated.
[0030] In this embodiment, the analysis items can be semantically categorized into at least one of the following: preservation status items, risk factor items, monitoring status items, and disposal-related items. For example, the percentage of decayed area of load-bearing components converted from electronic disease survey archives belongs to the preservation status item; the regional heavy rainfall warning converted from meteorological data belongs to the risk factor item; the circumferential ratio of the vertical displacement of the body converted from IoT sensing displacement sensors belongs to the monitoring status item; and the implementation status of the protection and reinforcement project converted from project management data belongs to the disposal-related item.
[0031] Furthermore, the credibility level of each data stream is determined based on its data source, and may include a first credibility level, a second credibility level, and a third credibility level, which decrease sequentially. For example, data streams corresponding to field-measured data, archival data, and officially published data (such as disease investigation results, digital mapping results of the physical body, and public early warnings from meteorological departments) are marked as first credibility level; data streams corresponding to business management data and administrative management data (such as project management data and administrative approval data) are marked as second credibility level; and data streams corresponding to publicly crawled data from the internet (such as social media sentiment, news reports, and open-source intelligence) are marked as third credibility level. The aforementioned preset level can be the second credibility level, that is, the second credibility level is used as the boundary to distinguish low-credibility data streams.
[0032] The trust level is assigned simultaneously with the conversion of data streams into analytical items because multi-source heterogeneous data exhibits varying degrees of reliability in terms of collection methods, review processes, and traceability: on-site measurements and authoritative archival data undergo standardized collection and review, ensuring traceability and verifiability; publicly crawled data from the internet comes from diverse sources, lacks unified review, and is a mixture of truth and falsehood. After assigning trust levels, subsequent fusion analysis can differentiate and verify data from different sources, eliminating the need for uniform treatment of all data. Distortions carried by low-reliability data can thus be identified and constrained in subsequent processing. The trust level only indicates the reliability of the source and does not directly indicate whether the anomalies generated by it are distorted; therefore, fusion analysis is still required to determine the presence of anomalies.
[0033] Step S120: Perform a first fusion analysis on each analysis item to determine whether there are any abnormal items.
[0034] like Figure 2 As shown, the first fusion analysis, the subsequent second fusion analysis, and the third fusion analysis are fusion analyses with different participating analysis items but the same judgment method. That is, the three fusion analyses use the same analysis mechanism, the only difference being the set of analysis items participating in the analysis. For example, the fusion analysis may include: associating each analysis item participating in the analysis with the same cultural relic object by object identifier, and classifying and aligning them according to the analysis item type to obtain the abnormal characterization data of the cultural relic object; determining abnormal analysis indicators based on the abnormal characterization data; when the abnormal analysis indicators meet the preset abnormal judgment conditions, determining that there is an abnormal item, and recording the analysis item on which the abnormal item is based and its source data stream, so as to determine the credibility level of the abnormal item.
[0035] Specifically, obtaining the anomaly representation data and the anomaly analysis indicators may include: aligning the preservation status, risk factors, monitoring status, and treatment-related items of the same cultural relic object on a time axis based on the unique code of the cultural relic; extracting core entities such as the cultural relic itself, risk type, influencing factors, and prevention and control measures from the aligned analysis items through entity extraction; and then sorting out the logic between entities through relation modeling to form the correlation between the cultural relic itself, risk factors, degree of impact, and prevention and control plan. The correlation constitutes the anomaly representation data. The anomaly analysis indicators are determined by the degree of deviation of the correlation relative to a preset normal state or historical baseline. They may include at least one of the following: preservation status change indicators, risk factor matching indicators, monitoring deviation indicators, and treatment adaptation indicators, which respectively represent the magnitude of change in preservation status relative to the historical baseline, the degree of matching between risk factors and the vulnerability of the cultural relic, the degree of deviation of monitoring volume relative to the normal range, and the degree of adaptation between existing treatment measures and current risks. The preset anomaly determination conditions are formed by embedding rules from cultural relic protection laws and standards, historical treatment cases and expert experience, and may include risk level classification rules and early warning trigger conditions; when the anomaly analysis index hits the preset anomaly determination conditions, an anomaly is determined to exist, and the anomaly may include at least one of disease development items, external risk items, monitoring anomalies and mismatched treatment measures items.
[0036] It should be noted that performing a first fusion analysis on all analysis items before deciding whether to verify them aims to limit subsequent replacement verification to cases where anomalies are indeed found. Understandably, replacement verification requires selecting replacement data streams and performing two additional fusion analyses, incurring corresponding computational costs; indiscriminately performing verification on all artifacts would waste computational resources. Using the first fusion analysis to determine the presence of anomalies as a prerequisite ensures that verification calculations are triggered only when necessary. Furthermore, even if anomalies are identified, whether that anomaly needs verification still depends on the credibility level of the data stream it is based on.
[0037] For example, the aforementioned preset anomaly judgment conditions can be represented as one or more configurable rules. For instance, if the percentage increase in the decayed area of the same load-bearing wooden component exceeds a set proportion in two adjacent survey periods, and a heavy rainfall warning is issued for the area during the same period, then a disease development item is determined to exist. Similarly, if the cumulative change in the body displacement monitoring volume within a set time window exceeds the normal fluctuation range of that monitoring point, then a monitoring anomaly item is determined to exist. The set proportions and normal fluctuation ranges in the above rules are all configurable thresholds and can be specified separately according to the cultural relic category and disease type.
[0038] To facilitate understanding of the above fusion analysis process, taking a wooden ancient building (denoted as cultural relic object A, unique code A) of a national key cultural relic protection unit as an example, the analysis items participating in the first fusion analysis and their sources and confidence levels are as follows: Analysis item a1 (preservation status item), sourced from the electronic archive data stream of disease investigation, first confidence level, value: the proportion of decayed area of load-bearing wooden columns increased from 8% in the previous year to 15%, and the moisture content was 28%; Analysis item a2 (risk factor item), sourced from the official public data stream of the meteorological department, first confidence level, value: An orange alert for continuous heavy rainfall in the region over the next 72 hours has been issued; Analytical item a3 (risk factor item), sourced from publicly crawled internet data streams, is at the third level of credibility, and its value is a rumor circulating on social media that the building's 'beam frame has collapsed'; Analytical item a4 (monitoring status item), sourced from IoT displacement sensor data streams, is at the first level of credibility, and its value is a 2.1mm increase in the circumferential vertical displacement of the main beam; Analytical item a5 (disposal-related item), sourced from project management data streams, is at the second level of credibility, and its value is that the waterproofing and reinforcement project is in the pending approval stage and has not yet been implemented.
[0039] In the first fusion analysis, a1 to a5 are first associated with cultural relic object A by code A and aligned by type; entity extraction yields the ontology (load-bearing wooden components), risk type (continuous heavy rainfall, decay, increased displacement), and prevention and control measures (waterproofing and reinforcement); relational modeling forms a chain of associations: decay of wooden components (preservation status), continuous heavy rainfall (risk factor), accelerated disease development (impact level), and lack of waterproofing and reinforcement (treatment gap), which matches the pre-set association rule for disease development in ancient buildings during heavy rain; the calculated preservation status change index (both decay area and moisture content increase), risk factor matching index (high match between rainfall and decay), and monitoring deviation index (displacement exceeding the normal range) all meet the risk level triggering conditions, thus identifying the abnormal item as a disease development item, and recording the analysis items a1 to a5 and their corresponding source data streams as the basis for it.
[0040] Step S130: When an anomaly exists, determine whether to execute a replacement verification strategy based on the credibility level of the source data stream of the analysis item on which the anomaly is based.
[0041] like Figure 3As shown, considering that the same anomaly is often supported by data streams from multiple sources and with various trust levels, this embodiment first calculates its trust level composition before making a judgment. Specifically, this may include: determining the source data streams of each analysis item on which the anomaly is based, calculating the trust level of each source data stream, and obtaining the trust level composition of the anomaly; when the trust level of each source data stream in the trust level composition is not lower than a preset level, the anomaly is entirely based on trusted data, its existence is directly confirmed, and no replacement verification strategy is executed, thereby saving unnecessary verification computation; when there is a source data stream in the trust level composition with a trust level lower than the preset level, the data stream is determined as a low-trust data stream, and a replacement verification strategy is executed on it. When there are multiple low-trust data streams, the replacement verification strategy can be executed on each low-trust data stream in order of trust level from low to high.
[0042] Understandably, if an anomaly is entirely supported by data streams of a preset level or higher, its validity is unaffected by low-confidence data and can be confirmed without verification. Verification is only necessary when the supporting data contains low-confidence data streams, which could distort the anomaly. Gating based on the confidence level structure limits the anomalies to those supported by genuinely low-confidence data, further narrowing the verification scope based on the initial fusion analysis. Simultaneously, it identifies the low-confidence data streams requiring verification as subsequent replacement verification targets.
[0043] Following the previous example, the credibility level of the disease development item is composed of first (a1), first (a2), third (a3), first (a4), and second (a5). Among them, the publicly crawled data stream of a3 is at the third credibility level, which is lower than the preset second credibility level. Therefore, the data stream of a3 is determined to be a low credibility data stream, and a replacement verification strategy is performed on it.
[0044] Step S140: Execute the replacement verification strategy and select a replacement data stream from the multi-source heterogeneous data streams.
[0045] In an optional implementation of this embodiment, selecting a replacement data stream may specifically include: preferentially selecting a data stream that corresponds to the same cultural relic object as the low-confidence data stream, is of the same analysis item type, has a higher confidence level than the low-confidence data stream, and whose time interval between the time marker and the low-confidence data stream is no greater than a preset time window, as a replacement data stream; when no data stream meets the aforementioned conditions, selecting a data stream that corresponds to the same analysis item type as the low-confidence data stream, has a higher confidence level than the low-confidence data stream, and corresponds to other cultural relic objects belonging to the same category as the cultural relic object, as a replacement data stream; when multiple data streams meet the conditions, selecting the one with the earlier order in order of confidence level from high to low and time marker from recent to distant, as the replacement data stream. Continuing the previous example, for a3 (risk factor item), among the data streams that are the same cultural relic object A, are also risk factor items, have a confidence level higher than the third confidence level, and fall within the time window, the data stream of the official on-site verification risk report (first confidence level, content: on-site verification did not find the beam frame collapsed, but confirmed the risk of rotten wooden pillars and water accumulation) can be selected as the replacement data stream R.
[0046] Selecting a replacement data stream provides a more reliable reference for verifying whether low-confidence data is distorted. This reference must correspond to the same cultural relic object, the same type of analysis item, and have a higher confidence level to be comparable. The cascading order described above, from the same cultural relic object and comparable time to other cultural relic objects of the same category as a fallback, balances the comparability and availability of references: high-confidence records of the same cultural relic object and similar time are the most comparable and are therefore selected first; when such records are unavailable, high-confidence records of the same type of cultural relic object are used as a fallback to avoid frequent lack of available references due to overly strict matching conditions. Through this selection, it is verified that the replacement data stream is a high-confidence record already existing in the database and does not rely on constructed or assumed data.
[0047] It should be noted that the value of the preset time window is closely related to the timeliness of the analysis item type: for preservation items with slow evolution, such as decaying wooden components and weathered bricks and stones, the replacement data stream is still comparable even if it is several months apart from the data stream to be replaced. A larger time window (e.g., 3 to 12 months) can be used to avoid the window being too narrow, frequently failing to find comparable high-reliability records and being forced to enter a pending review phase. For monitoring items with rapid changes, such as body displacement and environmental temperature and humidity, an excessively large time window may cause old monitoring records from several months ago to be selected as replacement data streams, replacing newly captured danger signals and resulting in false filtering. Therefore, a smaller time window (e.g., several days to several weeks) is preferable. In a production environment, configurations can be made separately according to the analysis item type (the above ranges are examples; the actual range should be determined based on the data update cycle).
[0048] Step S150: Replace the analysis item corresponding to the low-confidence data stream with the analysis item corresponding to the replacement data stream, perform a second fusion analysis, and determine whether the anomaly still holds true.
[0049] like Figure 4 As shown, the analysis item corresponding to the replacement data stream R is replaced with the analysis item corresponding to the low-confidence data stream, while the remaining analysis items remain unchanged. The analysis is then re-performed using the same fusion analysis mechanism as S120. The second fusion analysis is used to determine whether the original anomaly item still holds true after replacing the low-confidence data with data of a higher confidence level. Continuing the previous example, a3 (internet crawling item) is replaced with R (official verification risk factor item), while a1, a2, a4, and a5 remain unchanged. The fusion analysis is then redone. Since the decay of the wooden pillar (a1), continuous heavy rainfall (a2), and increased displacement (a4) are sufficient to support the disease development item, the disease development item still holds true in the second fusion analysis.
[0050] Step S160: Retain the analysis items corresponding to the low-confidence data stream and add the analysis items corresponding to the replacement data stream together to perform the third fusion analysis to determine whether the anomaly still holds true.
[0051] Unlike the second fusion analysis, the third fusion analysis does not remove the analysis item corresponding to the low-confidence data stream. Instead, it retains the analysis item while adding the analysis item corresponding to the replacement data stream, allowing both to participate in the same fusion analysis. Understandably, the same replacement data stream R is used in two stages in this scheme: in the second fusion analysis, its corresponding analysis item replaces the analysis item corresponding to the low-confidence data stream to verify whether the establishment of an anomaly depends on the low-confidence data; in the third fusion analysis, its corresponding analysis item and the analysis item corresponding to the low-confidence data stream participate in the analysis simultaneously to verify whether the low-confidence data still satisfies the anomaly determination condition when higher-confidence data is involved.
[0052] The second and third fusion analyses are set up because a single replacement is insufficient to distinguish between two types of low-confidence data: one type is low-confidence data that does not determine the validity of an anomaly, and the other type is low-confidence data with distorted values that incorrectly satisfy the anomaly detection conditions. The second fusion analysis replaces the low-confidence data stream's analysis items with those of the replacement data stream: if the anomaly still holds after replacement, it indicates that the anomaly is supported by the remaining data and is unrelated to the participation of the low-confidence data; if the anomaly does not hold after replacement, it indicates that the anomaly's validity is related to the low-confidence data. The third fusion analysis incorporates the replacement data stream's analysis items while retaining the low-confidence data stream's analysis items: if the anomaly still holds, it indicates that even with the participation of data of a higher confidence level, the low-confidence data still satisfies the anomaly detection conditions, meaning its value is inconsistent with the higher-confidence data and still satisfies the anomaly detection conditions; if the anomaly does not hold, it indicates that the low-confidence data is insufficient to independently satisfy the anomaly detection conditions. Therefore, the second and third fusion analyses provide judgment criteria on two aspects: whether the anomaly depends on the low-confidence data, and whether the low-confidence data still satisfies the anomaly judgment criteria when higher-confidence data is involved. The combination of the two can distinguish between low-confidence data that does not play a decisive role, low-confidence data with distorted values, and real anomalies.
[0053] Step S170: Based on the results of the second and third fusion analyses, determine whether to filter out low-confidence data streams and whether to retain outliers.
[0054] In an optional implementation of this embodiment, the following steps may be taken: First, if the anomaly still holds true in the second fusion analysis, it indicates that the anomaly can be established by the remaining data without relying on the low-confidence data stream. The anomaly is retained and identified as a confirmed anomaly, and the low-confidence data stream is not filtered out. Continuing the previous example, if the disease development item still holds true in the second fusion analysis, the anomaly is confirmed, and the public opinion data stream is not filtered out. Second, if the anomaly does not hold true in the second fusion analysis but still holds true in the third fusion analysis, it indicates that the anomaly no longer holds true after being replaced with data of a higher confidence level. However, with the participation of the data of the higher confidence level, the low-confidence data still satisfies the anomaly determination condition. That is, the value of the low-confidence data is inconsistent with the data of the higher confidence level and satisfies the anomaly determination condition. Based on this, the low-confidence data stream is determined to be a data stream with distorted values. The low-confidence data stream is filtered out, and the anomaly is also filtered out. Third, if an anomaly does not hold true in the second fusion analysis and also does not hold true in the third fusion analysis, it indicates that the anomaly does not hold true after being replaced with data of a higher confidence level. Furthermore, the low-confidence data is insufficient to satisfy the anomaly judgment condition when data of a higher confidence level is involved, and it is not sufficient to determine whether its value is distorted. Therefore, the anomaly is retained and identified as an anomaly to be reviewed. The low-confidence data stream is not screened out and is handed over to a human for further review.
[0055] It should be understood that both data stream filtering and anomaly filtering are irreversible. False filtering of genuine anomalies could delay the handling of potential cultural relic hazards. Therefore, filtering is only performed when the second fusion analysis fails but the third fusion analysis still succeeds. Cases meeting only one of these criteria or neither are identified as confirmed anomalies or anomalies requiring review, ensuring that the handling aligns with the judgments provided by the two fusion analyses. This achieves a balance between filtering low-reliability distorted data and avoiding the false filtering of genuine anomalies.
[0056] To facilitate understanding of the above handling process, let's take another ancient site (cultural relic object B) as an example: Analysis item b1 (risk factor item) originates from publicly crawled internet data streams (third-trusted), with the value being online reports of construction damage and large-scale collapse of the site itself. The other analysis items b2 (monitoring status item, first-trusted, no abnormal displacement detected) and b3 (preservation status item, first-trusted, recent inspection records normal) are both normal. In the first fusion analysis, b1 matches the association rule for human disturbance leading to site damage, confirming the existence of an external risk item. The official on-site verification report data stream of the same type and first-trusted nature (with the value "no construction damage found during on-site verification, site intact") is selected as the replacement data stream R'. In the second fusion analysis, after replacing b1 with R', the external risk item is no longer valid. In the third fusion analysis, after retaining b1 and adding R', the external risk item still exists. Therefore, the publicly crawled internet data stream containing b1 is removed, and this external risk item is also filtered out. For example, taking cultural relic object C as an example, its low-confidence risk factor item c1 is suspected to be slight water seepage. After the second fusion analysis replaced c1 with high-confidence data, the monitoring anomaly item was not established. After the third fusion analysis retained c1 and added the replacement data stream, the monitoring anomaly item was also not established. Therefore, the monitoring anomaly item was determined to be an anomaly item to be reviewed.
[0057] Based on the above implementation methods, the following further explains the handling of missing replacement data streams, the dynamic maintenance of trust levels, and the output of anomaly analysis results, which can be understood as a supplement to the above implementation methods in long-term operation scenarios.
[0058] In an optional implementation of this embodiment, if there is still no replacement data stream after selection in S140 (for example, in a scenario of preventing theft in an ancient tomb, an external risk item supported by a third trusted open-source intelligence data stream has no data stream of the same type with a higher level of trust available for replacement in the same cultural relic object and similar cultural relic objects), the abnormal item is identified as an abnormal item to be reviewed, and a supplementary collection request for the low-trust data stream is generated; the supplementary collection result corresponding to the supplementary collection request is converted into an analysis item and used as the input for the next round of fusion analysis.
[0059] For example, supplementary data collection requests can be dispatched as on-site verification tasks, targeted retrieval of official reports, or encrypted monitoring sampling. The collection results are converted into high-confidence analysis items by S110 and used as input for the next round of fusion analysis. This allows the next round of fusion analysis to re-verify the original anomalies to be reviewed, forming a closed-loop process of judgment, supplementary data collection, and re-judgment. The reason for not screening or filtering when there is no replacement data stream, but instead proceeding to supplementary data collection and review, is that there is a lack of a higher-confidence level reference for determining whether low-confidence data is distorted. Obtaining this reference through supplementary data collection and then verifying it in the next round ensures that screening and filtering are only performed when the aforementioned dual conditions are met.
[0060] like Figure 5 As shown, in an optional implementation of this embodiment, the reliability level of each data stream can also be dynamically maintained. Specifically, this includes: counting the number of confirmed anomalies generated by the corresponding analysis item in each data stream during several rounds of fusion analysis, and the number of filtered anomalies; increasing the reliability level of the data stream when the ratio of confirmed anomalies to filtered anomalies is greater than a first threshold; and decreasing the reliability level of the data stream when the ratio is less than a second threshold; wherein the first threshold is greater than the second threshold. This dynamic maintenance is implemented because if the reliability level is determined solely by the static source, data streams with varying quality within the same source will be assigned the same reliability level. Calibrating the reliability level by the relative number of confirmed and filtered anomalies generated by each data stream in each round allows the reliability level to gradually reflect the actual reliability of the data stream as the process continues. Therefore, the credibility level of data streams such as publicly crawled data from the Internet is no longer determined solely by the static source: in multi-round fusion analysis, the credibility level of data streams with more confirmed anomalies generated by their analysis items is gradually improved, while the credibility level of data streams with more anomalies filtered out is reduced or even no longer participates in subsequent fusion analysis, thus allowing for dynamic calibration of source credibility with quantifiable statistics.
[0061] It should be noted that the first threshold and the second threshold mentioned above need to be set together. The interval between the second threshold and the first threshold is the interval in which the trust level is not adjusted: The first threshold determines the triggering condition for increasing the trust level. The larger the value of the first threshold, the more confirmed anomalies are required relative to the filtered anomalies to reach the threshold, and the more rounds are required for the data stream to increase its trust level; the smaller the value of the first threshold, the more rounds the data stream can increase its trust level. The second threshold determines the triggering condition for decreasing the trust level. The larger the value of the second threshold, the more situations where the ratio is lower than the threshold, and the higher the frequency of the data stream being decreased in trust level; the smaller the value of the second threshold, the more filtered anomalies are required relative to the confirmed anomalies for the data stream to decrease its trust level. To ensure that the confidence level of the data stream remains constant when the ratio is between the second and first thresholds, the first threshold is greater than the second threshold. For example, the first threshold can be 2.0 and the second threshold can be 0.5. That is, when the ratio is higher than 2.0, the confidence level is increased; when it is lower than 0.5, the confidence level is decreased; and when it is between 0.5 and 2.0, it remains unchanged. The above values are examples and can be adjusted according to the statistical distribution of confirmed and filtered anomalies in the data stream.
[0062] In an optional implementation of this embodiment, an anomaly analysis report can also be generated based on confirmed anomalies and / or anomalies pending review. The anomaly analysis report includes the confirmed anomalies and / or anomalies pending review, the corresponding data streams, and the reliability level of their respective data streams. The anomaly analysis report can be output to the comprehensive information management platform for cultural relic protection for security risk assessment, providing traceable assessment evidence for daily management of cultural relics, hazard investigation, repair project initiation, and protection decisions.
[0063] It should be noted that the division of the steps in the above method is for illustrating each processing procedure. In actual implementation, they can be combined or split, and the logical order between the steps can be adjusted without contradiction. The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0064] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0065] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention; various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for fusing multi-source heterogeneous data for cultural relic protection, characterized in that, include: The multi-source heterogeneous data streams corresponding to the cultural relics are converted into analysis items for cultural relic anomaly analysis. Each analysis item is marked with the confidence level of its source data stream, which is determined based on the data source. Perform a first fusion analysis on each analysis item to determine if any outliers exist; When an anomaly exists, the decision on whether to execute a replacement verification strategy is based on the credibility level of the source data stream upon which the anomaly is based. The replacement verification strategy includes: selecting, from the source data streams of the analysis items on which the anomaly is based, the data stream that corresponds to the same cultural relic object and the same analysis item type as the low-confidence data stream and has a higher confidence level than the low-confidence data stream, as the replacement data stream; Replace the analysis item corresponding to the low-confidence data stream with the analysis item corresponding to the replacement data stream, perform a second fusion analysis, and determine whether the anomaly item still holds true. The analysis items corresponding to the low-confidence data stream are retained, and the analysis items corresponding to the replacement data stream are added together for a third fusion analysis to determine whether the anomaly is still valid. Based on the results of the second and third fusion analyses, it is determined whether to filter out the low-confidence data streams and whether to retain the anomalies.
2. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 1, characterized in that, The step of determining whether to filter out the low-confidence data stream and whether to retain the anomalies based on the results of the second and third fusion analyses includes: If the anomaly still holds in the second fusion analysis, the anomaly is retained, and the low-confidence data stream is retained. If the anomaly does not hold true in the second fusion analysis but still holds true in the third fusion analysis, the low-confidence data stream is filtered out, and the anomaly is also filtered out. If the anomaly is not valid in the second fusion analysis and is not valid in the third fusion analysis, the anomaly is retained and identified as an anomaly to be reviewed, and the low-confidence data stream is not filtered out.
3. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 2, characterized in that, The trust levels include a first trust level, a second trust level, and a third trust level that decrease sequentially. The data streams corresponding to on-site measured data, archival data, and officially published data are marked as the first trust level; the data streams corresponding to business management data and administrative management data are marked as the second trust level; and the data streams corresponding to publicly crawled data from the Internet are marked as the third trust level.
4. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 1, characterized in that, The analysis items include at least one of the following: preservation status items, risk factor items, monitoring status items, and disposal-related items; The original fields of each data stream are converted into analysis items with unified object identifiers, time identifiers, analysis item types, and analysis item values, and the analysis items are bound to the identifiers and trust levels of their source data streams.
5. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 1, characterized in that, The first fusion analysis, the second fusion analysis, and the third fusion analysis are fusion analyses with different analysis items but the same judgment method. The fusion analysis includes: Align the analysis items involved in the analysis according to the object identifier and analysis item type to obtain the abnormal characterization data of the cultural relic object; Anomaly analysis indicators are determined based on the anomaly characterization data. When the anomaly analysis indicators meet the preset anomaly judgment conditions, the existence of the anomaly item is determined, and the analysis item on which the anomaly item is generated and its source data stream are recorded.
6. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 1, characterized in that, Selecting the replacement data stream includes: Prioritize selecting a data stream that corresponds to the same cultural relic object and analysis item type as the low-confidence data stream, has a higher confidence level than the low-confidence data stream, and whose time interval between the time marker and the low-confidence data stream is no greater than a preset time window, as the replacement data stream. If no replacement data stream meets the conditions, select a data stream that corresponds to the same analysis item type as the low-confidence data stream, has a higher confidence level than the low-confidence data stream, and corresponds to other cultural relic objects belonging to the same category as the cultural relic object, as the replacement data stream. When multiple data streams meet the conditions, the one with the earlier time stamp is selected as the replacement data stream, in descending order of trust level and in ascending order of time stamp.
7. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 3, characterized in that, The decision to execute a replacement verification strategy is based on the credibility level of the source data stream upon which the anomaly is based, including: If the confidence level of the source data stream of the analysis item on which the anomaly is based is not lower than the preset level, the existence of the anomaly is confirmed and the replacement verification strategy is not executed. If there is a data stream with a confidence level lower than a preset level in the source data stream of the analysis item on which the anomaly is based, the data stream is identified as the low-confidence data stream, and the replacement verification strategy is executed. The preset level is the second trust level.
8. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 6, characterized in that, If the replacement data stream still does not exist after selection, the anomaly is identified as an anomaly to be reviewed, and a supplementary collection request is generated for the low-confidence data stream; the supplementary collection results corresponding to the supplementary collection request are converted into analysis items and used as input for the next round of fusion analysis.
9. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 3, characterized in that, The number of confirmed anomalies generated by the corresponding analysis items of each data stream in several rounds of fusion analysis and the number of filtered anomalies are counted. When the ratio of the number of confirmed anomalies to the number of filtered anomalies is greater than a first threshold, the credibility level of the data stream is increased. When the ratio is less than a second threshold, the trust level of the data stream is reduced; wherein the first threshold is greater than the second threshold.
10. The multi-source heterogeneous data fusion method for cultural relic protection according to claim 9, characterized in that, Also includes: An anomaly analysis report is generated based on the confirmed anomalies and / or the anomalies to be reviewed. The anomaly analysis report includes the confirmed anomalies and / or the anomalies to be reviewed, the corresponding data streams, and the confidence level of their corresponding data streams.
Citation Information
Patent Citations
Heterogeneous multi-source data fusion processing method, device and system for cultural relic protection
CN115391314A