Ancient book and historic site element screening method based on multi-dimensional scale

By employing a multi-dimensional benchmarking approach, combined with cross-validation of historical materials and expert interviews, and utilizing clustering algorithms to identify the comprehensive value of ancient books and historical sites, the limitations of single-dimensional judgment are overcome, enabling precise screening of elements related to ancient books and historical sites and flexible and efficient decision-making.

CN121808616AInactive Publication Date: 2026-04-07XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-04-07
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN121808616A_ABST
    Figure CN121808616A_ABST
Patent Text Reader

Abstract

The invention discloses an ancient book and historic site element screening method based on a multi-dimensional scale. According to the method, the limitation of single-dimensional judgment in traditional screening is effectively avoided by constructing a multi-dimensional evaluation system and combining a depth analysis algorithm. According to the method, firstly, the authenticity of data is ensured through historical cross check and expert interviews, then potential associations among all dimensions are mined, and candidate elements which are balanced in multiple dimensions or have unique values are accurately recognized. The method not only can comprehensively reflect the comprehensive value of the ancient books and historic sites, but also can prevent the overall significance from being ignored due to the shortages of individual dimensions, so that the screening result better meets the actual requirements, and a reliable basis is provided for subsequent protection, research or utilization. Meanwhile, a closed-loop optimization mechanism is formed in the whole process, the deep analysis algorithm can timely find out anomalies or unreasonable parts in the data, adjustment and improvement of subsequent steps are promoted, the screening process is more flexible and efficient, and decision making is more scientific and reasonable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of ancient book and historical site protection technology, specifically a method for screening ancient book and historical site elements based on a multi-dimensional scale. Background Technology

[0002] The selection of elements for ancient books and historical sites is a crucial step in cultural heritage protection and research. Through systematic analysis and scientific evaluation, it involves identifying, extracting, and determining key elements with core value and representative characteristics from a wealth of historical information. This process aims to identify the fundamental elements constituting the historical, artistic, and scientific value of ancient books or historical sites, such as the version history, content, and author information of ancient books, or the architectural style, structural features, materials, craftsmanship, and historical context of historical sites. The selection process requires the comprehensive application of knowledge from multiple disciplines, including textual research, archaeology, architectural history, and art history. Based on clear value assessment standards, it eliminates secondary or later altered information, thereby accurately identifying the core carriers of authenticity and integrity. This process is not only the foundation for establishing archives, developing protection plans, and implementing restoration projects, but also a scientific prerequisite for deepening cultural interpretation and public education.

[0003] However, existing technologies often have limitations in the selection of ancient books and historical sites due to their reliance on a single dimension. They either rely solely on the age of the site or the degree of preservation, making it difficult to fully reflect the comprehensive value of the elements. Furthermore, the data verification process lacks cross-checking of multiple historical sources and quantitative integration of expert opinions, which can easily lead to data distortion or bias, thus limiting the scientific nature and accuracy of decision-making. Summary of the Invention

[0004] The purpose of this invention is to provide a method for screening elements of ancient books and historical sites based on a multi-dimensional scale in order to solve the problems mentioned above.

[0005] The technical solution adopted in this invention is as follows: a method for screening ancient book and historical site elements based on a multi-dimensional scale, the method comprising the following steps:

[0006] S1: Determine the core boundaries such as the type of target object, geographical range, and time span based on user needs. The subsequent S2 dimension system will be adjusted accordingly based on this objective.

[0007] S2: The evaluation criteria in S3 will be further refined based on six core dimensions: historical value, cultural representativeness, preservation integrity, research value, social impact, and utilization potential. The weight of each dimension will be allocated in conjunction with the objectives of S1.

[0008] S3: Design operable assessment levels and scores for each dimension of S2, and the subsequent data collection in S4 must strictly follow these details.

[0009] S4: Collect the original data of each candidate element under the S3 rules, verify the authenticity of the data through cross-verification of historical materials, expert interviews and other methods, and ensure that the data meets the evaluation requirements of S3, so as to provide a reliable foundation for the in-depth analysis of S5.

[0010] S5: Map candidate elements to the multidimensional space of S2, identify outlier elements through clustering algorithms, and explore hidden relationships between dimensions. The results will serve as a supplementary basis for the comprehensive score in S6, with priority given to marking unique elements.

[0011] S6: Weight the scores of each dimension of S3 according to the weights of S2, sum them, and adjust the results by combining the outlier features mined in S5 to generate a candidate element ranking table, providing preliminary results for the verification of S7.

[0012] S7: Invite domain experts to verify the rationality of the ranking results in S6, adjust the dimensional weights in S2 or the evaluation details in S3 based on the feedback, and repeat steps S3 to S6 until the results meet expectations.

[0013] In a preferred embodiment, in step S1, the filtering direction is first determined based on user needs, including three core types: protection priority, academic research, and cultural tourism development. The geographical scope needs to be precise down to the provincial administrative region or a specific cultural area, and the time span needs to be locked within a specific dynasty or year range, such as from the pre-Qin period to the Ming and Qing dynasties or before 1912. The dimensional system of subsequent steps will be dynamically adjusted according to the objective: if the objective is cultural tourism development, the social awareness dimension will be strengthened; if it is academic research, the citation rate dimension will be emphasized. Initial weights are also assigned to each objective to ensure the relevance of subsequent steps.

[0014] In a preferred embodiment, step S2 involves identifying six core dimensions: historical value, cultural representativeness, preservation integrity, research value, social impact, and utilization potential. The weight of each dimension is determined using the analytic hierarchy process (AHP) and allocated in conjunction with the objectives of S1: under the protection objective, preservation integrity 25%, historical value 20%, research value 20%, cultural representativeness 15%, social impact 10%, and utilization potential 10%; under the academic research objective, research value 30%, historical value 25%, preservation integrity 15%, cultural representativeness 15%, social impact 10%, and utilization potential 5%. Each dimension needs to have a clearly defined composition; for example, historical value includes the distance from the original date and its relevance to historical events, and preservation integrity includes the proportion of the original structure and the degree of damage.

[0015] In a preferred embodiment, in step S3, historical value is categorized by era into: pre-Qin and earlier (10 points), Qin-Han to Sui-Tang (8 points), Song-Yuan (6 points), Ming-Qing (4 points), and modern times (2 points); preservation integrity is categorized by original proportion and degree of damage into: intact (10 points), relatively good (8 points), average (5 points), endangered (3 points), and damaged (1 point); research value is categorized by core journal citation rate into: ≥50 times (10 points), 30-49 times (8 points), 15-29 times (6 points), and 5-14 times (4 points). <5 times 2 points; Social influence is categorized by media coverage level: National level ≥10 times 10 points, Provincial level 5-9 times 8 points, Municipal level 3-4 times 6 points, County level 1-2 times 4 points, No coverage 2 points; Cultural representativeness is categorized by intangible cultural heritage level: National level 10 points, Provincial level 8 points, Municipal level 6 points, County level 4 points, Not listed 2 points; Utilization potential is categorized by annual tourist volume: ≥100,000 10 points, 50,000-90,000 8 points, 10,000-40,000 6 points, <10,000 4 points, No utilization record 2 points. Each level requires clear judgment criteria, such as the intact level requiring 100% original structure and no damage.

[0016] In a preferred embodiment, step S4 specifically includes:

[0017] The first step includes: automated matching for cross-validation of historical materials. For each candidate element's key attributes (such as age, preservation status, and research value-related information), data from multiple sources such as official histories, local chronicles, academic monographs, and archaeological reports are collected. Based on historical materials issued by national authoritative institutions, the attribute descriptions of each source are compared using text similarity algorithms. The number of sources that reach a consensus on the attribute is counted, and a consistency score is calculated. If the score is lower than the threshold set by S3, it is marked as data to be verified.

[0018] The second step includes: quantitative integration of expert interviews, inviting 5-7 authoritative experts in the field to evaluate the marked data to be verified one by one. Each expert gives a credibility score and correction suggestions based on their own professional judgment. At the same time, weights are assigned according to the experts' qualifications and project participation experience. The opinions of all experts are integrated using a weighted average method to obtain the final corrected value of the data.

[0019] The third step includes data correction and consistency checks. The corrected data is then reintegrated into the original dataset and subjected to multi-source cross-validation again to ensure that the consistency score of all data meets the standard required by S3 before it can be output as reliable data for the S5 deep analysis stage.

[0020] In a preferred embodiment, the formula for the consistency score of cross-validation of historical materials in step S4 is:

[0021] ;

[0022] In the formula:

[0023] CS represents the consistency score.

[0024] N consistent It refers to the number of historical sources from which a consensus is reached regarding a certain attribute.

[0025] N total It is the total number of historical materials involved in the verification.

[0026] Avg sim It is the average text similarity between each consistent source and the benchmark historical material;

[0027] The formula for weighted integration of expert opinions is:

[0028] ;

[0029] In the formula:

[0030] ES represents the expert integrated score.

[0031] W i It is the weight of the i-th expert.

[0032] S i It is the rating of the i-th expert on the credibility of the data.

[0033] n is the total number of experts participating in the evaluation.

[0034] In a preferred embodiment, step S5 specifically includes:

[0035] Step 5-1: Based on the core dimensions determined in S2, standardize the scoring data of each element collected in S3:

[0036] The raw scores for each dimension are normalized using min-max to transform them into values ​​in the [0,1] interval, thus eliminating dimensional differences.

[0037] The dimension weights assigned by S2 are retained to prepare for subsequent weighted calculations.

[0038] Step 5-2: Perform K-means clustering to divide the feature groups, use the elbow rule to determine the number of clusters k, and divide the standardized feature data into 5 clusters, each cluster representing a group of features with similar characteristics;

[0039] Output the center vector CentroidC of each cluster as a reference for subsequent outlier detection.

[0040] Step 5-3: Isolation Forest for Outlier Detection: For each feature, calculate the outlier score AS using the Isolation Forest algorithm. i ;

[0041] Based on the clustering results: if element i belongs to cluster C, and its distance to the cluster center is much greater than the average distance within the cluster, and AS... iIf the value is greater than 0.7, it is marked as an outlier feature.

[0042] Step 5-4: For the dimensional system of S2, calculate the weighted correlation coefficient between any two dimensions to identify hidden relationships;

[0043] Output the top 3 dimension pairs with the strongest correlation as a reference for adjusting the S6 score.

[0044] Step 5-5: Organize the outlier elements and dimension correlation results into supplementary information; these results will serve as additional basis for the S6 comprehensive score.

[0045] In a preferred embodiment, in step S5, the outlier comprehensive scoring formula, which combines cluster distance and outlier score, is used to quantify the degree of uniqueness of the elements:

[0046] ;

[0047] In the formula:

[0048] OS i The outlier comprehensive score of element i

[0049] V i Represents the standardized multidimensional feature vector of element i

[0050] Centroid C The center vector of cluster C to which element i belongs.

[0051] Represents the Euclidean distance from element i to the cluster center.

[0052] AvgDist C This represents the average distance from all elements within cluster C to the center.

[0053] AS i The isolated forest anomaly score represents element i.

[0054] W anomaly This represents the outlier weighting coefficient (set according to the S1 target, such as 1.2 for the protection target and 1.0 for the cultural tourism target).

[0055] The formula for calculating the weighted Pearson correlation coefficient, which takes into account the dimensional weights, is as follows:

[0056] ;

[0057] In the formula:

[0058] This represents the weighted correlation coefficient between dimensions X and Y, ranging from -1 to 1;

[0059] WX represents the S2 weight of dimension X;

[0060] WY represents the S2 weight of dimension Y;

[0061] xj represents the standardized score of the j-th element in dimension X;

[0062] yj represents the standardized score of the j-th element in dimension Y;

[0063] xˉ represents the average standardized score of dimension X;

[0064] yˉ represents the average standardized score of dimension Y.

[0065] In a preferred embodiment, in step S6, a weighted summation method is used to calculate the comprehensive score. The scores of each dimension in S3 are multiplied by their corresponding weights in S2 and then summed. For example, a historical value of an element (8 points) multiplied by a weight of 25% yields 2 points, and preservation integrity (10 points) multiplied by a weight of 25% yields 2.5 points. The scores for all dimensions are calculated sequentially and then summed. Based on the outlier feature adjustment results mined in S5, outlier elements receive an additional 5% score on top of their comprehensive score. A candidate element ranking table is generated, arranged from highest to lowest comprehensive score. The initial screening quantity is determined according to user needs, such as selecting the top 50 for protection objectives and the top 30 for academic research objectives.

[0066] In a preferred embodiment, in step S7, 5-7 leading experts in the field are invited to participate in the verification, including historians, cultural relic protection experts, and scholars of ancient book collation. The experts review the ranking results one by one and provide suggestions for modification. Based on the feedback, the dimensional weights of S2 or the evaluation details of S3 are adjusted, such as adjusting the social influence weight from 10% to 15%, or adjusting the endangered status score for preservation integrity from 3 points to 4 points. After adjustment, steps S3 to S6 are repeated until the expert acceptance rate reaches over 90%, finally determining the screening results and establishing a dynamic adjustment mechanism to periodically update the screening list based on new data.

[0067] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0068] 1. This invention effectively avoids the limitations of single-dimensional judgment in traditional screening by constructing a multi-dimensional evaluation system and combining it with the S5 deep analysis algorithm. It first ensures the authenticity of the data through cross-verification of historical materials and expert interviews, and then uses the S5 algorithm to mine the potential correlations between various dimensions, accurately identifying candidate elements that perform well across multiple dimensions or possess unique value. This approach not only comprehensively reflects the overall value of ancient books and historical sites but also avoids overlooking their overall significance due to shortcomings in individual dimensions, making the screening results more aligned with practical needs and providing a reliable basis for subsequent protection, research, or utilization.

[0069] 2. In this invention, the dimensional weights can be adjusted according to different screening objectives, and the S5 algorithm can adapt to these dynamically changing needs, outputting targeted analysis results. Whether it's cultural heritage protection departments identifying priority protection objects, academic institutions searching for samples with research value, or cultural tourism departments developing unique resources, this method can quickly obtain a screening list that meets the objectives. Simultaneously, the entire process forms a closed-loop optimization mechanism. S5's in-depth analysis can promptly identify anomalies or irrationalities in the data, driving adjustments and improvements in subsequent steps, making the screening process more flexible and efficient, and ensuring more scientific and rational decision-making. Attached Figure Description

[0070] Figure 1 This is a schematic diagram illustrating the process principle of the present invention. Detailed Implementation

[0071] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0072] Reference Figure 1 A method for filtering ancient books and historical sites based on a multi-dimensional scale, comprising the following steps:

[0073] S1: Determine the core boundaries such as the type of target object, geographical range, and time span based on user needs. The subsequent S2 dimension system will be adjusted accordingly based on this objective.

[0074] S2: The evaluation criteria in S3 will be further refined based on six core dimensions: historical value, cultural representativeness, preservation integrity, research value, social impact, and utilization potential. The weight of each dimension will be allocated in conjunction with the objectives of S1.

[0075] S3: Design operable assessment levels and scores for each dimension of S2, and the subsequent data collection in S4 must strictly follow these details.

[0076] S4: Collect the original data of each candidate element under the S3 rules, verify the authenticity of the data through cross-verification of historical materials, expert interviews and other methods, and ensure that the data meets the evaluation requirements of S3, so as to provide a reliable foundation for the in-depth analysis of S5.

[0077] S5: Map candidate elements to the multidimensional space of S2, identify outlier elements through clustering algorithms, and explore hidden relationships between dimensions. The results will serve as a supplementary basis for the comprehensive score in S6, with priority given to marking unique elements.

[0078] S6: Weight the scores of each dimension of S3 according to the weights of S2, sum them, and adjust the results by combining the outlier features mined in S5 to generate a candidate element ranking table, providing preliminary results for the verification of S7.

[0079] S7: Invite domain experts to verify the rationality of the ranking results in S6, adjust the dimensional weights in S2 or the evaluation details in S3 based on the feedback, and repeat steps S3 to S6 until the results meet expectations.

[0080] In step S1, the selection direction is first determined based on user needs, including three core types: protection priority, academic research, and cultural tourism development. The geographical scope needs to be precise down to the provincial administrative region or specific cultural area, and the time span needs to be locked into a specific dynasty or year range, such as from the pre-Qin period to the Ming and Qing dynasties or before 1912. The dimensional system for subsequent steps will be dynamically adjusted according to the objectives: if the objective is cultural tourism development, the social awareness dimension will be strengthened; if it is academic research, the citation rate dimension will be emphasized. Initial weights are also assigned to each objective to ensure the relevance of subsequent steps.

[0081] In step S2, six core dimensions are identified: historical value, cultural representativeness, preservation integrity, research value, social impact, and utilization potential. The analytic hierarchy process (AHP) is used to determine the weight of each dimension, which is then allocated in conjunction with the objectives of S1: Under the protection objective, preservation integrity 25%, historical value 20%, research value 20%, cultural representativeness 15%, social impact 10%, and utilization potential 10%; under the academic research objective, research value 30%, historical value 25%, preservation integrity 15%, cultural representativeness 15%, social impact 10%, and utilization potential 5%. Each dimension needs to have a clearly defined composition; for example, historical value includes the time period and its relevance to historical events, and preservation integrity includes the proportion of the original structure and the degree of damage.

[0082] In step S3, historical value is categorized by era: Pre-Qin and earlier (10 points), Qin-Han to Sui-Tang (8 points), Song-Yuan (6 points), Ming-Qing (4 points), and modern times (2 points); preservation integrity is categorized by original proportions and degree of damage: well-preserved (10 points), relatively good (8 points), average (5 points), endangered (3 points), and damaged (1 point); research value is categorized by core journal citation rate: ≥50 times (10 points), 30-49 times (8 points), 15-29 times (6 points), 5-14 times (4 points), and <5 times (2 points); social... The impact of a cultural event is categorized by media coverage level: National level (≥10 times, 10 points); Provincial level (5-9 times, 8 points); Municipal level (3-4 times, 6 points); County level (1-2 times, 4 points); No coverage, 2 points. Cultural representativeness is categorized by intangible cultural heritage level: National level (10 points); Provincial level (8 points); Municipal level (6 points); County level (4 points); Not listed, 2 points. Utilization potential is categorized by annual tourist volume: ≥100,000 (10 points); 50,000-90,000 (8 points); 10,000-40,000 (6 points); <10,000 (4 points); No utilization record, 2 points. Each level requires specific criteria for judgment; for example, a "well-preserved" level requires 100% original structure with no damage.

[0083] Step S4 specifically includes:

[0084] The first step includes: automated matching for cross-validation of historical materials. For each candidate element's key attributes (such as age, preservation status, and research value-related information), data from multiple sources such as official histories, local chronicles, academic monographs, and archaeological reports are collected. Based on historical materials issued by national authoritative institutions, the attribute descriptions of each source are compared using text similarity algorithms. The number of sources that reach a consensus on the attribute is counted, and a consistency score is calculated. If the score is lower than the threshold set by S3, it is marked as data to be verified.

[0085] The second step includes: quantitative integration of expert interviews, inviting 5-7 authoritative experts in the field to evaluate the marked data to be verified one by one. Each expert gives a credibility score and correction suggestions based on their own professional judgment. At the same time, weights are assigned according to the experts' qualifications and project participation experience. The opinions of all experts are integrated using a weighted average method to obtain the final corrected value of the data.

[0086] The third step includes data correction and consistency checks. The corrected data is then reintegrated into the original dataset and subjected to multi-source cross-validation again to ensure that the consistency score of all data meets the standard required by S3 before it can be output as reliable data for the S5 deep analysis stage.

[0087] In step S4, the formula for the consistency score of cross-validation of historical materials is:

[0088] ;

[0089] In the formula:

[0090] CS represents the consistency score.

[0091] N consistent It refers to the number of historical sources from which a consensus is reached regarding a certain attribute.

[0092] N total It is the total number of historical materials involved in the verification.

[0093] Avg sim It is the average text similarity between each consistent source and the benchmark historical material;

[0094] The formula for weighted integration of expert opinions is:

[0095] ;

[0096] In the formula:

[0097] ES represents the expert integrated score.

[0098] W i It is the weight of the i-th expert.

[0099] Si It is the rating of the i-th expert on the credibility of the data.

[0100] n is the total number of experts participating in the evaluation.

[0101] Step S5 specifically includes:

[0102] Step 5-1: Based on the core dimensions determined in S2, standardize the scoring data of each element collected in S3:

[0103] The raw scores for each dimension are normalized using min-max to transform them into values ​​in the [0,1] interval, thus eliminating dimensional differences.

[0104] The dimension weights assigned by S2 are retained to prepare for subsequent weighted calculations.

[0105] Step 5-2: Perform K-means clustering to divide the feature groups, use the elbow rule to determine the number of clusters k, and divide the standardized feature data into 5 clusters, each cluster representing a group of features with similar characteristics;

[0106] Output the center vector CentroidC of each cluster as a reference for subsequent outlier detection.

[0107] Step 5-3: Isolation Forest for Outlier Detection: For each feature, calculate the outlier score AS using the Isolation Forest algorithm. i ;

[0108] Based on the clustering results: if element i belongs to cluster C, and its distance to the cluster center is much greater than the average distance within the cluster, and AS... i If the value is greater than 0.7, it is marked as an outlier feature.

[0109] Step 5-4: For the dimensional system of S2, calculate the weighted correlation coefficient between any two dimensions to identify hidden relationships;

[0110] Output the top 3 dimension pairs with the strongest correlation as a reference for adjusting the S6 score.

[0111] Step 5-5: Organize the outlier elements and dimension correlation results into supplementary information; these results will serve as additional basis for the S6 comprehensive score.

[0112] In step S5, the outlier comprehensive scoring formula, which combines cluster distance and outlier score, is used to quantify the degree of uniqueness of the elements:

[0113] ;

[0114] In the formula:

[0115] OS i The outlier comprehensive score of element i

[0116] V i Represents the standardized multidimensional feature vector of element i

[0117] Centroid C The center vector of cluster C to which element i belongs.

[0118] Represents the Euclidean distance from element i to the cluster center.

[0119] AvgDist C This represents the average distance from all elements within cluster C to the center.

[0120] AS i The isolated forest anomaly score represents element i.

[0121] W anomaly This represents the outlier weighting coefficient (set according to the S1 target, such as 1.2 for the protection target and 1.0 for the cultural tourism target).

[0122] The formula for calculating the weighted Pearson correlation coefficient, which takes into account the dimensional weights, is as follows:

[0123] ;

[0124] In the formula:

[0125] This represents the weighted correlation coefficient between dimensions X and Y, ranging from -1 to 1;

[0126] WX represents the S2 weight of dimension X;

[0127] WY represents the S2 weight of dimension Y;

[0128] xj represents the standardized score of the j-th element in dimension X;

[0129] yj represents the standardized score of the j-th element in dimension Y;

[0130] xˉ represents the average standardized score of dimension X;

[0131] yˉ represents the average standardized score of dimension Y.

[0132] In step S6, a weighted summation method is used to calculate the comprehensive score. The scores for each dimension in S3 are multiplied by their corresponding weights in S2 and then summed. For example, a historical value of 8 points multiplied by a weight of 25% yields 2 points, and preservation integrity of 10 points multiplied by a weight of 25% yields 2.5 points. The scores for all dimensions are calculated sequentially and then summed. Based on the outlier feature adjustment results mined in S5, outlier elements receive an additional 5% score on top of their comprehensive score. A candidate element ranking table is generated, arranged from highest to lowest comprehensive score. The initial screening quantity is determined according to user needs, such as selecting the top 50 for protection objectives and the top 30 for academic research objectives.

[0133] In step S7, 5-7 leading experts in the field are invited to participate in the verification, including historians, cultural relic protection experts, and scholars of ancient book collation. The experts review the ranking results one by one and provide suggestions for modification. Based on the feedback, the dimensional weights of S2 or the evaluation details of S3 are adjusted, such as increasing the weight of social impact from 10% to 15%, or increasing the score for the endangered status of preservation integrity from 3 points to 4 points. After adjustment, steps S3 to S6 are repeated until the expert acceptance rate reaches over 90%, thus finalizing the selection results and establishing a dynamic adjustment mechanism to regularly update the selection list based on new data.

[0134] As can be seen from the above, this invention effectively avoids the limitations of single-dimensional judgment in traditional screening by constructing a multi-dimensional evaluation system and combining it with the S5 deep analysis algorithm. It first ensures the authenticity of the data through cross-verification of historical materials and expert interviews, and then uses the S5 algorithm to mine the potential correlations between various dimensions, accurately identifying candidate elements that perform well across multiple dimensions or possess unique value. This approach not only comprehensively reflects the overall value of ancient books and historical sites but also avoids overlooking their overall significance due to shortcomings in individual dimensions, making the screening results more aligned with practical needs and providing a reliable basis for subsequent protection, research, or utilization.

[0135] In this invention, the dimensional weights can be adjusted according to different screening objectives, and the S5 algorithm can adapt to these dynamically changing needs, outputting targeted analysis results. Whether it's cultural heritage protection departments identifying priority protection objects, academic institutions searching for samples with research value, or cultural tourism departments developing unique resources, this method can quickly obtain a screening list that meets the objectives. Simultaneously, the entire process forms a closed-loop optimization mechanism. S5's in-depth analysis can promptly identify anomalies or irrationalities in the data, driving adjustments and improvements in subsequent steps, making the screening process more flexible and efficient, and ensuring more scientific and rational decision-making.

[0136] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0137] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for screening ancient book and historical site elements based on a multi-dimensional scale, characterized in that: The method includes the following steps: S1: Determine the core boundaries of the filtering object type, geographical range, and time span based on user needs. The subsequent S2 dimension system will be adjusted accordingly based on this objective. S2: Sorting out six core dimensions: historical value, cultural representativeness, preservation integrity, research value, social impact, and utilization potential. Combining the objectives of S1, the weights of each dimension are allocated. The evaluation criteria in S3 will be further refined for each dimension. S3: Design operable assessment levels and scores for each dimension of S2, and the subsequent data collection in S4 must strictly follow these details; S4: Collect the original data of each candidate element under the S3 rules, verify the authenticity of the data through cross-verification of historical materials and expert interviews, and ensure that the data meets the evaluation requirements of S3, so as to provide a reliable basis for the in-depth analysis of S5. S5: Map candidate elements to the multidimensional space of S2, identify outlier elements through clustering algorithms, and explore hidden relationships between dimensions. The results will serve as a supplementary basis for the comprehensive score of S6, with priority given to marking unique elements. S6: The scores of each dimension of S3 are weighted and summed according to the weights of S2. The results are adjusted by combining the outlier features mined in S5 to generate a candidate element ranking table, providing preliminary results for the verification of S7. S7: Invite domain experts to verify the rationality of the ranking results in S6, adjust the dimensional weights in S2 or the evaluation details in S3 based on the feedback, and repeat steps S3 to S6 until the results meet expectations.

2. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S1, the filtering direction is first determined based on user needs, including three core types: protection priority, academic research, and cultural tourism development; the geographical scope needs to be accurate to the provincial administrative region or specific cultural region, and the time span needs to be locked to a specific dynasty or year range.

3. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S2, six core dimensions are identified: historical value, cultural representativeness, preservation integrity, research value, social influence, and utilization potential. The weight of each dimension is determined using the analytic hierarchy process (AHP) and allocated in conjunction with the objectives of S1: Under the protection objective, preservation integrity is 25%, historical value is 20%, research value is 20%, cultural representativeness is 15%, social influence is 10%, and utilization potential is 10%; under the academic research objective, research value is 30%, historical value is 25%, preservation integrity is 15%, cultural representativeness is 15%, social influence is 10%, and utilization potential is 5%.

4. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S3, historical value is categorized by era into: Pre-Qin and earlier (10 points), Qin-Han to Sui-Tang (8 points), Song-Yuan (6 points), Ming-Qing (4 points), and modern times (2 points); preservation integrity is categorized by original proportion and degree of damage into: intact (10 points), relatively good (8 points), average (5 points), endangered (3 points), and damaged (1 point); research value is categorized by core journal citation rate into: ≥50 times (10 points), 30-49 times (8 points), 15-29 times (6 points), 5-14 times (4 points), and <5 times (2 points). Social influence is categorized by media coverage level as follows: National level (≥10 times, 10 points); Provincial level (5-9 times, 8 points); Municipal level (3-4 times, 6 points); County level (1-2 times, 4 points); No coverage, 2 points. Cultural representativeness is categorized by intangible cultural heritage level as follows: National level (10 points); Provincial level (8 points); Municipal level (6 points); County level (4 points); Not listed, 2 points. Utilization potential is categorized by annual tourist volume as follows: ≥100,000 (10 points); 50,000-90,000 (8 points); 10,000-40,000 (6 points); <10,000 (4 points); No utilization record, 2 points.

5. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: Step S4 specifically includes: The first step includes: automated matching for cross-validation of historical materials. For the key attributes of each candidate element, data from multiple sources such as official histories, local chronicles, academic monographs, and archaeological reports are collected. Based on the historical materials issued by national authoritative institutions, the attribute descriptions of each source are compared through text similarity algorithms. The number of sources that reach a consensus on the attribute is counted, and a consistency score is calculated. If the score is lower than the threshold set by S3, it is marked as data to be verified. The second step includes: quantitative integration of expert interviews, inviting 5-7 authoritative experts in the field to evaluate the marked data to be verified one by one. Each expert gives a credibility score and correction suggestions based on their own professional judgment. At the same time, weights are assigned according to the experts' qualifications and project participation experience. The opinions of all experts are integrated using a weighted average method to obtain the final corrected value of the data. The third step includes data correction and consistency checks. The corrected data is then reintegrated into the original dataset and subjected to multi-source cross-validation again to ensure that the consistency score of all data meets the standard required by S3 before it can be output as reliable data for the S5 deep analysis stage.

6. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S4, the formula for the consistency score of cross-validation of historical materials is: ; In the formula: CS represents the consistency score. N consistent It refers to the number of historical sources from which a consensus is reached regarding a certain attribute. N total It is the total number of historical materials involved in the verification. Avg sim It is the average text similarity between each consistent source and the benchmark historical material; The formula for the weighted integration score of expert opinions is: ; In the formula: ES represents the expert integrated score. W i It is the weight of the i-th expert. S i It is the rating of the i-th expert on the credibility of the data. n is the total number of experts participating in the evaluation.

7. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: Step S5 specifically includes: Step 5-1: Based on the core dimensions determined in S2, standardize the scoring data of each element collected in S3: The raw scores for each dimension are normalized using min-max to transform them into values ​​in the [0,1] interval, thus eliminating dimensional differences. The dimension weights assigned by S2 are retained to prepare for subsequent weighted calculations; Step 5-2: Perform K-means clustering to divide the feature groups, use the elbow rule to determine the number of clusters k, and divide the standardized feature data into 5 clusters, each cluster representing a group of features with similar characteristics; Output the center vector CentroidC of each cluster as a reference for subsequent outlier detection; Step 5-3: Isolation Forest for Outlier Detection: For each feature, calculate the outlier score AS using the Isolation Forest algorithm. i ; Based on the clustering results: if element i belongs to cluster C, and its distance to the cluster center is much greater than the average distance within the cluster, and AS... i If the value is greater than 0.7, it is marked as an outlier feature; Step 5-4: For the dimensional system of S2, calculate the weighted correlation coefficient between any two dimensions to identify hidden relationships; Output the top 3 dimension pairs with the strongest correlation as a reference for adjusting the S6 score; Step 5-5: Organize the outlier elements and dimension correlation results into supplementary information; these results will serve as additional basis for the S6 comprehensive score.

8. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S5, the outlier comprehensive scoring formula, which combines cluster distance and outlier score, is used to quantify the uniqueness of the elements: ; In the formula: OS i The outlier comprehensive score of element i V i Represents the standardized multidimensional feature vector of element i Centroid C The center vector of cluster C to which element i belongs. Represents the Euclidean distance from element i to the cluster center. AvgDist C This represents the average distance from all elements within cluster C to the center. AS i The isolated forest anomaly score represents element i. W anomaly This represents the outlier weighting coefficient; The formula for calculating the weighted Pearson correlation coefficient, which takes into account the dimensional weights, is as follows: ; In the formula: This represents the weighted correlation coefficient between dimensions X and Y, ranging from -1 to 1; W X This represents the S2 weight of dimension X; W Y This represents the S2 weight of dimension Y; xj represents the standardized score of the j-th element in dimension X; yj represents the standardized score of the j-th element in dimension Y; xˉ represents the average standardized score of dimension X; yˉ represents the average standardized score of dimension Y.

9. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S6, a weighted summation method is used to calculate the comprehensive score. The scores of each dimension in S3 are multiplied by the corresponding weights in S2 and then summed. For example, the historical value of a certain element is 8 points, which is multiplied by a weight of 25% to get 2 points, and the preservation integrity is 10 points, which is multiplied by a weight of 25% to get 2.5 points. The scores of all dimensions are calculated and summed in turn. Combined with the outlier feature adjustment results mined in S5, the outlier elements are given an additional 5% score on the basis of the comprehensive score. Generate a candidate element ranking table, sorting them from highest to lowest based on the comprehensive score. Determine the initial screening quantity according to user needs, such as selecting the top 50 under the protection objective and the top 30 under the academic research objective.

10. The method for screening ancient book and historical site elements based on a multi-dimensional scale as described in claim 1, characterized in that: In step S7, 5-7 authoritative experts in the field are invited to participate in the verification, including historians, cultural relic protection experts, and scholars of ancient book collation; the experts review the sorting results one by one and put forward modification opinions.