Information extraction management analysis and comprehensive database system based on chinese archaeological report

By designing a comprehensive database system based on Chinese archaeological reports, the problems of non-automated information extraction, inconsistent data formats, and missing analysis modules in existing database systems were solved. This enabled efficient and automated archaeological data processing and analysis, improving the overall efficiency and accuracy of the database system.

CN117473108BActive Publication Date: 2025-12-30FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311406027.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-12-30
Estimated Expiration
2043-10-27

AI Technical Summary

Technical Problem

Existing database systems lack automated archaeological information extraction processes, have inconsistent data types and formats, and lack problem-oriented information extraction and data analysis modules, making it difficult to effectively merge and analyze archaeological data.

Method used

A comprehensive database system for information extraction, management, and analysis based on Chinese archaeological reports was designed. It includes an information extraction module, a data analysis module, and a data storage module. It employs image segmentation, text processing, and feature extraction techniques, combined with methods such as elliptic Fourier analysis, K-means clustering, and BR coefficient calculation, to achieve automated information extraction and data analysis.

Benefits of technology

It has enabled automated processing and analysis of archaeological data, improved the efficiency of database construction, ensured data quality consistency, supported problem-oriented analysis, and improved the accuracy and efficiency of data retrieval and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117473108B_ABST
    Figure CN117473108B_ABST
Patent Text Reader

Abstract

The application discloses a Chinese archaeological report-based information extraction management analysis and comprehensive database system and relates to the technical field of data management systems.The technical scheme is as follows: the system comprises an information extraction module, a data analysis module and a data storage module;the information extraction module is used for picture segmentation, text processing and feature extraction;the data analysis module is used for information preprocessing and analysis;the data storage module is used for classified storage of the information analyzed by the data analysis module and the information extracted by the information extraction module and enables data search;the classified storage adopts a three-level directory index form for storage.The system is based on Chinese archaeological reports and simultaneously comprises a comprehensive archaeological information database system of the whole process of information extraction, information storage and data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management system technology, and more specifically, to a comprehensive database system for information extraction, management, and analysis based on Chinese archaeological reports. Background Technology

[0002] In archaeological work, the organization, statistics, and analysis of data, as well as the management of various cultural relics, are complex tasks that require a great deal of manpower and resources. Even then, it is still not possible to quickly and comprehensively grasp the situation of the data. If database technology is introduced into archaeological work, it can not only greatly simplify the data organization work, but also integrate the planning, management, and progress of the entire archaeological work into a systematic process, turning the preparation, implementation, organization, and research of the work into a closely connected workflow.

[0003] In recent years, with the development of information technology, the demand for establishing large-scale archaeological databases has been increasing. The construction of large-scale archaeological databases can greatly improve the efficiency of archaeological research and broaden its perspective. In traditional archaeological research, researchers often need to compare and analyze artifacts from different sites based on empirical knowledge. However, research based on large-scale quantitative archaeological databases can conduct horizontal and vertical comparisons between sites and artifacts on a larger spatiotemporal scale, revealing patterns about ancient societies that cannot be grasped by experience alone.

[0004] Currently, this type of method has gradually become a new trend in archaeological research worldwide in recent years. In 2021, Martine Robbeets and dozens of other linguists, geneticists, and archaeologists jointly published an article in *Nature* titled "Triangulation Supports Agricultural Spread of the Transeurasian Languages." Based on 172 discrete archaeological features from 252 archaeological sites in China, Japan, South Korea, and Russia, the article constructed a Bayesian phylogenetic tree of archaeological cultures for the first time using the Monte Carlo Markov chain method. This successful combination of biological and archaeological methods provides a good example for the construction of subsequent archaeological databases: that is, the construction of quantitative archaeological databases should be based on subsequent research objectives and methodological design, rather than simply an unconscious accumulation of all recorded materials from the excavation process.

[0005] This type of new database system can be called a "quantitative archaeological database system". Compared with traditional archaeological databases, its characteristics can be summarized in the following two points: 1) Large data volume, including a large amount of archaeological site information over a wide range of time and space, rather than archaeological sites in a specific region or period; 2) The data entry and export formats are convenient for subsequent data analysis, and are often stored in the form of tables rather than pictures or text.

[0006] However, existing database systems still have the following technical problems:

[0007] 1. Current database systems lack automated processes for extracting archaeological information.

[0008] Currently, there is no Chinese-based quantitative archaeological database system available for researchers to use, and traditional database construction still relies on manual collection of information from archaeological reports.

[0009] 2. Current database systems lack a database framework with a unified format and data types.

[0010] Currently, international archaeological database systems are often organized in a bottom-up manner, with the main data source being field data shared by researchers. This results in inconsistent data types and formats, varying data quality, and difficulties in merging and analyzing archaeological data from different sites.

[0011] 3. Current database systems lack embedded problem-oriented information extraction and data analysis modules.

[0012] Current archaeological database systems primarily aim to share firsthand data from field excavations, but they do not address actual scientific questions, and quantitative research methods are often disconnected from the database system itself. Summary of the Invention

[0013] The purpose of this invention is to provide a comprehensive database system for information extraction, management, and analysis based on Chinese archaeological reports. This system is based on Chinese archaeological reports and includes a comprehensive archaeological information database system covering the entire process of information extraction, information storage, and data analysis.

[0014] The above-mentioned technical objective of the present invention is achieved through the following technical solution: an integrated database system for information extraction, management, and analysis based on Chinese archaeological reports, including an information extraction module, a data analysis module, and a data storage module;

[0015] The information extraction module is used for image segmentation, text processing, and feature extraction.

[0016] The data analysis module is used to preprocess and analyze information.

[0017] The data storage module is used to classify and store the information analyzed in the data analysis module and the information extracted in the information extraction module, and it can also perform data search; the classification and storage adopts the form of a three-level directory index.

[0018] The data analysis module includes image information preprocessing and text information preprocessing for information preprocessing.

[0019] The information analysis includes elliptical Fourier analysis, single-site level analysis, and multi-site level analysis.

[0020] The single-site level analysis includes the analysis of relationships between artifacts, the analysis of relationships between features, and the analysis of relationships between artifacts and features;

[0021] The method for analyzing the relationships between the artifacts is to use the K-means clustering algorithm to classify the artifacts unearthed at the site as a whole.

[0022] The method for analyzing the relationship between features is as follows: using the Fisher exact test function to test the correlation between the two, the option to calculate and statistically test the Point-biserial correlation coefficient, and the option to calculate and statistically test the Pearson correlation coefficient to obtain the correlation between discrete and continuous features;

[0023] The relationship analysis method between the artifacts and features is as follows: Principal component analysis is used to reduce the dimensionality of the images and discrete features, and feature weights are calculated based on the principal component variance interpretation degree to evaluate the classification efficiency of different archaeological features.

[0024] The multi-site level analysis includes cultural factor analysis and BR coefficient calculation;

[0025] The analysis method for cultural factors is as follows: First, select a group of artifacts with typical cultural significance and good preservation from a site for examination, while removing missing artifacts; then, select archaeological cultures for comparison from the database system; finally, determine the cultural affiliation of each artifact in the selected site artifact group.

[0026] The BR coefficient calculation is used to help users quickly obtain the cultural similarity between the site and surrounding sites, and to visually display the cultural geography of the site in the form of a heat map.

[0027] Furthermore, the specific method for image information preprocessing is as follows:

[0028] S1: After converting the binarized artifact line map in PNG format output by the information extraction module into JPG format, the gap filling algorithm is used to fill the interior of the line map to improve the contour recognition accuracy. Finally, the processed images are read in batches into the image processing module based on R language.

[0029] S2: After the user sets a threshold, the image parser automatically identifies the outer contour of the image. The identified outer contour is the x and y coordinates of each pixel.

[0030] S3: After obtaining the coordinates, calculate the area of ​​the contour, and then standardize all contours according to their area size;

[0031] Furthermore, the specific method for preprocessing the text information is as follows:

[0032] S1: Merge similar quantitative archaeological features output by the information extraction module based on the user's analysis needs;

[0033] S2: Perform data cleaning on the merged data.

[0034] Furthermore, the specific method of data cleaning in S2 is to filter out features with insufficient information that cannot be used for analysis.

[0035] Furthermore, the specific method of data cleaning in S2 is as follows: by setting a filtering threshold, features with a frequency lower than the set threshold are removed; after removal, the original data is used directly for analysis, and then the k-nearest neighbor data imputation method is used to fill in the missing values.

[0036] Furthermore, the specific method for determining cultural affiliation is as follows: First, calculate the distances between the coefficients of the object images after elliptical Fourier transform and the distances between the objects unearthed from all archaeological cultural sites; then, sort all the distances and determine the cultural attributes based on the set proportional threshold.

[0037] Furthermore, the specific steps for calculating the BR coefficient are as follows: First, calculate the proportion of each specific pottery type in the total amount of pottery unearthed at the two sites; then sum the absolute values ​​of the difference between the percentages of the two sites, and the corrected BR coefficient ranges from [0,1]; then, the user categorizes the artifacts of interest and selects the comparison sites to be compared; after the selection is completed, the data analysis module calculates the relative abundance of each type of artifact at the target site and all comparison sites, and performs BR coefficient calculation, finally outputting the BR coefficient vector of each comparison site and the target site.

[0038] Furthermore, the specific method of image segmentation in the information extraction module is as follows:

[0039] S1: Convert each page of the archaeological report PDF file into an image stored in array format and process them separately;

[0040] S2: Convert the image to grayscale and then convert it into a binary image with a black background and white outline;

[0041] S3: For each binary image, use a dilation algorithm to expand its outline range to bridge gaps caused by PDF clarity or scanning accuracy issues;

[0042] S4: Detect small regions in a binary image based on the connectivity of pixels to remove noise and redundant information.

[0043] S5: Perform a closing operation on the extracted image to close the outline;

[0044] S6: Performs edge-based contour detection on the processed image, and automatically cuts the image into rectangular cropping ranges using the outermost x and y coordinates of the contours, outputting it in the form of a compressed package.

[0045] Furthermore, the specific method of text processing in the information extraction module is as follows:

[0046] S1: Import the image and use regular expressions to match the legend text within it;

[0047] S2: Split the legend text into three categories of data: image number, object name, and object number;

[0048] S3: After obtaining the list of artifacts, use the figure number to return the entire text and extract the description information of the artifact.

[0049] Furthermore, the specific method of feature extraction in the information extraction module is as follows:

[0050] S1: The positions of the object number column and description column are specified by the user, and the column containing the original description information of each object is extracted;

[0051] S2: The user sets the features of interest to be exported from the corpus, including continuous features and discrete features;

[0052] S3: The information extraction module performs information matching on continuous features according to user needs, uses regular expressions to capture relevant numeric character information and Chinese unit information, and then merges and outputs the two; at the same time, it uses regular expressions to capture discrete features, assigning a value of 1 for successful matching and 0 for unsuccessful matching, and organizes them into a table to output the final quantified archaeological feature matrix in CSV format.

[0053] In summary, the present invention has the following beneficial effects:

[0054] 1. The information extraction process of this system is fully automated. The entire process is based on the most popular and authoritative information sharing method in the Chinese archaeological community, which greatly improves the efficiency of database construction.

[0055] 2. This system adopts a top-down organizational approach, using an automated extraction process to extract a large amount of information from published archaeological reports as the foundation of the database, while also supporting users to upload new materials.

[0056] 3. As a problem-oriented database platform, this system is the first to integrate a data analysis module into the database framework. At the same time, it has made certain innovations in the analysis methods, embedding analysis functions such as elliptical Fourier profile analysis, BR coefficient calculation, principal component analysis, and Kmeans clustering into the system. This can help users to easily analyze archaeological data and place it in the archaeological big data for research.

[0057] 4. Employing elliptic Fourier analysis minimizes computational load and significantly improves image comparison accuracy with small sample sizes compared to deep learning-based image classification algorithms that require large training sets. Therefore, the image search system in the data storage module will also be based on distance calculation using Fourier coefficients, greatly saving server space and improving retrieval efficiency and accuracy. Attached Figure Description

[0058] Figure 1 This is a flowchart of the information extraction, management, analysis, and integrated database system based on Chinese archaeological reports in this embodiment of the invention.

[0059] Figure 2 This is a flowchart of the information extraction module in an embodiment of the present invention;

[0060] Figure 3 This is a flowchart of the image segmentation process in an embodiment of the present invention;

[0061] Figure 4 This is a flowchart of the data analysis module in an embodiment of the present invention;

[0062] Figure 5 This is a flowchart of the contour standardization process in an embodiment of the present invention;

[0063] Figure 6 This is a structural diagram of the data storage module in an embodiment of the present invention. Detailed Implementation

[0064] The following is in conjunction with the appendix Figure 1-6 The present invention will be described in further detail below.

[0065] Example: A comprehensive database system for information extraction, management, and analysis based on Chinese archaeological reports, including an information extraction module, a data analysis module, and a data storage module. The database is a built-in system database, with supplementary materials including a basic database of archaeological report PDFs and a database of archaeological terminology. Users can utilize these built-in resources to process data, or upload their own data to the system for analysis, enriching the system's built-in datasets, such as... Figure 1 As shown, users can directly upload archaeological reports in PDF format to the information extraction module for information extraction. They can also directly import self-processed, formatted data into the data analysis and data storage modules. Furthermore, users can re-extract and re-analyze data from the data storage module at any time. The information extraction module outputs extracted images and tables to the data storage module, and the data analysis module similarly stores the analyzed images and tables in the data storage module. Users can interactively edit the data and download it as needed. Compared to traditional archaeological database management systems, this embodiment's integrated management and analysis database system, through its information extraction and data analysis modules, extends the database system's function beyond simply storing and sharing data. It opens it up to researchers as a comprehensive data processing, integration, analysis, and research platform.

[0066] The information extraction module of this embodiment is developed based on Python 3.9.16 and incorporates a complete contour detection and image segmentation algorithm based on the format of an archaeological report. Related dependency libraries include OpenCV-Python 4.7.0.72, scikit-image 0.21.0, and imageio 2.31.1 for image segmentation, Streamlit 1.24.1 for creating interactive pages, and PyMuPDF 1.22.5 for reading PDF files. The information extraction module mainly includes three parts: image segmentation, text processing, and discrete / continuous feature extraction. Input information is the original PDF file or a high-quality scanned file of the archaeological report. Output includes a PNG file of artifact and relic line atlases from the archaeological report, a CSV file of artifact information, and a CSV file of a quantified archaeological feature matrix. Users can set relevant parameters for image, text, and feature extraction according to their needs, interactively edit text and table information within the system, and download the processed data.

[0067] The specific method of the information extraction module is as follows:

[0068] 1. Image Segmentation Pipeline: For user-uploaded archaeological report PDF files, each page is converted into an image stored in array format for separate processing. First, the images are converted to grayscale and then to binary images with a black background and white outline. For each binary image, a dilation algorithm is used to expand its outline range to connect gaps caused by PDF clarity or scanning progress issues. Then, small regions in the image are detected based on the connectivity of pixels in the binary image to remove image noise and redundant information such as text, numbers, and symbols (not archaeological images). Next, a closing operation is performed on the extracted images to close the outlines. Finally, edge-based automatic outline detection is performed on the processed images, and the images are automatically cropped using the outermost x and y coordinates of the outline as a rectangular cropping range and output as a compressed package. Users can set the convolution kernel weight matrix of the dilation algorithm and adjust the outline area of ​​small regions according to the actual situation of the report to improve the quality of image segmentation. The output image numbers are user-inputted numbers or report name + page number + image number, allowing for rapid classification and processing in subsequent data analysis. The final output is a text file saved in PNG format, which is then stored in a zip archive for users to download.

[0069] 2. Text processing pipeline: First, import the images and use regular expressions to match the legend text. Split the legend text into three categories: image number, artifact name (e.g., li, II-style pan, etc.), and artifact number (e.g., H:91:3, M18:1-4).

[0070] For image numbers, the commas and tildes indicate characters to be skipped or omitted. For example, images numbered 1, 4-9 need to be expanded to 1, 4, 5, 6, 7, 8, 9 to match each artifact's information. For artifact names, regular expressions are used to locate non-Roman numerals (artifact type numbers) and Chinese characters (artifact names themselves) in the description and assign them to each image number label. Artifact numbers need to be logically broken down according to the archaeological report. Generally, it's: site unit + colon + artifact number + artifact piece number. For example, the second pendant of a complete jade beaded necklace (piece 7) unearthed in ash pit 13 would typically be labeled H13:7-2. Some reports also include stratigraphic numbers or excavation dates before the site unit. This allows the system to accurately separate artifact numbers of different patterns and ultimately match them with image numbers and artifact names for output.

[0071] After obtaining the list of artifacts, the system extracts the descriptive information of each artifact from the entire text using its figure number. Since archaeological report writing rules often list the image number after the description (e.g., (Figure xx, figure number)), the artifact description can be located by combining the figure name and figure number. This system allows users to set the extraction format according to the specific format of their archaeological reports, and users can redefine the text extraction and recognition methods based on the report's writing style. The final output is a CSV format artifact information table, with each row corresponding to one unearthed artifact. The output information includes the name, page number, figure number, and all detailed descriptions in the report for each artifact.

[0072] 3. Discrete / Continuous Feature Extraction Pipeline: This module includes a thousand-word archaeological corpus for information extraction, allowing users to select between continuous and discrete variables and export them directly as matrix tables.

[0073] This module primarily uses algorithms based on regular expressions and natural language processing methods, and incorporates a large-scale archaeological corpus of thousands of words. This corpus effectively matches information from archaeological artifact registration forms and converts it into a data matrix format suitable for large-scale quantitative analysis. The corpus includes information on 38 categories of unearthed artifacts, such as artifact name, measurements, color, material, craftsmanship, decoration, and form. It contains 1006 built-in Chinese words for information matching, which should meet the needs of most archaeological research in my country.

[0074] The information extraction process of the system algorithm is mainly as follows: First, the user specifies the positions of the artifact number column and the description column, and the column containing the original description information of each artifact is extracted. The user then sets the features of interest to be exported from the corpus. Based on the user's selection, the features are first divided into continuous and discrete categories. Continuous features (metric traits) mainly refer to the measurement values ​​describing the artifacts, including length, width, and height; for incomplete artifacts, this also includes residual length, residual width, and residual height. The system can perform information matching on continuous features according to user needs, using regular expressions to extract relevant numeric character information and Chinese unit information, and then merge the two for output. Discrete features (nonmetric traits) mainly refer to binary variables describing the archaeological characteristics of the artifacts. For example, the shape of the artifact's belly can be divided into a series of types such as straight belly, drooping belly, round belly, and folded belly. This system embeds an archaeological database consisting of over 1000 discrete features in 25 categories. The system will automatically match information according to the user's needs and use regular expressions to capture discrete features. If the match is successful, it will be assigned a value of 1, and if the match is unsuccessful, it will be assigned a value of 0. The system will then organize the data into a table and output the final quantified archaeological feature matrix in CSV format.

[0075] The data analysis module in this embodiment is developed based on R4.1.3 and incorporates a complete data analysis workflow based on extracted image and table information. Related dependencies include Momocs 1.4.0 for image contour morphology processing, stats 4.1.3 for K-means clustering, DMwR 0.4.1 for KNN missing value imputation, and the gstat 2.0-9 package for Kriging interpolation. The data analysis module mainly includes three sections: image segmentation, text processing, and discrete / continuous feature extraction.

[0076] 1. Image information preprocessing

[0077] 1.1 Contour Standardization

[0078] For the input image information, the binarized PNG image of the artifact output by the information extraction module is converted to JPG format. Then, a gap-filling algorithm is used to fill the interior of the image to improve the accuracy of contour recognition. Finally, the processed images are batch-read into the image processing module based on R language. Using the JPG image parser built into the momocs package, the system automatically re-identifies the outer contour of the image after the user sets a threshold. The identified outer contour is the x and y coordinates of each pixel. After obtaining the coordinates, the contour area is calculated, and all contours are then standardized according to their area. Based on the characteristics of archaeological artifacts, the system automatically identifies the pixel where the midline of the contour intersects with the lowest edge as the starting point of the contour. Figure 5 Point 1 in 3) and calculate the x and y coordinate deviations of each pixel from the previous pixel in turn for subsequent elliptic Fourier transform. Figure 5 )

[0079] 1.2 Elliptic Fourier Analysis

[0080] Elliptic Fourier analysis is a Fourier method for fitting complex closed contours. The Fourier coefficients can be independent of the contour location and their magnitudes can be standardized. Let T be the perimeter of the contour, which becomes the period of the signal. A setting... Let it be a pulse. Then its formula is as follows.

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087] This method increases the number of coefficients through harmonics, allowing for more accurate contour description with fewer harmonics compared to other contour extraction methods. The system defaults to describing the archaeological report line drawing contour with 20 harmonics, therefore the output image will be converted to include a1, a2...a... n b1, b2...b n c1, c2...c n d1, d2...d n The vector consists of 80 coefficients. Based on this vector, a complete contour can be reconstructed through inverse transformation, containing all the original information of the contour for subsequent analysis. Compared to direct analysis based on pixel data, converting a line drawing with less noise directly into a contour, although losing some information about the surface and cross-section of the artifact, is beneficial. Considering that differences in archaeological artifacts are often reflected solely by their outer contours, and that traditional archaeological classification of artifacts is largely based on their shape, this method minimizes computational load and significantly improves the accuracy of image comparison with small sample sizes compared to deep learning-based image classification algorithms that require large training sets. Therefore, the image search system in the data storage module will also be based on distance calculations using Fourier coefficients, greatly saving server space and improving retrieval efficiency and accuracy.

[0088] 2. Text Information Preprocessing

[0089] For the input text information, the quantitative archaeological feature matrix output by the information extraction module is input into the data analysis module. First, the user selects the variables to be analyzed. Before starting the analysis, similar features need to be merged based on the user's analytical needs to improve statistical efficiency. Firstly, in archaeological fieldwork, different standards often result in different synonymous descriptions; for example, the terms "round belly" and "bulging belly" often describe the same type of feature on certain artifacts. Secondly, in analytical work, different analytical resolutions often require manual adjustment of these features. For example, descriptions of reddish-brown, reddish-brown, and red should be merged when analyzed alongside gray and black, but treated as different variables when comparing within the red category. After merging similar terms, the merged data is cleaned to filter out features with insufficient information for analysis. The user can set a filtering threshold to remove features with a frequency below a specific threshold. After removal, the user can choose to use the original data directly for analysis; the system also has a built-in k-nearest neighbor data imputation method to fill in missing values. The k-nearest neighbor algorithm uses the k nearest neighbors to fill in unknown (NA) values ​​in the dataset. For each case with any NA value, it searches for its k most similar cases and uses the values ​​of those cases to fill in the unknowns. The final output is a preprocessed feature matrix.

[0090] 3. Analysis at the single site level

[0091] The quantitative archaeological data analysis module at the single-site level is mainly used to help users process information within a single site. It includes K-means clustering for exploring the relationships between artifacts, correlation analysis for exploring the relationships between features, and principal component analysis for exploring the relationships between artifacts and features.

[0092] 3.1 K-means Clustering - Relationships between Relics

[0093] K-means clustering is an unsupervised classification method widely used in archaeological classification research. In traditional field archaeology database systems, artifact classification is often based on user-inputted background information. This system utilizes the K-means classification method to provide users with a way to classify artifacts unearthed from a site as a whole. It can perform unsupervised classification of artifacts based on the contour information of the Fourier coefficients after elliptic Fourier transform, helping users quickly organize and classify large amounts of archaeological material. The K-means algorithm's processing flow is as follows: 1) Randomly assign the Fourier coefficients of all artifacts to K non-empty classes; 2) Calculate the mean of each class and use this mean to represent the corresponding class; 3) Reassign each object to the nearest class based on its distance from the center of each class; 4) Return to step 2) until the criterion function converges. The above analysis method is implemented based on the `fviz_nbclust` function in the R language. Users can use the system's automatically calculated optimal `k` or manually input `k` to determine the final number of classifications.

[0094] 3.2 Correlation Analysis - Relationships between Features

[0095] To explore the correlations between discrete features, this system provides users with correlation analysis functions. For discrete features versus discrete features, the Fisher exact test is available to examine the correlation between them. For discrete features versus continuous features, Point-biserial correlation coefficient calculation and statistical testing are available. For continuous features versus continuous features, Pearson correlation coefficient calculation and statistical testing are available. Users select variables of interest in the quantified archaeological feature matrix, and the system can automatically identify the variable type and select the corresponding statistical method, outputting statistics reflecting the strength of the correlation and exact p-values.

[0096] 3.3 Principal Component Analysis - Relationship between Relics and Features

[0097] Because the coefficients of the quantified archaeological feature matrix and images transformed by Fourier transform have high dimensionality, practical research often requires combining the analysis of image and discrete feature information. Therefore, this system provides a function to reduce the dimensionality of images and discrete features based on principal component analysis (PCA), and calculate feature weights based on the variance explained by the principal components, effectively evaluating the classification efficiency of different archaeological features. PCA is a commonly used multivariate statistical method, introduced into archaeological statistical analysis in the 1980s and widely applied in pottery typology research based on measurement data. PCA selects the linear combination with the largest variance to reduce the dimensionality of multivariate data; this linear combination is called the first principal component, and so on. To help users determine the contribution of different archaeological feature descriptions to artifact classification, the weights of each feature are further calculated based on the principal components obtained from PCA. The weight of a feature is defined as the product of its maximum absolute loading (weight) among all principal components and the variance explained by the specific PC corresponding to that maximum loading. By dividing by the original weights and then standardizing the weights of all coordinates, a relative weight value ranging from 0 to 1 is finally obtained for comparison. Generally speaking, some archaeological features will show a stronger signal than others. Quantifying these features can help users filter and refine their descriptions of artifacts and quickly identify artifacts with special cultural significance within a site.

[0098] 4. Multi-site level analysis

[0099] The multi-site quantitative archaeological data analysis module primarily assists users in making horizontal comparisons between sites. Users can upload images and text information about their archaeological sites and extract relevant information. In this module, their archaeological data is matched against the system's built-in large dataset to obtain a comprehensive assessment of the site's condition. In previous studies, this work often required significant effort and expertise from researchers. This system introduces a fully automated site analysis method that effectively calculates the cultural similarity between a site and its surrounding sites, and can estimate the proportion of cultural elements at a site based on user-defined comparative cultures. Ultimately, this helps users explore cultural relationships between sites at a macro level.

[0100] 4.1 Analysis of Cultural Factors

[0101] For image datasets of archaeological sites, this system can help users analyze the cultural composition of a particular site. The specific process is as follows: First, the user selects a group of artifacts from a site that has typical cultural significance and is well-preserved for verification. These artifacts are generally those that have been restored and are relatively complete after being unearthed from the entire site. Missing artifacts should be removed at this step to improve the accuracy of identification. Next, the user selects archaeological cultures from the database system for comparison. Archaeological cultures are generally a series of sites with similar cultural features. The system includes several clusters of Neolithic archaeological cultures with a large number of sites. After the user completes the selection, the system will determine the cultural affiliation of each artifact in the user-defined site artifact group. The specific process is as follows: First, the distance between the coefficients of the artifact image after elliptic Fourier transform and all artifacts unearthed from archaeological culture sites is calculated sequentially. Different distance calculation strategies, such as Euclidean distance, maximum distance, minimum distance, and Manhattan distance, can be used based on user settings. Then, all distances are sorted, and cultural attributes are determined based on the proportion threshold set by the user. For example, the user can set the cultural affiliation of the artifact to the cultural affiliation of the closest 10% of the artifacts as the cultural attribute of that artifact.

[0102] During the identification process, both discrete and continuous strategies can be used for accumulation. For example, if an artifact unearthed from a site is compared with 100 artifacts from a comparative culture, and the 10 artifacts (10%) with the closest distance (i.e., the most similar morphology) are found to include 3 from the Longshan culture and 7 from the Liangzhu culture, using a discrete classification method, the Liangzhu culture with the highest frequency among the 10 artifacts will be used as the artifact's cultural affiliation. Alternatively, a continuous classification method can be used, where cultural affiliation is included as a score in the overall score; in this case, the artifact would be recorded as 0.3 Longshan culture + 0.7 Liangzhu culture. Finally, the system will perform the above operations on all artifacts in the group, accumulate all cultural identification results, calculate the percentage, and output a bar chart showing the percentage of cultural factors for the artifact group from that site.

[0103] 4.2 BR Coefficient Calculation

[0104] This system helps users quickly obtain the cultural similarity between a site and surrounding sites by providing a quantitative archaeological feature matrix, and visually displays the cultural geography of the site in the form of a heat map. This function is accomplished through BR coefficient calculation and Kriging interpolation. The Brainerd-Robinson coefficient (BR coefficient) is the most commonly used indicator in archaeological research to assess the similarity of artifact composition between two archaeological sites. The BR coefficient first calculates the proportion of each specific pottery type in the total amount of pottery unearthed at the two sites, and then sums the absolute values ​​of the difference between the percentages of the two sites. The corrected BR coefficient ranges from [0,1]. For example, in site A, type I pottery accounts for 45% and type II pottery accounts for 55%, while in site B, type I pottery accounts for 35%, type II pottery accounts for 20%, and type III pottery accounts for 45%. Then its BR coefficient should be 1 - ((0.45 - 0.35) + (0.55 - 0.20) + (0.45 - 0)) / 2 = 0.125. Since pottery is generally used as the classification standard for BR coefficient calculation, users need to specify the categories of artifacts they are interested in. Artifact classification can use the unsupervised K-means classification results from the previous analysis or user-inputted classification labels. Users also need to select comparison sites. After selection, the system automatically calculates the relative abundance of each artifact category for the target site and all comparison sites, and performs BR coefficient calculations. The final output is the BR coefficient vector for each comparison site and the target site. If the user selects spatial analysis and inputs site coordinates, this BR coefficient vector can be output as a heatmap on a map using Kriging interpolation, visually demonstrating the degree of cultural similarity between different regions and the target site from a point-to-area perspective.

[0105] In this embodiment, the data storage module is independent of the information extraction module and the data analysis module. Data storage includes the portion of data that has undergone information extraction and analysis on the user's local machine, and a cloud database available for download and comparison by the user. Both types of data are categorized and stored using a three-level directory index.

[0106] 1. First-level directory

[0107] The first-level directory is the site directory, which records relevant information for each site. This information includes the site's geographical location, such as province, city, county, and village; and its geographical coordinates, such as 120°21'E, 40°14'N. It also includes the site's chronological information, which is often divided into two categories: absolute chronology, storing all carbon-14 dating data for the site; and relative chronology, indicating the archaeological cultural period of the site. The directory also includes information on the archaeological cultural attributes, allowing users to set and adjust relevant cultural classifications according to the research focus. Each folder in the first-level directory corresponds to a folder named after the site, containing the above information.

[0108] 2. Second-level directory

[0109] The second-level directory is the site directory, which records relevant information about each site within each site. The information mainly includes the site type, such as ash pits, tombs, house foundations, or excavation squares, trenches, etc.; stratigraphic relationships, mainly including the stratigraphic location of the site and its superposition and disruption relationships, etc.; and relative ages inferred from the carbon-14 dating of artifacts in the site or archaeological reports.

[0110] 3. Third-level directory

[0111] The third-level directory is the artifact catalog, recording relevant information for each artifact unearthed from each site. The image information includes standard artifact line drawings from the archaeological report, along with photographs and other images, as well as Fourier coefficients of the line drawing outlines obtained through the information extraction and data analysis module. The discrete feature information includes complete descriptions directly extracted from the archaeological report, and discrete feature matrices obtained through the information extraction module. The continuous feature information mainly consists of continuous feature matrices representing various measurements of the artifacts.

[0112] 4. Data Search

[0113] Data search requires users to first define the search scope. The system will then automatically match the input string within the corresponding directory and list relevant files. In the third-level directory, both image search and character search can be used simultaneously; character search works the same way. If the user chooses image search, a line drawing image must be entered. The system will automatically process this image in the information extraction workflow, extracting its outline and performing an elliptic Fourier transform to obtain Fourier coefficients. These coefficients are then compared with all other Fourier coefficients in the system, and n similar images (user-defined) are displayed. This innovative image search scheme significantly improves the accuracy of archaeologists' comparison of artifact information, making it possible to find similar images based on archaeological line drawings.

[0114] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.

Claims

1. An information extraction management analysis and comprehensive database system based on Chinese archaeological reports, characterized in that: The information extraction module, the data analysis module and the data storage module are included. ​ The information extraction module is used for picture segmentation, text processing and feature extraction. The data analysis module is used for information preprocessing and information analysis. The data storage module is used for classified storage of information analyzed by the data analysis module and information extracted by the information extraction module and can perform data search. The classified storage adopts a form of three-level directory index. The information preprocessing in the data analysis module includes picture information preprocessing and text information preprocessing. The information analysis includes elliptical Fourier analysis, single-site layer analysis and multi-site layer analysis. The single-site layer analysis includes artifact-artifact relationship analysis, feature-feature relationship analysis and artifact-feature relationship analysis. The artifact-artifact relationship analysis mode is that the K-means clustering algorithm is used to classify the unearthed artifacts in the site. The feature-feature relationship analysis mode is that the correlation between discrete features and continuous features is obtained by using the fisher exact test function to test the correlation between the two, the point-biserial correlation coefficient calculation and statistical test option and the Pearson correlation coefficient calculation and statistical test option. The artifact-feature relationship analysis mode is that the principal component analysis is used to reduce the dimension of the picture and the discrete feature, and the feature weight is calculated according to the principal component variance explanation degree to evaluate the classification efficiency of different archaeological features. The multi-site layer analysis includes cultural factor analysis and BR coefficient calculation. The cultural factor analysis analysis mode is that a typical cultural and well-preserved object group in a site is selected for testing, and missing objects are excluded, then an archaeological culture is selected from the database system for comparison, and finally each object in the set object group of the site is judged for cultural attribution. The BR coefficient calculation is used to help users quickly obtain the cultural similarity between the site and the surrounding sites, and intuitively display the cultural geographical relationship of the site in the form of a heat map.

2. The Chinese-based archaeological report information extraction management analysis and synthesis database system according to claim 1, characterized in that: The specific method of the picture information preprocessing is: S1: After converting the binary artifact line graph in png format output by the information extraction module into jpg format, the line graph is filled inside by using the hole filling algorithm to improve the contour recognition accuracy, and finally the processed pictures are read into the picture processing block based on R language; S2: The user sets the threshold value, and the picture parser automatically recognizes the outer contour of the picture. The recognized outer contour is the x, y coordinates of each pixel point; S3: After obtaining the coordinates, the contour area is calculated, and all contours are standardized according to the area size.

3. The Chinese-based archaeological report information extraction management analysis and synthesis database system according to claim 1, characterized in that: The specific method of the text information preprocessing is: S1: Similar quantitative archaeological features output by the information extraction module are merged based on the user's analysis needs; S2: The merged data is cleaned.

4. The Chinese-based archaeological report information extraction management analysis and synthesis database system according to claim 3, characterized in that: The specific way of data cleaning in S2 is to filter out features with little information that cannot be used for analysis.

5. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 3, characterized in that: The specific way of data cleaning in S2 is: by setting a filtering threshold, features with a frequency lower than the set threshold are removed; after removal, the original data is directly used for analysis, and then the k-nearest neighbor data interpolation method is used to fill in the missing values.

6. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 1, characterized in that: The specific way of cultural attribution is: first, calculate the distance between the coefficients of the elliptical Fourier transform of the artifact picture and all the unearthed artifacts of the archaeological cultural sites; then sort all the distances, and determine the cultural attribute based on the set proportion threshold.

7. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 1, characterized in that: BR The specific steps of coefficient calculation are: first, calculate the percentage of each specific pottery type in the total amount of unearthed pottery in two sites; then sum the absolute value of the difference between the percentages of the two sites, and the interval of the corrected BR coefficient is [0, 1]; then the user classifies the artifacts of interest and selects the comparison sites to be compared; after the selection is completed, the data analysis module calculates the relative abundance of each type of artifact in the target site and all comparison sites, and calculates the BR coefficient, and finally outputs the BR coefficient vector of each comparison site and the target site.

8. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 1, characterized in that: The specific way of picture segmentation in the information extraction module is: S1: Convert each page of the archaeological report pdf file to a picture stored in array format for separate processing; S2: Convert the picture to a grayscale picture and convert it to a binary picture with a black background and white outline; S3: For each binary picture, use the dilation algorithm to expand its outline range to link the pores caused by pdf clarity or scanning accuracy problems; S4: Detect small areas in the picture based on the connectivity of binary picture pixels to remove picture noise and redundant information; S5: Perform a closed operation on the extracted picture to close the outline; S6: Perform automatic edge-based outline detection on the processed picture, and use the outermost x and y coordinates of the outline as the rectangular cropping range to automatically cut the picture and output it in a compressed package.

9. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 1, characterized in that: The specific way of text processing in the information extraction module is: S1: Import the picture and use regular expressions to match the legend text in it; S2: Split the legend text into three types of data: picture number, artifact name, and artifact number; S3: After obtaining the artifact list, use the figure number to extract the description information of the artifact in the entire text.

10. The Chinese archaeological report information extraction management analysis synthesis database system according to claim 1, characterized in that: The specific way of feature extraction in the information extraction module is: S1: The user specifies the positions of the artifact number column and the description column to extract the column containing the original description information of each artifact; S2: The user sets the features of interest in the corpus that need to be exported, including continuous features and discrete features; S3: The information extraction module matches the continuous features according to the user's needs, uses regular expressions to grab relevant numerical character information and Chinese unit information, and then merges them for output; at the same time, regular expressions are used to grab discrete features, and if the match is successful, assign a value of 1, and if the match is not successful, assign a value of 0, and then arrange them into a table to output the final quantitative archaeological feature matrix in csv format.

Citation Information

Patent Citations

  • Archaeological excavation data based relic site early terrain three-dimensional reconstruction method

    CN108009314A

  • A method for accurately relocate an early archaeological excavation site

    CN109241222A